Diagnosing Bugs
Diagnosis loop for hard bugs and performance regressions. Use when the user says 'diagnose'/ 'debug this', or reports something broken/throwing/failing/slow.
Diagnosing Bugs
A discipline for hard bugs. Skip phases only when explicitly justified.
Phase 1 — Build a feedback loop
This is the skill. Everything else is mechanical. If you have a tight pass/fail signal for the bug — one that goes red on this bug — you will find the cause.
Spend disproportionate effort here. Be aggressive. Be creative. Refuse to give up.
Ways to construct one — try them in roughly this order
- Failing test at whatever seam reaches the bug — unit, integration, e2e.
- Curl / HTTP script against a running dev server.
- CLI invocation with a fixture input, diffing stdout against a known-good snapshot.
- Headless browser script (Playwright / Puppeteer) — drives the UI, asserts on DOM/console/network.
- Replay a captured trace. Save a real network request / payload / event log to disk; replay it through the code path in isolation.
- Throwaway harness. Spin up a minimal subset of the system that exercises the bug code path with a single function call.
- Property / fuzz loop. If the bug is "sometimes wrong output", run 1000 random inputs and look for the failure mode.
- Bisection harness. If the bug appeared between two known states, automate "boot at state X, check, repeat" so you can
git bisect runit. - Differential loop. Run the same input through old-version vs new-version and diff outputs.
Tighten the loop
Treat the loop as a product. Once you have a loop, tighten it:
- Can I make it faster?
- Can I make the signal sharper?
- Can I make it more deterministic?
Phase 2 — Reproduce + minimise
Run the loop. Watch it go red — the bug appears.
Confirm:
- The loop produces the failure mode the user described
- The failure is reproducible across multiple runs
- You have captured the exact symptom
Minimise
Once it's red, shrink the repro to the smallest scenario that still goes red. Cut inputs, callers, config, data, and steps one at a time.
Phase 3 — Hypothesise
Generate 3–5 ranked hypotheses before testing any of them.
Each hypothesis must be falsifiable: state the prediction it makes.
Format: "If <X> is the cause, then <changing Y> will make the bug disappear / <changing Z> will make it worse."
Show the ranked list to the user before testing.
Phase 4 — Instrument
Each probe must map to a specific prediction from Phase 3. Change one variable at a time.
Tool preference:
- Debugger / REPL inspection if the env supports it.
- Targeted logs at the boundaries that distinguish hypotheses.
- Never "log everything and grep".
Tag every debug log with a unique prefix, e.g. [DEBUG-a4f2].
Phase 5 — Fix + regression test
Write the regression test before the fix — but only if there is a correct seam for it.
If a correct seam exists:
- Turn the minimised repro into a failing test at that seam.
- Watch it fail.
- Apply the fix.
- Watch it pass.
- Re-run the Phase 1 feedback loop against the original scenario.
Phase 6 — Cleanup + post-mortem
Required before declaring done:
- Original repro no longer reproduces
- Regression test passes
- All
[DEBUG-...]instrumentation removed - Throwaway prototypes deleted
- The hypothesis that turned out correct is stated in the commit / PR message