mcpskills.net
SkillsMCPsAgentsPrompts
mcpskills.net — A curated directory of AI agent Skills and MCP servers
TermsPrivacy
← Back to Skills
Engineering

Diagnosing Bugs

Diagnosis loop for hard bugs and performance regressions. Use when the user says 'diagnose'/ 'debug this', or reports something broken/throwing/failing/slow.

by Matt PocockRepository →Source →

Diagnosing Bugs

A discipline for hard bugs. Skip phases only when explicitly justified.

Phase 1 — Build a feedback loop

This is the skill. Everything else is mechanical. If you have a tight pass/fail signal for the bug — one that goes red on this bug — you will find the cause.

Spend disproportionate effort here. Be aggressive. Be creative. Refuse to give up.

Ways to construct one — try them in roughly this order

  1. Failing test at whatever seam reaches the bug — unit, integration, e2e.
  2. Curl / HTTP script against a running dev server.
  3. CLI invocation with a fixture input, diffing stdout against a known-good snapshot.
  4. Headless browser script (Playwright / Puppeteer) — drives the UI, asserts on DOM/console/network.
  5. Replay a captured trace. Save a real network request / payload / event log to disk; replay it through the code path in isolation.
  6. Throwaway harness. Spin up a minimal subset of the system that exercises the bug code path with a single function call.
  7. Property / fuzz loop. If the bug is "sometimes wrong output", run 1000 random inputs and look for the failure mode.
  8. Bisection harness. If the bug appeared between two known states, automate "boot at state X, check, repeat" so you can git bisect run it.
  9. Differential loop. Run the same input through old-version vs new-version and diff outputs.

Tighten the loop

Treat the loop as a product. Once you have a loop, tighten it:

  • Can I make it faster?
  • Can I make the signal sharper?
  • Can I make it more deterministic?

Phase 2 — Reproduce + minimise

Run the loop. Watch it go red — the bug appears.

Confirm:

  • The loop produces the failure mode the user described
  • The failure is reproducible across multiple runs
  • You have captured the exact symptom

Minimise

Once it's red, shrink the repro to the smallest scenario that still goes red. Cut inputs, callers, config, data, and steps one at a time.

Phase 3 — Hypothesise

Generate 3–5 ranked hypotheses before testing any of them.

Each hypothesis must be falsifiable: state the prediction it makes.

Format: "If <X> is the cause, then <changing Y> will make the bug disappear / <changing Z> will make it worse."

Show the ranked list to the user before testing.

Phase 4 — Instrument

Each probe must map to a specific prediction from Phase 3. Change one variable at a time.

Tool preference:

  1. Debugger / REPL inspection if the env supports it.
  2. Targeted logs at the boundaries that distinguish hypotheses.
  3. Never "log everything and grep".

Tag every debug log with a unique prefix, e.g. [DEBUG-a4f2].

Phase 5 — Fix + regression test

Write the regression test before the fix — but only if there is a correct seam for it.

If a correct seam exists:

  1. Turn the minimised repro into a failing test at that seam.
  2. Watch it fail.
  3. Apply the fix.
  4. Watch it pass.
  5. Re-run the Phase 1 feedback loop against the original scenario.

Phase 6 — Cleanup + post-mortem

Required before declaring done:

  • Original repro no longer reproduces
  • Regression test passes
  • All [DEBUG-...] instrumentation removed
  • Throwaway prototypes deleted
  • The hypothesis that turned out correct is stated in the commit / PR message