Most debugging with an AI agent goes like this: it reads the error, guesses a cause, changes something, and asks you to try again. Matt Pocock's diagnosing bugs skill stops that guessing. Before the agent may form a single theory, it has to build a command that reproduces the bug on demand. Everything after that, from narrowing the cause to proving the fix, runs against that command.
What it does
Six phases, in order:
- Build a feedback loop. A failing test, a curl script, a headless
browser script, a replayed request, a
git bisect runharness. Then make it faster and more reliable. The skill calls this phase "the skill". - Reproduce and minimise. Confirm it fails the way you described, then cut the repro down until every remaining piece is needed.
- Hypothesise. Three to five ranked causes, each with a prediction that could prove it wrong. It shows you the list before testing it.
- Instrument. One change at a time. Every debug log gets a tag like
[DEBUG-a4f2]so they can all be found and removed afterwards. - Fix, with a regression test, written before the fix, where there's a sensible place for it.
- Clean up. Rerun the original repro, remove the tagged logs, and write down which theory was right.
The skill is blunt about why phase 1 comes first:
If you catch yourself reading code to build a theory before this command exists, stop: jumping straight to a hypothesis is the exact failure this skill prevents.
From SKILL.md by Matt Pocock, MIT.
An example
You report: "The export button sometimes produces an empty CSV." The agent writes a small script that calls the export endpoint 50 times and counts empty files, and runs it: 7 out of 50 are empty. Only then does it list causes (a race between two queries, a timeout, a cache), each with the change that would prove it. It tests the race first because the script makes it cheap to check.
When to use it, and when not to
Use it on bugs that have already cost you time: ones that come and go, ones that only happen with real data, and slowdowns.
For a plain error with an obvious line number, it's more process than you need. If you're learning to debug and want the agent to explain each step as it goes, coding.kitty's own debug with me skill (coming soon) is built for that. Pick one of the two: they answer to the same requests.
What we checked
We read every file in the folder at the commit linked above: SKILL.md,
agents/openai.yaml (a display name for Codex) and
scripts/hitl-loop.template.sh. There are no allowed-tools. The script is a
short bash file with two helpers: one shows you an instruction and waits for
Enter, the other asks a question and prints your answer back for the agent to
read. It doesn't download or send anything; its example step opens
localhost:3000. The skill itself only reaches your own dev server or test
environment, and it tells the agent to redact secrets before showing any
output. The repo is MIT licensed. Our full checklist is in the
guide to checking a skill before you install it.