Code quality skillCommunity

Diagnosing Bugs Skill

Matt Pocock's debugging skill for hard bugs: your AI agent builds a command that reproduces the failure before it guesses at any cause.

Works in
  • Claude Code
  • Codex

By Matt Pocock · mattpocock/skills · MIT

Last updated

Install

Run

npx skills add mattpocock/skills --skill diagnosing-bugs

This installs the repo's latest version. We reviewed commit c55ee46; use the Manual tab to install exactly that.

What it can touch

Runs scripts
Yes. It ships scripts your agent can run, so read them before installing.
Needs network
No.
allowed-tools
Not set. It asks for no extra tool permissions; your tool's usual prompts apply.
Licence
MIT
Last reviewed
How we check a skill is safe
Check it yourself

Agent skills checklist

0 of 10 checked

Who made it

Before reading a line of it, know whose code you are about to run.

What's inside

The part people skip. Read what your agent will read.

What it can reach

Give it the least access that still does the job.

Keeping it that way

What you checked today is only what runs tomorrow if you pin it.

Your ticks are saved in this browser only.

Most debugging with an AI agent goes like this: it reads the error, guesses a cause, changes something, and asks you to try again. Matt Pocock's diagnosing bugs skill stops that guessing. Before the agent may form a single theory, it has to build a command that reproduces the bug on demand. Everything after that, from narrowing the cause to proving the fix, runs against that command.

What it does

Six phases, in order:

  1. Build a feedback loop. A failing test, a curl script, a headless browser script, a replayed request, a git bisect run harness. Then make it faster and more reliable. The skill calls this phase "the skill".
  2. Reproduce and minimise. Confirm it fails the way you described, then cut the repro down until every remaining piece is needed.
  3. Hypothesise. Three to five ranked causes, each with a prediction that could prove it wrong. It shows you the list before testing it.
  4. Instrument. One change at a time. Every debug log gets a tag like [DEBUG-a4f2] so they can all be found and removed afterwards.
  5. Fix, with a regression test, written before the fix, where there's a sensible place for it.
  6. Clean up. Rerun the original repro, remove the tagged logs, and write down which theory was right.

The skill is blunt about why phase 1 comes first:

If you catch yourself reading code to build a theory before this command exists, stop: jumping straight to a hypothesis is the exact failure this skill prevents.

From SKILL.md by Matt Pocock, MIT.

An example

You report: "The export button sometimes produces an empty CSV." The agent writes a small script that calls the export endpoint 50 times and counts empty files, and runs it: 7 out of 50 are empty. Only then does it list causes (a race between two queries, a timeout, a cache), each with the change that would prove it. It tests the race first because the script makes it cheap to check.

When to use it, and when not to

Use it on bugs that have already cost you time: ones that come and go, ones that only happen with real data, and slowdowns.

For a plain error with an obvious line number, it's more process than you need. If you're learning to debug and want the agent to explain each step as it goes, coding.kitty's own debug with me skill (coming soon) is built for that. Pick one of the two: they answer to the same requests.

What we checked

We read every file in the folder at the commit linked above: SKILL.md, agents/openai.yaml (a display name for Codex) and scripts/hitl-loop.template.sh. There are no allowed-tools. The script is a short bash file with two helpers: one shows you an instruction and waits for Enter, the other asks a question and prints your answer back for the agent to read. It doesn't download or send anything; its example step opens localhost:3000. The skill itself only reaches your own dev server or test environment, and it tells the agent to redact secrets before showing any output. The repo is MIT licensed. Our full checklist is in the guide to checking a skill before you install it.

FAQ

What is a feedback loop, in debugging?

A command you can run that goes red while the bug is there and green once it's fixed: a failing test, a curl against your dev server, a script that diffs output. The skill spends most of its effort building one, because with it every later step is quick to check.

Does it work for performance problems?

Yes. For a slowdown it measures first (a timing harness, a profiler, a query plan) and bisects from there, instead of adding logs. "Performance regressions" is in its description for that reason.

What's the script in the skill folder?

A bash template for bugs only a person can reproduce, like one that needs clicking through a UI. The agent copies it, fills in the steps, and runs it; you follow the prompts in your terminal and type what you saw, and the answers go back to the agent.

Will it paste my secrets into the chat?

It's told not to. The skill asks the agent to replace every secret with <REDACTED> before showing commands or output, to build loops around environment variables, and to quote only the lines of a captured request that matter.

See all skills