Code quality skillCommunity

Verification Before Completion Skill

A Superpowers skill that stops your AI agent saying "done" or "tests pass" until it has run the command that proves it and read the output.

Works in
  • Claude Code
  • Codex
  • Cursor
  • Gemini CLI

By Jesse Vincent (obra) · obra/superpowers · MIT

Last updated

Install

Run

npx skills add obra/superpowers --skill verification-before-completion

This installs the repo's latest version. We reviewed commit 8ca22db; use the Manual tab to install exactly that.

What it can touch

Runs scripts
No. Instructions only.
Needs network
No.
allowed-tools
Not set. It asks for no extra tool permissions; your tool's usual prompts apply.
Licence
MIT
Last reviewed
How we check a skill is safe
Check it yourself

Agent skills checklist

0 of 10 checked

Who made it

Before reading a line of it, know whose code you are about to run.

What's inside

The part people skip. Read what your agent will read.

What it can reach

Give it the least access that still does the job.

Keeping it that way

What you checked today is only what runs tomorrow if you pin it.

Your ticks are saved in this browser only.

AI agents are quick to say "Done! All tests pass." Sometimes they ran the tests. Sometimes they ran them before the last change, or ran the linter and assumed the rest. This skill from Jesse Vincent's Superpowers collection has one rule: no claim that work is finished, fixed or passing until the agent has run the command that proves it, in the same message, and read what it printed.

What it does

Before any claim about the state of the work, the agent goes through five steps: work out which command would prove the claim, run all of it, read the full output and exit code, check that the output actually says what it's about to claim, and only then say it, with the evidence.

The skill spells out what counts as proof:

ClaimProofNot proof
Tests passTest output showing 0 failuresAn earlier run, "should pass"
Build succeedsThe build command exiting 0The linter passing
Bug fixedThe original symptom, now gone"Code changed, assumed fixed"
Sub-agent finishedThe diff showing its changesThe sub-agent saying so

It also lists warning signs in the agent's own wording, like "should", "seems to", or "Great!" before anything has been checked. The core of it is one line:

If you haven't run the verification command in this message, you cannot claim it passes.

From SKILL.md by Jesse Vincent, MIT.

An example

You ask the agent to fix a failing date test. It edits the code. Without the skill, the next message is often "Fixed! The test should pass now." With it, the next message runs the test, shows 12 passed, 0 failed, then runs the whole suite because the change touched a shared helper, and only then says it's fixed, quoting both results.

When to use it, and when not to

Use it everywhere you let an agent change code, and especially before commits and pull requests. It costs a few seconds per task and catches the most common way agents mislead people.

There's little reason to leave it out, but if your project has no tests, no build and nothing to run, it has nothing to check with.

What we checked

We read every file in the skill's folder at the commit linked above: just SKILL.md. There are no scripts, no network access and no allowed-tools, and nothing asks the agent to skip a confirmation or read anything outside your project. The Superpowers repo is MIT licensed and updated most weeks. Our full checklist is in the guide to checking a skill before you install it.

It works well with the Karpathy guidelines skill, which asks the agent to define what "done" means before it starts.

FAQ

Do I need the rest of Superpowers to use it?

No. It's one self-contained SKILL.md. Superpowers is a larger set of skills by the same author, installed as a plugin, and this is one of them, but it doesn't depend on the others.

Why does it matter if my agent says "done" too early?

Because you believe it. If an agent says the tests pass without running them, you commit broken code and find out later, often from someone else. The skill makes it show you the output that backs up the claim.

Does it run my tests for me?

It makes the agent run whatever command proves the claim, using the tools your agent already has. It has no scripts of its own. If your agent needs permission to run commands, you'll still be asked.

It reads quite strict. Is that a problem?

It's written that way on purpose: it lists the excuses an agent might give for skipping a check and answers each one. If it's too much for your taste, you can edit your copy, though the strictness is what makes it work.

See all skills