Skill creator is Anthropic's skill for making skills. You tell your agent what
you want a skill to do, and it takes you through writing the SKILL.md,
testing it on real prompts with and without the skill, reviewing the results
with you in a browser, and improving it. It can also tune the skill's
description so your agent loads it when it should.
What it does
The loop it runs has five steps:
- Capture what you want. What the skill should do, when it should load, what good output looks like. If you've just done the task in the chat, it pulls the steps from there.
- Write the draft. A
SKILL.mdwith a name and a description, plus any scripts or reference files, kept under about 500 lines. - Test it. Two or three realistic prompts, each run once with the skill and once without, so you can compare.
- Review with you. A local viewer opens in your browser with the outputs side by side, pass rates, time and tokens, and a box for your comments.
- Improve and repeat. It reads your comments and rewrites the skill, then runs the tests again.
It also explains its own approach to writing instructions, which is worth reading whether or not you use it:
Try to explain to the model why things are important in lieu of heavy-handed musty MUSTs.
From SKILL.md by Anthropic, Apache 2.0.
An example
You say: "Turn what we just did into a skill for writing release notes." It reads back the steps from your chat, asks what should trigger it ("write release notes", "changelog for v2"), drafts the skill, and proposes three test prompts for you to approve. After the runs, you see the notes it wrote with the skill next to the ones it wrote without, and you comment on each.
When to use it, and when not to
Use it when you're writing a skill you'll rely on, or improving one that loads at the wrong times. The with-and-without comparison is the quickest way to find out whether a skill helps at all.
For a five-line personal skill, it's more process than you need; write the
SKILL.md by hand.
What we checked
We read every file in the folder at the commit linked above: SKILL.md, three
sub-agent briefs in agents/, references/schemas.md, the eval viewer
(eval-viewer/generate_review.py and viewer.html), the HTML template in
assets/, the Python files in scripts/, and the Apache 2.0
LICENSE.txt. There are no allowed-tools. What it can do:
- Runs scripts. Python for benchmarks, reports, packaging and validation.
run_eval.pyandimprove_description.pystartclaude -pas a child process. For each test query,run_eval.pywrites a temporary file to your project's.claude/commands/and deletes it afterwards. - Uses the network.
claude -ptalks to Anthropic's API on your account. The viewer loads fonts from Google Fonts and a spreadsheet library from cdn.sheetjs.com, the latter pinned with a hash. - Runs a local server. The viewer serves on
127.0.0.1port 3117. Before it starts, it stops whatever is already listening on that port (it useslsofto find it), so don't run it while something you care about is using 3117.
None of that is hidden, but it isn't confined to the skill's workspace either.
run_eval.py launches Claude from your project root with your full process
environment, so that child session can see your project and any environment
variables you have set, and it writes its temporary command under
.claude/commands/. It's more than most skills do, so read the scripts
yourself before running it. Our full checklist is in the
guide to checking a skill before you install it,
and it applies to the skills you make too.