Anthropic's webapp testing skill teaches your AI agent to check a web app the way a person would: open it in a browser, click things, and look at what happened. The agent writes a short Python script with Playwright, starts your dev server if it isn't running, runs the script in a headless Chromium, and reads back screenshots, the page's HTML and the console log.
What it does
The skill gives the agent a decision tree:
- A static HTML file? Read the file to find selectors, then script it.
- A dynamic app that isn't running? Start it with the bundled
with_server.py, which launches one or more servers, waits for their ports to open, runs the test, and stops them afterwards. - Already running? Look before acting: load the page, wait for the network to go quiet, take a screenshot or dump the DOM, find the selectors, then write the actions.
It flags the mistake that causes most flaky scripts:
❌ Don't inspect the DOM before waiting for
networkidleon dynamic apps
From SKILL.md by Anthropic, Apache 2.0.
Three example scripts show how to list a page's buttons and links, drive a local HTML file, and capture console messages.
An example
You say: "Check the signup form shows an error for a bad email." The agent runs
with_server.py --server "npm run dev" --port 5173 with a script that opens
the form, types not-an-email, submits, takes a screenshot and prints any
console errors. It reads the result and reports whether the error message
appeared, with the screenshot as evidence.
When to use it, and when not to
Use it when you want your agent to see the result of a frontend change instead of assuming it worked: forms, navigation, a bug you can only see in the browser.
It's not a test framework. For tests that run in CI, ask the agent to write them with your project's own setup (Playwright Test, Vitest, Cypress). And if you don't have Python, it's the wrong tool.
What we checked
We read every file in the folder at the commit linked above: SKILL.md,
scripts/with_server.py, the three files in examples/ and the Apache 2.0
LICENSE.txt. There are no allowed-tools. What it can do:
- Runs scripts.
with_server.pyruns the server commands you (or the agent) give it through the shell, so it can run anything a terminal can. That's the point of it, but it means you should read the command before approving it. - Stays local. The helper and examples we reviewed connect only to
localhostports, and nothing in them downloads or sends anything, apart from Playwright's own browser download when you install it. Scripts your agent writes with the skill can point anywhere, so check their URLs too.
One line to be aware of: the skill tells the agent to run the helper with
--help and not read its source unless it needs to, to save context. That's
a reasonable instruction, not an attempt to hide anything; the script is about
100 lines and easy to read yourself. Our full checklist is in the
guide to checking a skill before you install it.
If you'd rather give your agent a browser it can drive step by step, see the Playwright MCP server.