A penetration test is an authorised attempt to break into your own system before a real attacker does. Automated pentesting runs the routine part of that attack on every change you ship, so a new hole is caught in hours rather than at next year's audit.
What a penetration test is
A penetration test, or pentest, is a simulated attack carried out with the owner's permission. A tester, often called an ethical hacker, uses the same techniques a criminal would: probing forms for injection, guessing at hidden pages, trying to read other users' data, poking at login and session handling. The difference is that they stop at proof, write down what they found and hand you a report with fixes.
A good manual pentest has three parts:
- Scope and rules of engagement: which systems are in bounds, which are not, what hours testing may run and who to call if something breaks. Testing a system without this written permission is an attack, not a test.
- Discovery and exploitation: mapping the application, finding weaknesses and, crucially, trying to exploit them to show the real impact. 'This field might be injectable' is a lead; 'I read the orders table through this field' is a finding.
- Reporting: each issue with its severity, how to reproduce it and how to fix it.
That exploitation step is what separates a pentest from a vulnerability scan. A scanner says 'this looks risky'. A tester shows what an attacker could actually do with it.
Why a yearly test falls behind
The weakness of a traditional pentest is timing. It is a snapshot: it tells you how secure the application was during the week it was tested. Most teams then keep shipping, sometimes many times a day.
Every commit can open a new hole. A developer adds an endpoint and forgets an authorisation check. A dependency update pulls in a library with a known vulnerability. Someone turns on verbose error pages to debug an issue and never turns them off. None of these existed when the tester looked, so none of them are in the report.
If the next test is twelve months away, a bug introduced the day after the last one can sit in production for a year. Attackers do not wait for your audit cycle; automated tools scan the internet constantly for exactly these common mistakes.
What automated pentesting does
Automated pentesting takes the repeatable part of a tester's work and runs it as software, on a schedule or on every change. The main technique for web applications is dynamic application security testing (DAST): a tool runs against the live, running application from the outside, the way an attacker would, with no access to the source code.
A typical DAST run:
- Crawls the application to find its pages, forms, parameters and API endpoints.
- Checks passively what comes back: missing security headers, cookies without the Secure or HttpOnly flags, stack traces, version numbers leaking in responses.
- Attacks actively, if you allow it: it sends crafted payloads for SQL injection, cross-site scripting, path traversal and similar bugs, and watches how the application reacts.
- Reports each alert with a risk level, and can fail the build when something serious turns up.
Some newer tools go further and try to chain steps together or confirm that a weakness is really exploitable, which is closer to what a human tester does. The principle is the same either way: the same attacks, run automatically, every time the code changes.
DAST is one layer among several. Static analysis (SAST) reads your source code for dangerous patterns, and dependency scanning checks your libraries against lists of known vulnerabilities. They catch different things: static analysis sees code paths a crawler never reaches, while DAST sees how the deployed system actually behaves, including its configuration.
A worked example: scanning on every change
Say you run a small shop that stores customer orders. Your pipeline builds the app, deploys it to a staging environment and runs the tests. You add a security scan as the next step, using OWASP ZAP, a free, open source DAST tool.
ZAP's baseline scan crawls the target for about a minute and reports what passive checks find. It sends no attacks, so it is quick and safe to run on every push:
name: security-scan
on:
push:
branches: [main]
jobs:
zap-baseline:
runs-on: ubuntu-latest
steps:
# deploy to staging runs in an earlier job
- name: Baseline scan of staging
run: |
docker run -t \
ghcr.io/zaproxy/zaproxy:stable \
zap-baseline.py \
-t https://staging.example.comThe script exits with a non-zero code when it raises a warning or a failure, which fails the job and blocks the merge until someone looks. You can give it a rules file that marks individual alerts as ignore, warn or fail, so a known, accepted issue does not break every build.
ZAP's full scan adds the active attacks. It takes longer and changes data as it goes, so run it nightly against staging rather than on every push, and never against production without agreeing it first:
on:
schedule:
- cron: "0 2 * * *"Now picture a Tuesday. A developer adds a search box to the order history page and builds the SQL query by joining strings. The nightly full scan sends a payload with a stray quote, sees a database error come back, and raises a high-risk SQL injection alert. The fix ships on Wednesday. Without the scan, that search box might have waited months for the next manual test.
Robots and humans together
Automation does not replace a human tester. It changes what the human spends time on.
Automated tools are good at:
- Known classes of bug with recognisable signatures: injection, cross-site scripting, missing headers, weak TLS settings.
- Consistency: they never get bored, skip a page or forget last month's checks.
- Speed and frequency: every commit, every night, at no extra cost per run.
Human testers are good at:
- Business logic flaws: applying a discount code twice, skipping the payment step, changing a price in a hidden field.
- Access control: noticing that changing order 1041 to 1042 in the URL shows someone else's order. A scanner has no idea which orders are yours.
- Chaining: combining three low-risk issues into one serious attack.
- Judgement: deciding whether an alert matters in this system, and how much.
The practical split is that the robot handles the routine checks continuously, and a person investigates the tricky, context-heavy parts and follows up on anything the robot flags but cannot confirm. Because the easy bugs are already caught, the human's time goes further.
Common mistakes
- Scanning only the login page. A scanner that cannot sign in sees almost nothing. Give it a test account so it can reach the pages real users use.
- Treating a clean scan as proof of security. It means the tool found nothing it knows how to find. Logic and access control bugs pass straight through.
- Letting alerts pile up. A scan that always fails with the same twenty low-risk warnings trains the team to ignore it. Triage them once, suppress the accepted ones and keep the build strict.
- Active scanning in production. Active attacks submit forms and write data. Point them at staging, which should look as much like production as possible.
- Scanning what you do not own. Only test systems you have written permission to test, including third-party services your app calls.
Key takeaways
- A pentest is an authorised, simulated attack that proves what an attacker could really do.
- A yearly manual test is a snapshot; code that changes every day needs checks that run every day.
- Automated pentesting runs the routine attacks on every change, with passive checks per push and active scans nightly against staging.
- Scanners catch known patterns; humans still find logic flaws, access control gaps and chained attacks.
- Use both: automation for coverage and speed, people for depth.