A red team plays the attacker and a blue team plays the defender, inside the same organisation and with permission. The exercise tests whether your security holds up when someone tries to break in, which a green dashboard can't tell you.
Why attack yourself
Most security controls are set up once and then trusted. A firewall rule written two years ago, an alert nobody has seen fire, a patching policy that exists on paper: each one looks fine until someone tests it. Often the first test comes from an attacker, which is the worst time to find a gap.
A red team and blue team exercise moves that first test to a time you choose. A group of authorised people behave like an adversary, and the people who defend the systems every day try to catch them. Afterwards both sides compare notes. The result is a list of weak spots found by people on your side, before anyone else finds them.
What the red team does
The red team's job is to think and act like a likely attacker. They pick a goal, such as reading the customer database or taking over an admin account, and try to reach it by whatever route a real attacker would use.
They use the same ways in as attackers do:
- Phishing: emails that trick someone into clicking a link, opening an attachment or typing their password into a fake login page.
- Stolen or reused passwords: credentials leaked from another site, guessed, or captured by that phishing page.
- Unpatched software: a known vulnerability in a server, library or appliance that nobody has updated yet.
- Misconfiguration: a storage bucket left public, a test server still on its default password, an admin panel open to the internet.
Once inside, the red team tries to move further: from one laptop to a file server, from a normal account to an admin one. They record every step, because the write-up of how they got in matters more than the fact that they did. Each route they try is part of your attack surface, the set of places an attacker can get in.
What the blue team does
The blue team defends. In many organisations they are the security operations team, the people who run monitoring and respond to incidents every day. During an exercise they often aren't told when the attack starts, so they have to notice it the way they would notice a real one.
Their work during an exercise follows the same order as any incident response:
- Detect: watch alerts, read logs and hunt for behaviour that doesn't fit, such as a login from a new country at 3am or a service account suddenly reading thousands of files.
- Contain: stop it spreading. Disable the stolen account, or cut the infected laptop off the network.
- Remove: get the attacker out completely, including any backdoor or extra account they created to come back later.
- Fix the cause: patch the bug, turn on multi-factor authentication, or change the rule that let the phishing email through.
The last step matters most. Kicking out the red team only proves the blue team can react. Closing the hole they used means the next attacker can't use it either.
A worked example
Here is how one exercise might go at a small software company.
The red team sends a convincing 'shared document' email to the finance team. One person enters their password on the fake login page. The red team signs in with it, but the company uses multi-factor authentication on email, so the login stalls there.
They try another route. The company's VPN appliance hasn't been updated in months and has a publicly known vulnerability. The red team uses it to get a foothold on the internal network without any password at all. From there they find a file share with a spreadsheet of service account passwords and use one to read the customer database. Goal reached.
Now the blue team's side of the story. Their tools flagged the phishing login attempt, and someone reset the password within an hour. But nothing alerted on the VPN exploit, and the database reads by a service account at an odd hour went unnoticed, because nobody had set up an alert for them.
The debrief turns that into concrete work:
- Patch the VPN appliance, and add it to the regular patching list so it isn't forgotten again.
- Move the passwords from the spreadsheet into a secrets manager and rotate them.
- Add an alert for service accounts reading unusual amounts of data.
- Keep the phishing detection as it is: it worked.
No real customer data left the building, and the company now knows about three gaps it didn't know it had.
Rules of engagement
An exercise is only safe because it has rules agreed before it starts. These usually cover:
- Scope: which systems, people and techniques are in bounds. Systems that customers rely on are often handled with extra care, or left out entirely.
- Authorisation: written sign-off from someone senior enough to give it. Without it, the red team's work looks exactly like a real attack, legally as well as technically.
- A referee: many exercises have a small neutral group, sometimes called the white team, who know both sides' plans, keep score and can stop the exercise if something breaks for real.
- Real incidents: if the blue team spots an outside attacker during the exercise, that comes first.
Red teaming and penetration testing
The two get confused because both involve people breaking into your systems on purpose. The difference is the question each answers.
| Penetration test | Red team exercise | |
|---|---|---|
| Question | What can be broken? | Would we notice? |
| Scope | One app or network | The whole organisation |
| Defenders | Often know about it | Often aren't told |
A penetration test tries to find as many vulnerabilities as possible in a defined target, and its report is a list of flaws to fix. A red team exercise picks a goal and tests the people, processes and tools that are meant to stop someone reaching it. Pen tests are often run as black box or white box tests, depending on how much the tester is told, and parts of them can be automated. A red team exercise is harder to automate, because it tests how people react.
Purple teaming
In a classic exercise the two teams work apart and only compare notes at the end. Purple teaming puts them in the same room. The red team runs one technique, the blue team checks whether they saw it, and if they didn't, both sides work out together what log or alert would have caught it. Then they try again.
This gives up the surprise of a real attack, but each gap is found and fixed while both sides are still looking at it. Many teams use a shared catalogue of attacker techniques, such as MITRE ATT&CK, so both sides can name exactly which behaviour was tested and which was detected.
Common mistakes
- Treating it as a contest: if the red team 'wins' and the report gets filed away, nothing has improved. The point is the fixes.
- Blaming the person who clicked: someone will fall for a good phishing email sooner or later. The useful question is why one click was enough to cause damage.
- Testing only the technology: attacks also go through people and processes, such as the help desk that resets passwords over the phone, or the alert that goes to an inbox nobody reads.
- Running it once: systems change and people leave, so an exercise from two years ago says little about today.
- Starting before the basics are in place: if there's no logging or patching yet, a red team will only confirm what you already know. Fix the obvious gaps first.
Key takeaways
- The red team acts like an attacker to find weak spots; the blue team defends, detects and responds.
- A good response means spotting the attack, containing it, removing the attacker and fixing how they got in.
- Exercises need clear scope, written authorisation and someone to referee.
- A pen test asks what can be broken; a red team exercise asks whether you would notice.
- The value is in the fixes afterwards, not in who won.