An app that feels instant with one user can crawl, or fall over, when a thousand arrive at once. Performance testing puts your app under realistic and unrealistic pressure on purpose, so you find out how it behaves before your users do.
What performance testing measures
A functional test asks 'does it work?'. A performance test asks 'how well does it work, and for how many people at once?'. You simulate many users sending requests, and you watch a handful of numbers.
- Response time: how long a request takes from the user's side, from sending it to getting the full answer.
- Throughput: how many requests the system finishes per second.
- Error rate: the share of requests that fail or time out.
- Resource use: CPU, memory, disk and network on each server, plus things like open database connections.
Why averages hide the problem
Response time is usually reported as percentiles, not an average. The p95 is the time that 95% of requests finish within; the p99 is the time 99% finish within. If most requests take 100 ms but one in twenty takes 3 seconds, the average still looks healthy while one user in twenty is staring at a spinner. A target such as 'p95 under 500 ms' describes what users feel far better than 'average under 200 ms'.
The numbers also only mean something together. Response time that stays flat while throughput climbs is good news. Response time that climbs while throughput stops growing means something is saturated: the system has hit a limit and requests are queueing for it.
Load tests and stress tests
'Performance testing' is the umbrella. Underneath it are a few test types, which differ mainly in how much traffic you send and what question you are asking.
Load testing
A load test sends the traffic you expect: a normal busy day, or the peak you plan for. You ramp up to that level, hold it for a few minutes, and check the results against your targets. It answers one question: at 100 users, or whatever your real peak is, do we stay fast and error-free?
Stress testing
A stress test keeps going past the expected load, to find where the system breaks and how. Maybe 10,000 users is unrealistic today, but you learn two useful things: roughly where the limit is, and what happens at it.
A system that degrades gracefully slows down a little, rejects some requests quickly with a clear error, and recovers once traffic drops. A system that fails badly runs out of memory, crashes, takes its neighbours down with it and needs a restart. Stress testing tells you which one you have.
Spike and soak tests
Two other test types change the shape of the traffic:
- A spike test jumps from quiet to very busy in seconds, like a sale going live or a link going viral. It tests whether autoscaling and caches cope with a sudden rush rather than a gentle ramp.
- A soak test holds ordinary load for hours. Problems that only appear over time show up here: memory that slowly leaks, logs filling a disk, connections that are never returned.
A worked example: testing a menu endpoint
Imagine a food ordering app whose GET /menu endpoint is hit on every visit. The team expects about 100 people browsing at once at lunchtime, and they agree a target: p95 under 500 ms, with fewer than 1% of requests failing.
Here is a load test for that target in k6, a popular open source tool where tests are written in JavaScript:
import http from 'k6/http';
import { check, sleep } from 'k6';
export const options = {
stages: [
{ duration: '1m', target: 100 },
{ duration: '3m', target: 100 },
{ duration: '1m', target: 0 },
],
thresholds: {
http_req_duration: ['p(95)<500'],
http_req_failed: ['rate<0.01'],
},
};
export default function () {
const res = http.get('https://staging.example.com/menu');
check(res, {
'status is 200': (r) => r.status === 200,
});
sleep(1);
}Each virtual user runs the default function in a loop: request the menu, check it came back with a 200, pause for a second like a real person reading. The stages ramp up to 100 users over a minute, hold for three, then ramp down. The thresholds turn the target into a pass or fail, so k6 exits with an error code when p95 goes over 500 ms or more than 1% of requests fail. That makes the same script usable as a check in a pipeline.
To turn it into a stress test, you keep the script and change the stages: ramp to 500, then 1,000, then 2,000 users, and watch where response time bends upwards and errors begin.
Point tests like this at a staging environment that resembles production, never at a live system without warning. A stress test is, from the server's side, very similar to an attack.
Finding and fixing bottlenecks
A bottleneck is the one part of the system that runs out of capacity first, so everything else waits for it. Adding more of anything else does not help: if the database is the limit, a second app server only sends it twice as many queries.
Say the menu test passes at 100 users, but at 500 the p95 climbs to over 2 seconds while the app server's CPU sits mostly idle. Something other than the app's own code is slow. Here is what is happening inside, from a calm start to the fix:
A slow query becomes a bottleneck under load
Step 1 of 7: At 100 users, each request reaches the app server and is answered quickly.
The test did not point at the index directly. It showed the symptom (slow responses, idle CPU), and the server's metrics narrowed it down to the database. That is the usual pattern: the test finds that something is slow, and monitoring tells you where.
Common bottlenecks
- Database queries: a missing index, a query inside a loop (the 'N+1' problem), or fetching far more rows than the page needs.
- Connection pools: too few database or HTTP connections, so requests queue for one.
- External calls: a slow payment or email service that every request waits on.
- CPU-heavy work in the request path, such as resizing images or generating reports, which is often better moved to a background job.
- Memory: leaks that grow under load until the process slows down or is killed.
Fix one thing at a time and rerun the same test. If you change three things at once, you will not know which one helped, or which one made things worse.
When to run performance tests
Bottlenecks are cheapest to fix early. A slow query found during development is a one-line index. The same query found on launch day is an outage.
In practice, teams tend to combine a few habits:
- A short, small load test in the CI/CD pipeline, with thresholds, so an obvious slowdown fails the build.
- A full load test before a release that changes something heavily used.
- Stress and spike tests before a known big event, like a launch or a sale.
- Soak tests now and then, or after changes to memory or connection handling.
Performance testing is not worth much effort on an internal tool used by five people, or on a prototype that may be thrown away. It earns its keep where slowness costs money or trust: checkout, sign-in, search, anything on the busiest path.
Common mistakes
- Testing on a laptop: a local database with ten rows behaves nothing like production with ten million. Test against realistic data and hardware, or the numbers mean little.
- Unrealistic users: virtual users that never pause and all hit the same URL test a cache, not your app. Add think time and a realistic mix of requests.
- Only reading the average: watch p95 and p99, and the error rate next to them.
- Overloading the load generator: if the machine running the test is out of CPU, the slowdown you measure is its own.
- Testing once: performance changes with every release. A test you run only before launch tells you nothing about next month's code.
Key takeaways
- Performance testing checks how fast and stable an app stays under many users, not just whether it works.
- Load tests check expected traffic against a target; stress tests push past it to find the breaking point and see whether the app fails gracefully.
- Measure response time as percentiles, alongside throughput, error rate, CPU and memory.
- A bottleneck is the first part to run out of capacity; find it with metrics, fix it, and rerun the same test.
- Test early and often, on realistic data, so slowdowns are cheap to fix.