Every test answers one question: does this software do what it should? White-box and black-box testing answer it from opposite sides, one by reading the code and one by only using the product, and a healthy project needs both.
Two ways to look at the same app
The names describe what the tester can see. Picture the application as a box.
- In white-box testing (also called clear-box or glass-box testing), the box is transparent. The tester can read the source code and writes tests based on how it works: which functions exist, which branches an
ifcan take, how many times a loop runs. - In black-box testing, the box is closed. The tester can only use the application the way a user or another system would: click, tap, fill in forms, send requests. Tests are based on what the application should do, and a test passes or fails purely on whether the output matches what was expected for a given input.
Neither is a kind of tool or framework. They are two approaches to deciding what to test, and the same test runner can run both.
White-box testing: testing the inside
White-box tests start from the code. You read a function, list the paths through it, and write a test for each one. The aim is that no line, branch or loop goes untested.
A few ideas come up again and again:
- Statement coverage: every line has run at least once during the tests.
- Branch coverage: every decision has gone both ways. An
ifwith noelsestill has two branches, the one where the condition is true and the one where it is skipped. - Loop testing: a loop that runs zero times, once, and many times often behaves differently in each case.
- Edge cases the code reveals: reading the code shows you that a list is indexed with
items[0], so you know to test an empty list.
This is why unit tests and integration tests are usually written white-box. A developer writing a unit test for a function they just wrote knows its internals, and naturally aims tests at the tricky parts. Coverage tools measure how much of the code those tests actually exercised and highlight what they missed.
The strength of white-box testing is thoroughness at a low level. It finds dead code, paths nobody thought to try, and bugs deep inside a function that would be hard to trigger from the outside. Its weakness is that it tests the code as written. If the developer misunderstood the requirement, a test written by reading their code will often confirm the misunderstanding rather than catch it.
Black-box testing: testing the outside
Black-box tests start from the requirements, the specification or the user's expectations. The tester does not know, or deliberately ignores, how the feature is built. They pick inputs, predict the correct outputs from the spec, and check what comes back.
Because there is no code to guide them, black-box testers rely on techniques for choosing inputs well:
- Equivalence partitioning: split the possible inputs into groups the spec treats the same way, and test one value from each group rather than hundreds.
- Boundary value analysis: bugs cluster at the edges of those groups, so test just below, on, and just above every limit the spec mentions.
- Invalid input: negative quantities, empty fields, huge values, the wrong type. The spec should say what happens, even if it is only an error message.
End-to-end tests that drive a browser, API tests that send HTTP requests to a running service, acceptance tests written from user stories and manual exploratory testing are all black-box in style.
Its strength is that it checks what users actually experience, and it is independent of the implementation: rewrite the whole feature and the black-box tests still apply unchanged. Its weakness is blind spots. Without seeing the code, you cannot know whether an unusual path inside it was ever run.
A worked example
Here is a small requirement: 'Delivery is free for members, and for orders of £50 or more. Otherwise it costs £4.99.'
A developer writes this:
def delivery_fee(total, is_member):
if is_member:
return 0.0
if total > 50:
return 0.0
return 4.99The white-box tests
Reading the code, a tester sees three paths: the member branch, the over-50 branch and the fallthrough. One test each gives full branch coverage:
def test_member_pays_nothing():
assert delivery_fee(10, True) == 0.0
def test_big_order_is_free():
assert delivery_fee(80, False) == 0.0
def test_small_order_pays():
assert delivery_fee(20, False) == 4.99All three pass, and a coverage report shows every line and branch was run. It looks finished.
The black-box tests
A tester who only has the requirement partitions the inputs (member, non-member under £50, non-member at £50 or over) and then tests the boundary the spec names:
def test_exactly_fifty_is_free():
assert delivery_fee(50, False) == 0.0
def test_just_under_fifty_pays():
assert delivery_fee(49.99, False) == 4.99The first test fails. The spec says '£50 or more', which is >=, but the code says >. The white-box tests missed it because they were written from the code, and the code was wrong. Full coverage only proved every line ran, not that every line was right.
The reverse happens too. If the function had a hidden branch the spec never mentions, say a special case for one postcode added during a bug fix, black-box tests would almost never reach it. A white-box tester reading the code would spot it straight away and test it.
How they compare
| White-box | Black-box | |
|---|---|---|
| Tester sees | The source code | Only inputs and outputs |
| Tests come from | How the code works | What the app should do |
| Typical tests | Unit, integration | End-to-end, API, acceptance |
| Good at finding | Untested paths, dead code | Wrong behaviour, missed requirements |
| Survives a rewrite | Often not | Yes |
Grey-box testing and security
Real testing is rarely purely one or the other. Grey-box testing sits in between: the tester has partial knowledge, such as the API documentation, the database schema or an architecture diagram, but still tests from the outside. Knowing that a search box feeds a SQL query, for instance, tells a grey-box tester exactly which inputs are worth trying.
The same three terms are used in security testing. A black-box penetration test gives the testers no inside information, the same starting point as an outside attacker, which shows what a stranger could find. A white-box test gives them source code, configuration and credentials, which lets them find more weaknesses in the same time. A grey-box test, often with an ordinary user account, models an attacker who is already partly inside. None is 'the best': they answer different questions about how exposed a system is.
Common mistakes
- Treating coverage as correctness. 100% coverage means every line ran, not that every result was checked against the requirement, as the delivery example shows.
- Only testing the happy path. Both approaches are weakest when nobody tests boundaries, empty inputs and invalid data.
- Writing white-box tests that mirror the code. If a test repeats the implementation's logic to compute the expected value, it will agree with any bug in it. Derive expected values from the requirement where you can.
- Black-box tests that peek inside. An end-to-end test that asserts on internal state, private fields or exact database rows breaks every time the implementation changes, losing the main benefit of testing from the outside.
- Picking one approach for everything. Fast white-box unit tests catch low-level mistakes early; slower black-box tests confirm the feature works for a real user. You want both, in proportion.
Key takeaways
- White-box testing looks inside the code and writes tests from how it works: every function, branch and loop.
- Black-box testing only uses the app, and writes tests from what it should do, judged on inputs and outputs.
- White-box tests find untested paths; black-box tests find behaviour that does not match the requirement.
- Coverage shows what ran, not what is correct.
- Grey-box testing mixes the two, and the same terms describe how much a security tester is told.