A race condition is a bug where the result depends on timing: two things run at once, both touch the same data, and whoever gets there first decides what happens. The code is often correct line by line, which is why these bugs get through code review and tests, then turn up in production under load.
What makes it a race
Three ingredients have to be present together:
- Concurrency: two or more pieces of work overlap in time. They might be threads in one program, two web requests handled by different servers, or two background workers.
- Shared state: they read and write the same thing, such as a variable, a database row, a file or a cache key.
- At least one write: if everyone only reads, the order doesn't matter.
Take away any one of the three and the race goes away.
The order the steps run in is called an interleaving. With two requests of three steps each, there are many possible interleavings, and most of them are fine. The bug lives in the few that aren't, which is why the same code can work a thousand times and fail on the next run.
Check, then act
The most common shape is 'check, then act': read a value, decide something from it, then write. Selling the last item in a shop is the classic case.
- Request A reads the stock: 1 left.
- Request B reads the stock: still 1, because A hasn't written yet.
- A sells the item and sets the stock to 0.
- B, trusting the 1 it read earlier, sells the same item and also sets the stock to 0.
The stock ends at 0, which looks right, but two customers have paid for one item. Nothing crashed and no error was logged. Here is that interleaving step by step:
Two requests selling the last item
Step 1 of 5: Request A reads the stock and sees one item left.
A close cousin is 'read, modify, write'. A line like count += 1 looks like one step, but it is three: read count, add 1, write it back. If two threads both read 5, both write 6, and one increment is lost.
A worked example in Python
This sketch runs two buyers on two threads. The sleep stands in for anything slow between the check and the write, such as a payment call or a network hop, and makes the bad interleaving happen nearly every time:
import threading
import time
stock = 1
orders = []
def buy(name):
global stock
if stock > 0: # check
time.sleep(0.01) # slow work
stock -= 1 # act
orders.append(name)
a = threading.Thread(target=buy, args=("A",))
b = threading.Thread(target=buy, args=("B",))
a.start(); b.start()
a.join(); b.join()
print(stock, orders)Both threads pass the if before either of them reaches stock -= 1, so it prints a stock of -1 and two orders. Remove the sleep and it will usually print the right answer, which is the problem: the bug is still there, it just shows up less often.
The fix is to make the check and the write one step that only one thread can be inside at a time. A lock does that:
lock = threading.Lock()
def buy(name):
global stock
with lock:
if stock > 0:
time.sleep(0.01)
stock -= 1
orders.append(name)Now the second thread waits at with lock: until the first has finished, then sees a stock of 0 and does nothing. It prints 0 and a single order. The section of code guarded by the lock is called a critical section.
How to fix one
Each fix stops two workers from acting on the same data at the same moment. Which one fits depends on where the shared state lives.
Locks
A lock (or mutex) lets one thread into a critical section at a time. It is the right tool for shared memory inside one process. Keep the locked section small: everything inside it runs one at a time, so slow work there turns a fast program into a queue. A lock in one process does nothing for a second server, so it can't protect a database row that several machines update.
Atomic operations
Often the simplest fix is to let the database do the check and the write in a single statement:
UPDATE products
SET stock = stock - 1
WHERE id = 42 AND stock > 0;The database locks the row while it updates it, so two of these can't both succeed on the last item. Check how many rows were updated: 1 means the sale went through, 0 means it sold out. Languages offer the same idea in memory as atomic counters and compare-and-swap operations.
Transactions and row locks
When the logic is too complex for one statement, read the row with SELECT ... FOR UPDATE inside a transaction. That locks the row until the transaction commits, so a second request trying to read it the same way waits its turn. A stricter isolation level, such as SERIALIZABLE, makes the database detect the conflict and fail one of the transactions, which your code then retries.
Optimistic concurrency
Instead of locking, give each row a version number. Read the row and its version, then write with WHERE id = 42 AND version = 7. If someone else changed the row first, the version no longer matches, zero rows are updated, and you read again and retry. This suits data that is read often and rarely updated at the same moment.
Queues
A queue removes the 'at once' part. Put every order for an item on one queue and let a single worker process them in turn: there is no overlap, so there is no race. The cost is throughput and a little latency. It is a common pattern for background jobs, where several workers might otherwise pick up the same task.
Races without threads
JavaScript runs your code on one thread, but it still has race conditions. Any await is a gap where other code can run. A search box that sends a request on every keystroke can show stale results if an older, slower response arrives after a newer one. The fix is to ignore or cancel responses that are no longer current.
The same is true across machines. Two servers behind a load balancer share a database, not memory, so a lock in one of them protects nothing. Races between services are one of the everyday difficulties of distributed systems.
Races also matter for security. A program that checks a file's permissions and then opens it leaves a gap in which an attacker can swap the file. This is called time-of-check to time-of-use, or TOCTOU.
Common mistakes
- Trusting a passing test: a race that fails one run in ten thousand will pass almost every test run. Tools help: Go has a built-in race detector (
go test -race), and ThreadSanitizer does the same for C and C++. - Assuming a transaction is enough: at the default isolation level in many databases, two transactions can both read the stock as 1 and both write 0. Use an atomic update, a row lock or a stricter isolation level.
- Locking too much: one big lock around everything makes the program correct and slow. Two locks taken in different orders by two threads can also leave each waiting for the other forever: a deadlock.
- Mixing up a race condition and a data race: a data race is two threads touching the same memory at once without synchronisation, with at least one writing. A race condition is broader: the shop example above happens across separate requests and a database, with no shared memory at all.
Key takeaways
- A race condition is a bug where timing decides the outcome: concurrent work, shared data and at least one write.
- 'Check, then act' and 'read, modify, write' are the two shapes to look out for.
- Race conditions show up only sometimes, so passing tests prove little.
- Fix them by making the check and the write one step: a lock, an atomic update, a row lock, a version check or a queue.
- A lock only protects what is in its own process: across servers, let the database or a queue do it.