Backend work is mostly about what happens between a request arriving and a response leaving: who is asking, how fast you can answer, and what breaks when many requests arrive at once. These are the ideas behind each question in the quiz, so if one caught you out, its section is below.
HTTP: the language of most APIs
Most web APIs speak HTTP, the same protocol your browser uses to load pages. A client sends a request made of a method (GET, POST, PUT, DELETE), a path, some headers and sometimes a body. The server sends back a status code, headers and usually a body, often JSON.
HTTP won for APIs because it is everywhere: every language has a client for it, it passes through firewalls and proxies, and caches and load balancers already understand it. Other protocols have their own jobs. FTP moves files, SMTP delivers email and SSH gives you a secure shell on a remote machine. None of them is a general way for two programs to ask each other for data.
A typical API call looks like this:
curl -i https://api.example.com/users/42The -i flag prints the status line and headers as well as the body, which is the quickest way to see what a server really answered.
Status codes: what the number tells you
Every HTTP response starts with a three-digit status code, and the first digit gives you the family:
- 2xx means it worked.
200 OKis the everyday success. - 3xx means go somewhere else, such as a redirect.
- 4xx means the client got something wrong.
404 Not Foundsays there is nothing at that path, and403 Forbiddensays it exists but you aren't allowed to have it. - 5xx means the server failed.
500 Internal Server Erroris the catch-all for an unhandled crash.
The difference between 4xx and 5xx matters in practice. A 4xx tells the caller to change their request; retrying the same one won't help. A 5xx says the fault is on the server's side, so a retry later might succeed. A common mistake is returning 200 with an error message in the body, which hides failures from every tool that reads status codes, from monitoring to the browser.
Authentication and authorisation
Authentication answers 'who are you?'. It is the step where a user or service proves its identity: a password, a one-time code, a signed token, or a sign-in through Google or another provider. Once it succeeds, the server knows which account is making the request.
Authorisation answers a different question: 'now that I know who you are, what may you do?'. That is where permissions, roles and ownership checks live. The two are easy to mix up because they happen one after the other, and the status codes follow the split: 401 Unauthorized really means 'not authenticated', while 403 Forbidden means 'authenticated, but not allowed'.
Encrypting data and logging activity are separate concerns again. HTTPS encrypts the connection so a password can't be read on the way, and logs record what happened, but neither proves who someone is.
Middleware: code in the middle of a request
Middleware is code that runs between the request arriving and your handler producing the response. Each piece gets the request, does one job, and either passes it on to the next piece or answers early. Logging, authentication checks, parsing JSON bodies, compression and rate limiting are all typically middleware.
In Express, a Node.js framework, a middleware function takes the request, the response and a next function:
function requireUser(req, res, next) {
if (!req.headers.authorization) {
return res.status(401).json({ error: "Sign in" });
}
next(); // hand over to the next step
}
app.get("/orders", requireUser, listOrders);The request reaches listOrders only if requireUser calls next(). Order matters: middleware runs in the order you register it, so a logger registered after the auth check never sees rejected requests.
Caching: faster reads
A cache keeps a copy of data somewhere quicker to reach than its source: in memory, in a store like Redis, or at the edge of a content delivery network. When a request needs data, the server checks the cache first. A hit returns the copy straight away; a miss falls through to the database, and the answer is usually stored for next time.
That is why caching improves read performance. It does nothing to make writes faster, and it often makes them more work, because every write now has to update or remove the cached copy.
The hard part is freshness. A cached value can go stale when the source changes, so every cache needs a rule for when copies expire, usually a time to live (TTL) or an explicit delete when the data is written. Cache data that is read far more often than it changes.
Load balancing: sharing the traffic
A load balancer sits in front of several copies of your server and distributes incoming traffic between them. Clients talk to one address; the load balancer picks a server for each request, using a rule such as round robin (take turns) or least connections (pick the least busy one).
It also runs health checks. If a server stops answering, the load balancer stops sending it traffic, so one crashed machine doesn't take the whole service down. Some load balancers also handle HTTPS, but that is an extra on top of spreading the load.
Scaling horizontally and vertically
When one server can't keep up, there are two ways to grow:
| Approach | What changes | The catch |
|---|---|---|
| Vertical | A bigger server: more CPU, more memory | There is a ceiling, and one machine is one point of failure |
| Horizontal | More servers of the same size | The app must work across many copies |
Scaling horizontally means adding more servers and splitting the work between them, which is exactly what a load balancer makes possible. It has no hard ceiling and survives a single machine failing. The cost is that your app has to be ready for it: a user's next request may land on a different server, so session data can't live in one server's memory. It belongs in a shared store, such as a database or cache, instead.
Message queues: decoupling services
A message queue lets one service hand work to another without waiting for it. The producer puts a message on the queue ('send the welcome email for user 42') and carries on. A consumer picks messages up when it is ready and processes them.
The main purpose is to decouple services. The producer doesn't need to know who handles the message, or whether that service is up right now. If the email service is down for a minute, messages wait in the queue instead of failing the sign-up. If sign-ups spike, the queue absorbs the burst and consumers work through it at their own pace, and you can add more consumers to go faster.
Queues hold messages for a while, but they aren't a long-term data store. The trade-off is that work becomes asynchronous: the user gets 'you're signed up' before the email is sent, and you have to handle a message that fails or arrives twice.
Race conditions: when timing changes the result
A race condition happens when two processes or threads work on the same data at the same time and the result depends on which one gets there first. The classic case is a lost update. Two requests add one to a counter that stands at 10, and each reads it before the other has written:
Two requests race to update a counter
Step 1 of 6: The counter stands at 10.
Nothing crashed and each request did the right thing on its own. The bug is in the timing, which is why race conditions often pass every test and only appear under real traffic.
The fix is to make the read and the write one step that nothing can interleave with. In SQL, let the database do the arithmetic:
UPDATE counters
SET value = value + 1
WHERE id = 1;Locks, which let one process in at a time, and transactions with the right isolation level do the same job.
Deadlocks: waiting forever
Locks fix race conditions, but they bring their own failure. A deadlock is when two transactions each hold something the other needs, and both wait for the other to let go. Neither can move, so without outside help they would wait forever.
Here is how it happens with two bank transfers:
- Transaction 1 locks account A, then asks for account B.
- At the same moment, transaction 2 locks account B, then asks for account A.
- Each is now waiting for a lock the other holds.
A deadlock isn't a crash or a slow query, though from outside it can look like one because requests hang. Most databases detect the cycle, pick one transaction as the victim and cancel it with an error, so your code should be ready to retry. The best prevention is to always take locks in the same order: if every transfer locks the lower account number first, the circle above can't form.
Key takeaways
- Most APIs speak HTTP, and the status code family tells you who is at fault: 4xx the client, 5xx the server.
- Authentication proves who you are; authorisation decides what you may do.
- Middleware runs between request and response, in the order you register it.
- Caching speeds up reads, load balancers spread traffic, and horizontal scaling adds servers for them to spread it across.
- Queues decouple services, locks prevent race conditions, and taking locks in a fixed order prevents deadlocks.