When one server can't keep up with your traffic, you have two ways to add power: make that server bigger, or add more servers beside it. The first is vertical scaling, the second is horizontal scaling, and which one you reach for decides how much your app can grow, what it costs and what happens when a machine fails.
What scaling means
Scaling is adding capacity so a system keeps working as the load on it grows. Load can mean more users, more requests per second, more data or heavier work per request.
A server struggles when one of its resources runs out. The usual suspects are:
- CPU: requests queue up waiting for processor time, so responses slow down.
- Memory: the app, its caches and the database all compete for RAM, and once it runs out the machine starts swapping to disk or killing processes.
- Disk I/O: a busy database reads and writes faster than the disk can serve.
- Network: the machine can only push so much data in and out.
Find out which one is the bottleneck before you scale. Adding CPU to a server that is starved of memory does nothing.
Vertical scaling: a bigger machine
Vertical scaling, or 'scaling up', means giving the one server you have more resources: more CPU cores, more RAM, faster disks, a faster network card. In the cloud this is usually changing the instance type, say from a machine with 2 virtual CPUs and 4 GB of RAM to one with 8 and 32.
Why it's the easy option
Nothing about your application changes. It runs on one machine before and after, so you don't need to rethink sessions, caches, file uploads or how the database is reached. For most apps it is a setting and a restart.
That makes it the right first move for a young project. It is cheap in engineering time, and a single modern server can handle far more than most new apps will ever see.
Where it stops
- One machine can only get so big. Once you're on the largest instance your provider offers, there is nowhere left to go.
- Cost climbs faster than capacity. The biggest machines cost a lot more, and you pay for that size all day, even at 3 a.m. when traffic is low.
- Resizing often means downtime. On many cloud platforms you stop the server, change its type and start it again. That is a planned outage, however short.
- It is still one machine. If it crashes, runs out of disk or needs a security patch, the whole app goes down with it. This is called a single point of failure, and no amount of extra RAM removes it.
Horizontal scaling: more machines
Horizontal scaling, or 'scaling out', means running several copies of your app on separate servers and sharing the traffic between them. To add capacity you add a server, and when traffic drops you remove one.
The load balancer
Users still need one address to visit, so a load balancer sits in front of the servers. Every request reaches the load balancer first, and it picks a server to pass it to, often by taking turns (round robin) or by choosing the least busy one.
It also runs health checks: it asks each server at a regular interval whether it's alive, and stops sending traffic to any that don't answer. That is what turns more servers into more reliability. Here is a server failing while the shop stays open:
A load balancer routing around a failed server
Step 1 of 7: A shopper's request arrives at the load balancer, not at any one server.
Why it takes more work
Spreading requests across machines only works if any server can answer any request. That rules out keeping anything important on one server's own memory or disk:
- Sessions: if a user's login lives in server 1's memory and their next request lands on server 3, they're logged out. Move sessions to a shared store such as Redis or the database, or use signed tokens the client sends with each request.
- Uploaded files: a file saved to server 2's disk doesn't exist on the others. Put files in shared object storage instead.
- Scheduled jobs: a nightly job running on every copy runs three times. Run it in one place, or use a lock.
An app that keeps no per-user data on the server between requests is called stateless, and making the app servers stateless is most of the work of scaling out.
The database is the hard part
Adding app servers is easy once they're stateless. The database is harder, because all those servers now share it, and data can't simply be copied around without the copies disagreeing.
The usual steps, roughly in order:
- Scale the database vertically, often for a long time.
- Add read replicas: copies that serve reads while one main database takes the writes. Replicas can lag slightly behind.
- Add a cache in front of hot data.
- Only when writes outgrow one machine, shard: split the data across several databases, for example by customer. Sharding is powerful and hard to undo, so leave it until you need it.
Once data lives on more than one machine, you're in distributed systems territory, with trade-offs like the ones the CAP theorem describes.
A worked example: an online shop grows
Picture a small online shop on a single cloud server running both the web app and the database.
In the first month, traffic is light and one small server is plenty.
Six months in, pages slow down at lunchtime. Monitoring shows the database is using most of the RAM and the CPU is pegged. The team resizes the server to a larger instance during a quiet hour. It costs a few minutes of downtime and no code changes, and the lunchtime slowdown goes away. This is what vertical scaling is good at.
A year in, a sale is coming, and the team expects several times the normal traffic for a few hours. Resizing again would mean paying for a huge machine all year for one afternoon, and a crash mid-sale would take the whole shop down. So they:
- Move the database to its own server, and scale that one vertically.
- Move sessions into Redis and product images into object storage, making the app stateless.
- Put a load balancer in front of three app servers.
- Turn on autoscaling, which adds app servers when CPU stays high and removes them when it drops.
On sale day the app tier grows to handle the rush and shrinks again afterwards. When one app server fails, the load balancer routes around it. The database, scaled up rather than out, takes the extra load on its bigger machine.
That's the pattern most teams end up with: scale up first, then scale out the parts that need it.
When to use which
| Vertical | Horizontal | |
|---|---|---|
| Effort | A setting change | App changes, load balancer |
| Limit | Largest machine available | Mostly cost and design |
| Failure | Whole app goes down | Others keep serving |
| Downtime to grow | Often a restart | None, add a server |
| Best for | Databases, early apps | Stateless web and API tiers |
Choose vertical scaling while the app is young and the load is steady, and for parts that are hard to split, which often means the database. Choose horizontal scaling when you need to survive a server failing, when traffic rises and falls sharply, or when you're near the top of what one machine can do.
Common mistakes
- Scaling out a stateful app: adding servers before moving sessions and files out leads to users being logged out at random and files that seem to vanish.
- Scaling before measuring: without knowing whether the bottleneck is CPU, memory, disk or the database, you may pay for capacity in the wrong place.
- Forgetting the database: ten app servers sharing one overloaded database are no faster than two.
- Assuming two servers means no downtime: if both depend on one database, one load balancer or one region, that is still a single point of failure.
- Splitting too early: a distributed setup has more moving parts to deploy and debug. If one bigger server would do, that's usually the cheaper choice.
Key takeaways
- Vertical scaling makes one server bigger; horizontal scaling adds more servers and shares the load between them.
- Vertical is simple and needs no code changes, but it has a hard ceiling and leaves a single point of failure.
- Horizontal scales further and survives a failed server, but needs a load balancer and a stateless app.
- The database is usually the last part to scale out, and the hardest.
- Most systems scale up first, then scale out when traffic or reliability demands it.