A system architect decides how the parts of a software system fit together before the team builds them: which services exist, how they talk, where the data lives and what happens when something breaks. They write less code than most engineers, but their decisions shape every line the rest of the team writes.
What a system architect does
Most developers work inside one part of a system, such as a feature or a service. An architect works across the gaps between those parts. Their job is to answer the questions no single developer owns:
- Structure: what are the main parts, and what is each one responsible for?
- Communication: do the parts call each other directly, share a database or send messages?
- Data: which part owns which data, and who is allowed to change it?
- Quality attributes: how fast, available, secure and cheap does it need to be? These are often called non-functional requirements, and they drive most architecture decisions.
- Change: which parts are likely to change, and how do we stop those changes rippling through everything else?
Much of the work is talking to people. Product managers know what is coming next; developers and operations staff know what hurts today. The output is a set of written decisions, plus the diagrams that explain them.
The title varies. Some companies have 'solution architects' for one product and 'enterprise architects' across the whole company. Others have no architects at all, and senior engineers and tech leads do the same thinking. What matters is that someone does it.
Why it's all boxes and arrows
Architecture diagrams look simple on purpose. A box is something that runs or stores data: a service, a database, a queue. An arrow is a dependency: this part calls, reads from or sends messages to that one. A label on the arrow says how: 'HTTPS/JSON', 'SQL', 'publishes OrderPlaced'.
A diagram gives the team a shared picture to argue about before anyone writes code, when changing your mind costs a whiteboard eraser rather than a rewrite.
A useful habit is to draw at different zoom levels, which is the idea behind the C4 model:
- Context: your system as one box, with the people and outside systems around it.
- Containers: the apps, services and databases inside it.
- Components: the main parts inside one of those containers.
- Code: classes and functions, usually left to the code itself.
Most teams only ever need the first two. A diagram that mixes levels, with a database next to a single class next to 'the cloud', is usually a sign the thinking is muddled too.
The phrases, translated
Architects have a vocabulary that sounds vague until you know what it is protecting you from.
Loose coupling
Two parts are coupled when changing or breaking one forces a change or a failure in the other. Loose coupling means keeping that dependency as small as possible: each part talks to the others through a narrow, stable interface and knows as little as it can about how they work inside.
An event bus (or message broker) is a common way to get there. Instead of the sign-up service calling the email service directly, it publishes an event, 'UserSignedUp', and moves on. Anything interested, such as an email service or an analytics service, subscribes and reacts in its own time. The sign-up service doesn't know they exist. Background jobs work on the same idea.
Here is what that buys you on a day when the email service is down:
Sign-up keeps working while the email service is down
Step 1 of 7: A new user signs up, and the request reaches the sign-up service.
The cost is that the flow is harder to follow. Nothing in the sign-up code tells you an email gets sent, and a message can be delayed, retried or delivered twice, so the email service has to cope with all three.
Cloud-native
Cloud-native describes software built to run on cloud platforms and make use of what they offer: small services packaged in containers, started and stopped automatically, scaled out by adding copies, and configured from outside rather than by hand. It is a reasonable goal when you need that flexibility. When you don't, it is a cost, because every extra moving part is something to deploy, monitor and pay for.
Extensible
Extensible means new behaviour can be added without rewriting what is there: a new payment provider plugs in behind the same interface, or a new consumer subscribes to an existing event. Design for the extensions you can see coming. Designing for every imaginable one is how a box called 'Future Scalability Layer' ends up on the diagram with nobody able to say what it does.
Thinking about scale and failure
Two questions sit behind most architecture decisions: what happens when ten times as many people use this, and what happens when part of it breaks?
How it scales
There are two basic ways to handle more load. Scaling up (vertical) means a bigger machine; it is simple but has a ceiling. Scaling out (horizontal) means more copies of the same service behind a load balancer. That goes much further, but only if each copy is stateless, keeping session data in a shared store rather than in its own memory.
The architect's job is to spot where the bottleneck will be. Very often it is the database, not the application servers, which is why caching and read replicas come up so often in these conversations.
How it fails
Once a system is spread across many machines, something is always failing somewhere; that is the everyday reality of distributed systems. Architects plan for it:
- Timeouts on every network call, so one slow service can't freeze everything waiting on it.
- Retries with backoff for failures that are likely to be temporary.
- No single points of failure where it matters: two copies of anything whose loss takes the whole system down.
- Graceful degradation: if recommendations are down, the shop still sells things, just without suggestions.
These choices trade off against each other and against cost. Running two of everything doubles the bill, so an architect decides which parts deserve it.
A worked example: a cat adoption site
Imagine a team building a site where shelters list cats and people apply to adopt them. Three developers, a few thousand visitors a day, a launch in three months.
A sensible architect starts with the requirements. Adopters browse and apply; shelters manage listings and photos; everyone gets email updates. Uptime matters, but a few minutes of downtime at night would not hurt anyone. That leads to a short list of decisions:
- Build one web application, not microservices. Three developers can't run ten services well, and a single deployable app is faster to build and debug.
- Keep listings, users and applications in one relational database, because the data is closely connected and needs consistent updates.
- Put photos in object storage and serve them through a CDN. They are large, and read far more often than they are written.
- Hand sign-in to a managed identity provider, so the team never stores passwords. It handles authentication, not authorisation: the app still decides what each user may do.
- Send emails through a queue and a background worker, so a slow email provider never slows down an application form.
Here is the whole system, the kind of diagram that comes out of that conversation:
Each decision gets a short written note of the context, the choice and its consequences, usually called an architecture decision record. Six months later, when someone asks why it isn't microservices, the answer is on file: three developers, a few thousand users, and a note on what would make the team revisit it.
Notice what is missing: no 'Future Scalability Layer'. If the site grows, the app can be scaled out, and a busy part such as search can be split off later, once it is clear where the pressure is.
Common mistakes
- Designing for scale you don't have: microservices, event buses and multi-region setups solve real problems, but they add a lot of complexity. Starting with a well-structured single app and splitting it later is usually cheaper than starting split.
- Buzzwords instead of reasons: 'It must be cloud-native' is not a requirement. 'We need to handle ten times the usual traffic on launch day without hiring an ops team' is, and you can test it.
- Diagrams that drift from reality: a diagram nobody updates becomes fiction. Keep a few high-level ones current rather than many detailed ones stale.
- Ivory-tower architecture: an architect who never touches the code loses track of what is hard. When developers nod along and then build something different, the design has failed.
- Skipping the failure question: many outages come from parts everyone assumed would always be there.
Key takeaways
- A system architect decides how the parts of a system fit together: what they are, how they talk, where data lives, and how it all scales and fails.
- Boxes and arrows give the team a shared picture that is cheap to change before any code exists.
- Loose coupling, cloud-native and extensible are means to an end; each has a cost and should be tied to a real requirement.
- Start with the simplest design that meets today's needs, and write down why, so it can change when the needs do.