Error handling is how your code behaves when something it depends on lets it down. Handle one call well and a failure becomes a clear message or a quiet retry. Handle a multi-step job badly and a failure leaves half-finished work that nobody notices until a customer complains.
What actually goes wrong
Most real code talks to things it doesn't control: an API over the network, a database, an email provider. Each of those can fail, and they fail in different ways.
- Transient failures go away on their own. A request times out, a database is briefly overloaded, a service returns '503 Service Unavailable'. Trying again a moment later often works.
- Permanent failures won't fix themselves. The card is declined, the input is invalid, the record doesn't exist. Retrying just repeats the same error.
- Silent failures are the worst kind. A call fails, nothing checks, and the code carries on as if it worked. The order page says 'thanks', but no order was saved.
Good error handling starts by telling these apart, because each one needs a different response: retry the first, report the second, and never allow the third.
Catching errors with try/catch
A try block wraps the risky code. If anything inside it throws, control jumps straight to the catch block, where you decide what to do instead of letting the error crash the program.
Here is a small TypeScript function that fetches a price from an API:
async function getTunaPrice(): Promise<number> {
try {
const res = await fetch(PRICE_URL, {
signal: AbortSignal.timeout(5000),
});
if (!res.ok) {
throw new Error(`Price API: ${res.status}`);
}
const body = await res.json();
return body.price;
} catch (err) {
logger.error("Price lookup failed", { err });
throw new UserFacingError(
"Prices are unavailable. Try again soon."
);
}
}Three habits are worth noticing:
- Check the result, not just for exceptions.
fetchdoesn't throw on a 500 response, so the code checksres.okitself. Many libraries report failure through a return value, and those failures never reach acatchunless you turn them into one. - Log the real error, show a useful one. The log keeps the details a developer needs. The user gets a message they can act on, not a stack trace.
- Put a limit on waiting. The timeout turns a call that would hang forever into an error you can handle.
Retrying safely
For transient failures, a retry is often the right response. Two rules keep retries from making things worse:
- Back off between attempts. Wait a little longer each time (say 1 second, then 2, then 4) and stop after a fixed number of tries. Retrying instantly in a tight loop piles more load on a service that is already struggling.
- Only retry what is safe to repeat. Reading a price twice is harmless. Charging a card twice is not. An operation that gives the same result however many times it runs is called idempotent, and only idempotent steps should be retried blindly. Payment providers usually support an idempotency key: send the same key with a retry and they recognise it as a repeat rather than a new charge.
Where try/catch runs out
try/catch works inside one running process. It can't help with a failure that takes the process down with it.
Picture an order checkout with four steps: create the order, charge the customer, save the result, send a confirmation email. Now suppose the server crashes, or is restarted by a deployment, straight after the charge succeeds.
- The
catchblock never runs, because there is no program left to run it. - Everything the code knew (which steps had finished, the payment id) lived in memory, and it is gone.
- When the service comes back, nothing tells it that a half-finished order exists. The customer has paid, and there is no saved order and no email.
You can patch this by hand: write progress to a database after every step, run a background job that finds stuck orders, work out which ones can be safely resumed. Every team that does this ends up building the same machinery, and the edge cases are where the bugs live.
Durable execution with Temporal
Temporal is an open-source platform that provides that machinery for you. It is built around the idea of durable execution: your code runs as if a crash can't interrupt it, because Temporal keeps enough of a record to pick up where it left off.
You split the job into two kinds of code:
- A workflow is the plan: the order of the steps and the decisions between them. It describes what should happen, and doesn't talk to the outside world itself.
- Activities are the individual steps that do real work and might fail: calling the payment API, writing to the database, sending the email.
Temporal records every step's result in the workflow's event history, stored on the Temporal service rather than in your server's memory. If a worker process dies, another worker picks the workflow up, replays the history to rebuild where it was, and carries on from the first step that hadn't finished. Completed steps are not run again.
Here is that recovery for the checkout, step by step:
A checkout workflow surviving a server crash
Step 1 of 8: Worker A runs the workflow. The order is created and Temporal records it.
What a workflow looks like
With Temporal's TypeScript SDK, the checkout reads like ordinary code. The activities are declared with a timeout and a retry policy:
import { proxyActivities } from "@temporalio/workflow";
import type * as acts from "./activities";
const {
createOrder, chargeCard, saveOrder, sendEmail,
} = proxyActivities<typeof acts>({
startToCloseTimeout: "30 seconds",
retry: {
initialInterval: "1 second",
backoffCoefficient: 2,
maximumAttempts: 5,
},
});
export async function checkout(o: Order) {
const id = await createOrder(o);
const payment = await chargeCard(id, o.total);
await saveOrder(id, payment);
await sendEmail(o.email, id);
}There is no retry loop and no progress table. If sendEmail fails because the email service is down, Temporal retries just that activity with backoff, without running the charge again. If the worker dies halfway through, the workflow resumes on another worker.
What Temporal asks of you
Durable execution isn't free. It comes with rules:
- Workflow code must be deterministic. Replay only works if running the workflow code again makes the same decisions in the same order. Anything that could come out differently on a second run, such as a network call, a file read or the current time, must either go in an activity or use the SDK's replay-safe version.
- Activities should still be idempotent. Temporal records an activity's result once it reports back. If a worker crashes after the payment API took the money but before the result was recorded, the charge activity runs again. Pass an idempotency key (the order id works well) so the payment provider treats it as the same charge.
- Some errors shouldn't be retried. A declined card will be declined on the fifth attempt too. Mark errors like that as non-retryable so the workflow can handle them straight away, for example by cancelling the order and telling the customer.
- It's another system to run. You need a Temporal service, either self-hosted or the managed Temporal Cloud, and workers that connect to it.
When to use which
try/catch and Temporal solve different problems, and most systems need both.
| Situation | Reach for |
|---|---|
| One call that might fail | try/catch, maybe a retry |
| Input the user can fix | Validate and report |
| Many steps that must all finish | A durable workflow |
A single request that reads data or does one write doesn't need a workflow engine. If it fails, the user sees an error and tries again. A durable workflow earns its place when a job has several steps with side effects in other systems, takes longer than one request, or would leave real damage if it stopped halfway: payments, sign-up flows that provision accounts, data pipelines, anything that waits hours or days between steps.
Temporal isn't the only way to get there. Message queues with retries, the outbox pattern and hand-built state machines all aim at the same goal. Temporal's appeal is that you write the steps as plain code and it takes care of the recording and resuming.
Common mistakes
- Swallowing errors. An empty
catch {}turns a loud failure into a silent one. If you catch an error, log it or handle it. - Catching too broadly. Wrapping a whole function in one try/catch hides where the failure came from. Catch close to the risky call, where you know what went wrong.
- Retrying everything. Retrying a validation error or a declined payment wastes time and can make things worse. Retry transient failures only.
- Retrying without idempotency. A retry that isn't safe to repeat can charge twice, send two emails or create duplicate records.
- Assuming the process will finish. Deploys, crashes and restarts happen mid-request. Any job that can't safely stop halfway needs its progress stored somewhere outside memory.
Key takeaways
- Tell failures apart: retry transient ones, report permanent ones, and never let one pass silently.
- try/catch handles errors inside a running process: log the detail, show a useful message, retry with backoff where it's safe.
- A crash between steps is beyond try/catch, because the process and everything in its memory are gone.
- Temporal records each step's result outside your server, so a workflow resumes after a crash without redoing finished steps.
- Error handling is designing for failure: decide in advance what happens when each step goes wrong.