Durable execution is a way of running code so that a crash, a restart or a long wait doesn't lose its place: the program picks up from the last step it finished, with its state intact. It matters most for work with several steps that touch the outside world, such as taking a payment and then shipping an order.
What goes wrong without it
Picture an order service. For each order it charges the customer, waits for the payment provider to confirm, prepares the order and sends it. In ordinary code that is four lines in one function, and the only record of how far it got lives in memory.
Now the process dies between the charge and the delivery. A deploy restarts it, the machine runs out of memory, or the host reboots. The local variables are gone, and when the service comes back nothing knows the order was halfway through. Either the order sits stuck, with the customer charged and nothing sent, or someone retries it from the start and charges the customer twice.
The longer a process runs, the more likely this is. Waiting for a payment confirmation can take minutes. A sign-up flow that sends a reminder after three days has to survive every deploy in those three days.
The usual fixes, and why they hurt
Teams usually build the safety net by hand:
- a status column on the order row ('charged', 'prepared', 'sent'), updated after every step;
- retries with backoff around every network call;
- a queue between steps, so work survives a restart;
- a cron job that sweeps for orders stuck in one status for too long.
Each piece works. Together, they scatter a four-line idea across handlers, tables, queues and scheduled jobs. The business logic, charge then wait then prepare then deliver, becomes a state machine you maintain yourself, and every new step adds another status, another retry and another edge case to test. If you have built background jobs before, this will look familiar.
What durable execution changes
With durable execution you write the process as one ordinary function, top to bottom, and the runtime keeps track of its progress. Each step's result is saved as soon as that step completes. If the process dies, another worker runs the function again, and every step that already has a saved result hands back that result instead of running a second time. The code reaches the first unfinished step and carries on from there.
Waiting is durable too. A workflow that sleeps for three days isn't holding a thread open for three days: the timer is stored on the server, and the workflow wakes up on whichever worker is free when it fires.
Several tools work this way, including Temporal (the open source tool in the video), Azure Durable Functions, AWS Step Functions and Cloudflare Workflows. Their APIs differ, but the idea is the same. The rest of this page uses Temporal.
How Temporal does it
Workflows and activities
Temporal splits your code in two:
- A workflow is the orchestration: the order of the steps, the waits and the decisions. It must be deterministic, which means that given the same inputs and the same step results it makes the same choices every time.
- An activity is one step that does real work in the outside world: charging a card, sending an email, calling an API. Activities are allowed to fail, and Temporal retries them according to a retry policy.
The event history
Your code runs in workers, processes you start and scale yourself. The Temporal service sits in the middle and keeps an event history for every workflow run: when it started, which activity was scheduled, what that activity returned, when a timer fired, which signal arrived. Workers poll the service for work and report each result back, so the history, not any worker's memory, is the record of how far each run has got.
Replay
When a worker picks up a workflow after a crash, it runs the workflow code again from the top. Each time the code asks for an activity, the SDK checks the history first. If the activity already completed, it returns the recorded result without calling it. This catch-up takes moments, because no activity is called, and once the code passes the end of the history it continues live.
Here is a crash in the middle of an order, and how the next worker resumes:
A worker crashes mid-order, and another resumes it
Step 1 of 9: An order arrives, and Worker A starts the workflow.
A worked example
Here is the order from the video as a Temporal workflow in Python. The activities are stubs; in a real service they would call the payment provider and the kitchen.
from datetime import timedelta
from temporalio import activity, workflow
from temporalio.common import RetryPolicy
@activity.defn
async def charge_customer(order_id: str) -> str:
# Call the payment provider here, with
# order_id as the idempotency key.
return f"charge-{order_id}"
@activity.defn
async def prepare_tuna(order_id: str) -> None:
print(f"Preparing tuna for {order_id}")
@activity.defn
async def deliver_tuna(order_id: str) -> None:
print(f"Delivering {order_id}")
@workflow.defn
class TunaOrder:
def __init__(self) -> None:
self.confirmed = False
@workflow.signal
def payment_confirmed(self) -> None:
self.confirmed = True
@workflow.run
async def run(self, order_id: str) -> str:
timeout = timedelta(seconds=30)
await workflow.execute_activity(
charge_customer,
order_id,
start_to_close_timeout=timeout,
retry_policy=RetryPolicy(maximum_attempts=5),
)
await workflow.wait_condition(
lambda: self.confirmed
)
await workflow.execute_activity(
prepare_tuna,
order_id,
start_to_close_timeout=timeout,
)
await workflow.execute_activity(
deliver_tuna,
order_id,
start_to_close_timeout=timeout,
)
return f"Order {order_id} delivered"It reads like the four steps it describes, and the durability comes from what Temporal does around it:
- A crash after the charge costs nothing. The next worker replays the history, gets the recorded charge back, and goes straight to waiting for confirmation.
wait_conditionpauses the workflow until the payment provider's webhook sends thepayment_confirmedsignal, whether that takes a second or a day. No worker sits blocked while it waits.- Activities are retried by default, and the policy on the charge caps it at five attempts, so a payment provider that is down for a minute doesn't fail the order.
To run it you also need a worker process that registers the workflow and activities, and a client that starts a TunaOrder. The Python getting-started guide in the useful links walks through both.
Common mistakes
Non-deterministic workflow code
Replay only works if the workflow makes the same decisions the second time. Calling random(), reading datetime.now() or making a network request directly in workflow code breaks that: on replay the value differs, the code can take another branch, and Temporal stops the run with a non-determinism error. Use the SDK's replay-safe versions (workflow.now(), workflow.random()), or move the work into an activity. The Python SDK runs workflow code in a sandbox that blocks some of these calls for you.
Assuming an activity runs exactly once
If a worker charges the card and then crashes before reporting the result, the history has no record of the charge, so Temporal retries the activity. Activities run at least once, not exactly once. Make the side effects idempotent: pass the payment provider an idempotency key such as the order id, so a repeat charge is recognised and ignored.
Changing a workflow while runs are in flight
A workflow that started last week replays against last week's history. If you reorder or remove steps in the code, old runs no longer match their histories. Temporal's versioning lets old runs keep the old path while new runs take the new one.
Passing large data through the workflow
Every input and result is stored in the event history, and the history has size limits. Pass ids and keep files and large payloads in your own storage.
When to use it
Durable execution fits processes that have several steps across services and must finish: payments and order fulfilment, multi-day onboarding flows, provisioning cloud resources, data pipelines, and AI agents that make many slow tool calls. It suits any process that waits on a human or on another system.
It is overkill for a request that reads or writes one database row, or for a fire-and-forget job where losing the odd run doesn't matter. It also has costs. You either operate a Temporal cluster with its own database or pay for Temporal Cloud. The determinism rules above take some getting used to. And every activity result is recorded before the workflow moves on, which adds a little latency to each step.
It sits alongside ideas you may already use. A queue moves work off the request path; durable execution keeps a multi-step process's place. A saga, where each step has a compensating action such as a refund if delivery fails, becomes an ordinary try and except in workflow code. For the wider picture of why machines and networks fail mid-task, see distributed systems.
Key takeaways
- Durable execution saves each step's result, so a crashed process resumes where it stopped instead of starting over or getting stuck.
- In Temporal, workflow code orchestrates and must be deterministic; activities do the real-world work and are retried.
- After a crash, a worker replays the event history: finished steps return their recorded results instead of running again.
- Activities run at least once, so make their side effects idempotent.
- It pays off for long, multi-step processes that must finish, and is overkill for a single request.