Temporal is a durable execution platform: you write a multi-step job as ordinary code, and Temporal makes sure it finishes, even when the network drops for days or the process running it crashes halfway through. Instead of starting over, the job carries on from the last step that succeeded.
Why long-running jobs break
Take a job that runs every hour: fetch a recipe from an API, check it, save it. As a script on a timer it looks like three lines of work. In practice every step talks to something that can fail.
If the network drops during the fetch, the script throws and that run is lost. So you add a retry loop. Then the machine restarts mid-retry, and you need to remember how far you got, so you store progress in a database. Then a slow run is still going when the next hour starts, so you add a lock to stop two runs overlapping.
None of that code is about recipes. It is plumbing, and every team writes its own slightly broken version. Temporal moves that plumbing out of your code and into a platform built for it.
Workflows and activities
Temporal splits your code into two kinds of function.
- A workflow is the plan: which steps run, in what order, and what happens between them. It can run for seconds or for months.
- An activity is one step that touches the outside world: an HTTP call, a database write, an email. Activities are where failures happen, so they are what Temporal retries.
Two more pieces make it run. The Temporal Service (which you host yourself, or rent as Temporal Cloud) keeps track of every workflow. Workers are your own processes: they poll a task queue on the service, run your workflow and activity code, and report back.
The split matters because of one rule: workflow code must be deterministic. Given the same inputs and the same activity results, it must make the same decisions every time. No network calls, no reading the clock directly, no random numbers. Anything unpredictable belongs in an activity. The next section shows why.
What happens when the connection drops
Retries with backoff
When an activity fails, Temporal retries it according to a retry policy. Unless you set your own, an activity gets the default policy: wait 1 second before the first retry, double the wait each time, cap the wait at 100 times the first interval (so 100 seconds), and keep trying with no limit on attempts.
That last part is what lets a job ride out a long outage. If the network is gone for two days, the fetch fails thousands of times, spaced at most 100 seconds apart. Nobody has to notice or restart anything. As soon as the connection comes back, the next attempt succeeds and the workflow moves on.
The event history
As a workflow runs, the Temporal Service writes each thing that happens to it (a step scheduled, a step finished with its result, a timer fired) to that workflow's event history.
If a worker crashes, or you deploy a new version and restart it, a worker picks the workflow up again by replaying it: it runs the workflow code from the top, but instead of calling activities that already finished, it takes their results from the history. The code reaches the point where it stopped, with the same local variables it had before, and carries on from there. That is why determinism matters: replay only works if the code makes the same decisions the second time.
Here is one run of an hourly recipe workflow through a network outage:
A recipe workflow riding out a network outage
Step 1 of 8: The schedule starts a run, and the start is recorded in the event history.
A worked example in Python
Here is that job with Temporal's Python SDK. First the activities, the steps that touch the outside world:
import os
import httpx
from temporalio import activity
from temporalio.exceptions import ApplicationError
from store import upsert_recipe # your own DB code
@activity.defn
async def fetch_recipe() -> dict:
async with httpx.AsyncClient() as http:
resp = await http.get(os.environ["RECIPES_URL"])
resp.raise_for_status()
return resp.json()
@activity.defn
async def validate_recipe(recipe: dict) -> dict:
if not recipe.get("ingredients"):
raise ApplicationError(
"Recipe has no ingredients",
non_retryable=True,
)
return recipe
@activity.defn
async def save_recipe(recipe: dict) -> None:
# Keyed on the recipe's id, so a repeat is harmless
await upsert_recipe(recipe["id"], recipe)Then the workflow, which only decides the order and how to retry:
from datetime import timedelta
from temporalio import workflow
from temporalio.common import RetryPolicy
with workflow.unsafe.imports_passed_through():
from activities import (
fetch_recipe,
save_recipe,
validate_recipe,
)
RETRY = RetryPolicy(
initial_interval=timedelta(seconds=1),
backoff_coefficient=2.0,
maximum_interval=timedelta(minutes=5),
)
@workflow.defn
class TunaRecipes:
@workflow.run
async def run(self) -> None:
opts = dict(
start_to_close_timeout=timedelta(seconds=30),
retry_policy=RETRY,
)
recipe = await workflow.execute_activity(
fetch_recipe, **opts
)
recipe = await workflow.execute_activity(
validate_recipe, recipe, **opts
)
await workflow.execute_activity(
save_recipe, recipe, **opts
)A few things to notice:
- The workflow reads like a plain function. There is no retry loop, no progress table and no 'where was I?' logic: Temporal handles all of it.
start_to_close_timeoutlimits how long a single attempt may take. If an attempt hangs, Temporal gives up on it and retries. The Python SDK makes you set this timeout, or a totalschedule_to_close_timeout.- The retry policy here waits at most five minutes between attempts. It sets no maximum number of attempts, so it keeps trying through an outage of any length.
- A recipe with no ingredients will never become valid by trying again, so
validate_reciperaises a non-retryable error and the run fails straight away instead of looping forever.
A worker process registers TunaRecipes and the three activities on a task queue and runs them. To run the workflow every hour, you create a schedule on the Temporal Service:
from datetime import timedelta
from temporalio.client import (
Client,
Schedule,
ScheduleActionStartWorkflow,
ScheduleIntervalSpec,
ScheduleSpec,
)
from workflows import TunaRecipes
async def main() -> None:
client = await Client.connect("localhost:7233")
every_hour = ScheduleIntervalSpec(
every=timedelta(hours=1)
)
await client.create_schedule(
"tuna-recipes-hourly",
Schedule(
action=ScheduleActionStartWorkflow(
TunaRecipes.run,
id="tuna-recipes",
task_queue="recipes",
),
spec=ScheduleSpec(intervals=[every_hour]),
),
)A schedule's default overlap policy is Skip: if the previous run is still going when the next hour comes round, the new run isn't started. That keeps it to one recipe at a time. It also means that during a two-day outage the stuck run carries on and the hourly runs in between are skipped, not stacked up. If you would rather have one run waiting to start as soon as the current one finishes, set the overlap policy to Buffer One.
Common mistakes
- Doing I/O in workflow code. A network call or a database read inside the workflow breaks replay, because it can return something different the second time. Put it in an activity. For time and randomness, use the SDK's own versions, such as
workflow.now()rather thandatetime.now(). - Assuming an activity runs exactly once. A worker can finish the work and crash before it reports back, so Temporal runs the activity again. Make activities idempotent: an upsert keyed on an id, not a blind insert.
- Retrying errors that can't succeed. A timeout is worth retrying; a malformed recipe or a failed permission check is not. Mark those as non-retryable, or the workflow will keep trying forever. The general ideas are in our error handling video.
- Passing big payloads. Every activity's input and output is stored in the event history. Pass ids and small records, and keep large files in storage of their own.
- Changing workflow code under running workflows. A workflow that started on the old code replays on the new code. If the new code makes different decisions, replay fails. Temporal's docs cover versioning for exactly this.
When to use it
Temporal fits processes with several steps across services that need to finish, even over hours or weeks: order fulfilment, payments and refunds, sign-up flows that wait for an email confirmation, data pipelines, and AI agents that call slow or flaky APIs.
It is overkill for a single request and response, or for a timer job with one step and nothing to remember. A message queue already gives you at-least-once delivery for one step. Temporal's value is tracking a whole sequence of steps and its state.
The costs are real. It is another system to run (the Temporal Service and its database) or to pay for as Temporal Cloud, and the determinism rule takes some getting used to.
Key takeaways
- Temporal runs workflows durably: a failed step is retried, and a crashed worker resumes from the event history instead of starting over.
- Workflows are deterministic plans; activities do the I/O and are retried, with unlimited attempts by default.
- Replay rebuilds a workflow's state from recorded results, so finished steps are not run again.
- Make activities idempotent, and mark errors that will never succeed as non-retryable.
- It earns its keep on long, multi-step processes across unreliable services, not on a single request.