A deterministic system gives the same result every time you give it the same input. An AI agent does not, and once you mix the two in one system, a crash can send the whole thing down a different path unless something remembers what already happened.
What deterministic means
A system is deterministic when the same input always produces the same steps and the same output. Run it once or a thousand times and nothing changes. Most everyday code works like this: a function that adds up a basket, a build script, a database query against the same data.
def basket_total(prices):
return sum(prices)
basket_total([3, 4, 5]) # 12, every timeThat predictability is what makes ordinary software testable. You can write a test that says 'this input gives 12', run it in CI a hundred times a day, and trust it. When a bug appears, you can replay the input and watch the same bug happen again.
What non-deterministic means
A system is non-deterministic when the same input can produce different results on different runs. The goal stays the same, but the output does not.
import random
def pick_greeting():
options = ["Hi", "Hello", "Hey"]
return random.choice(options) # may differ each callRandomness is the obvious source, but not the only one. Anything that depends on the outside world is non-deterministic from your code's point of view: the current time, a network call that might time out, a third-party API that returns different data tomorrow, or two threads racing each other.
Why AI agents are non-deterministic
An AI agent is a loop around a language model: it reads some input, decides what to do, calls a tool, reads the result and decides again. Several things make that loop unpredictable:
- The model samples its output. A language model picks each next word from a range of likely options, so the same prompt can produce a different answer on another run. Lowering the temperature makes answers more consistent, but many hosted models still don't promise identical output for identical input.
- Decisions compound. If the agent picks a different tool on step one, every step after that sees different data. A small wobble early on becomes a completely different path.
- The world changes underneath it. Tools hit real services, which can be slow, rate-limited or down. A time-out is just another result the agent has to react to, and it may react differently each time.
None of this is a bug. Flexibility is the reason to use an agent: it can handle an input nobody wrote a rule for. But it means you cannot treat an agent's output like the result of basket_total.
Where the two meet
Real systems are rarely one or the other. A typical AI-powered workflow is a chain of steps, some deterministic and some not. Take a support ticket system:
- Load the ticket and the customer's order history (deterministic).
- Ask an AI agent to classify the ticket: refund, bug report or question (non-deterministic).
- If it is a refund, call the payments API to issue it (a side effect).
- Email the customer (another side effect).
Now suppose the server running this crashes after step 2. The simple fix is to start again from the top. That is fine for a deterministic system, because running it again gives the same answer. Here it is a problem: the agent runs again, makes a fresh decision, and this time might call the ticket a bug report. The system has now taken two different paths for one ticket. If the crash happened after step 3 instead, restarting could refund the customer twice.
The underlying issue is that a restart throws away the one thing you cannot reproduce: the decision the agent already made.
Durable execution
Durable execution fixes this by recording what each step did as it happens, so that after a crash the system carries on from where it stopped instead of starting over. The work survives the process that was running it.
Temporal is an open-source platform built around this idea. You split your code into two kinds of function:
- A workflow is the orchestration: which steps run, in what order, and what happens on each branch. Workflow code must be deterministic.
- An activity is a single step that talks to the outside world: an AI model call, an API request, a database write. Activities are where non-deterministic work belongs.
As the workflow runs, Temporal writes every activity's result to an event history, stored by the Temporal service rather than in your server's memory. If the worker process dies, another one picks the workflow up and replays its code from the start. When the replay reaches an activity that already finished, Temporal does not run it again; it hands back the result from the history. The workflow rebuilds its exact state, then continues with the first step that had not finished yet.
Here is how the support ticket flow looks in Temporal's Python SDK. llm and payments stand in for your own clients:
from datetime import timedelta
from temporalio import activity, workflow
@activity.defn
async def classify_ticket(text: str) -> str:
# Non-deterministic: the model may vary
return await llm.classify(text)
@activity.defn
async def issue_refund(order_id: str) -> None:
await payments.refund(order_id)
@workflow.defn
class TicketWorkflow:
@workflow.run
async def run(self, text: str, order_id: str) -> str:
label = await workflow.execute_activity(
classify_ticket,
text,
start_to_close_timeout=timedelta(seconds=60),
)
if label == "refund":
await workflow.execute_activity(
issue_refund,
order_id,
start_to_close_timeout=timedelta(seconds=30),
)
return labelFollow a crash through it, one step at a time:
A workflow resuming after a crash
Step 1 of 6: The workflow starts and asks the AI agent to classify the ticket.
The agent is still as unpredictable as ever. What changed is that each of its decisions is made once, recorded, and then treated as fact for the rest of that run. The chaos stays inside one step instead of leaking into the whole system.
Temporal also retries a failed activity according to a retry policy, so a model call that times out is tried again without you writing the retry loop yourself.
Common mistakes
- Putting non-deterministic code in the workflow. Calling the model, reading the clock with
datetime.now()or usingrandomdirectly inside workflow code breaks replay: the second run makes different choices from the first, and Temporal reports a non-determinism error. Put outside calls in activities, and use the SDK's ownworkflow.now()andworkflow.random()when the workflow needs a time or a random value. - Assuming an activity runs exactly once. If a worker crashes midway through an activity, before its result is recorded, the activity runs again. Make side effects safe to repeat, for example by sending the payments API an idempotency key so a second refund request is ignored.
- Changing workflow code under running workflows. A workflow that started on the old code replays against the new code. If the steps no longer line up, replay fails. Temporal has versioning tools for this; plan for them before you ship a change.
- Expecting durable execution to make the agent consistent. It makes the system recoverable, not the model predictable. Two different tickets, or the same ticket in two separate runs, can still be classified differently. Validate the agent's output before acting on it.
When to use it
Durable execution earns its place when a process has several steps, takes long enough that a crash mid-way is realistic, and includes steps you cannot safely repeat: expensive model calls, payments, emails, anything a user would notice happening twice. Agent pipelines that run for minutes or hours fit this well.
It costs something. You run a Temporal service (self-hosted or the managed cloud), learn its programming model, and accept the deterministic-workflow rules above. For a single model call behind an API endpoint that can simply be retried, that is more machinery than the problem needs. Temporal is also one of several durable execution tools; the ideas here carry over to the others.
If you want to see how an agent's loop is put together in the first place, Building an AI Agent walks through it, and Context Engineering covers how to make each of its decisions better informed.
Key takeaways
- Deterministic means same input, same steps, same result, every time.
- AI agents are non-deterministic: the same task can produce a different answer or a different path on each run.
- Restarting a mixed system after a crash re-runs the agent, which can take a new path or repeat side effects.
- Durable execution records each step's result, so a workflow resumes where it stopped instead of starting over.
- In Temporal, keep workflow code deterministic and put model calls and other outside work in activities.