A large language model on its own can only talk: you send text, it sends text back. An AI agent gives that model a goal, a set of tools it can call and a loop that keeps it working until the job is done, which is the difference between describing how to book a flight and actually booking one.
Chatting is not doing
Ask a chat model to 'create an invoice for Acme' and the best it can do is write out what an invoice might look like. It can't look Acme up in your customer database, it can't call your billing system, and it can't check that the invoice was saved. Everything it knows comes from its training data and the conversation so far.
An agent closes that gap with three additions:
- A goal to work towards, rather than a single question to answer.
- Tools: functions your code exposes, such as a search, a database query, an API call or a payment.
- A loop that lets the model take a step, see what happened and decide on the next step, as many times as the task needs.
The model still does the reasoning. The tools are how that reasoning reaches the outside world. A useful way to hold it in your head: the model is the brain, and the tools are its hands.
The parts of an agent
Every agent, whichever framework or provider you use, is built from the same few parts working together.
The system prompt: the rulebook
The system prompt is the standing instruction the model sees before anything else. It sets out who the agent is, what it is for, which tools it may use and how, and where its limits are. Good system prompts are specific about boundaries: when to ask a human for help, which actions need approval first, and what the agent must never touch.
You are a billing assistant for a small agency.
Use find_customer before creating any invoice.
Never issue a refund. Ask a human instead.
If a customer is not found, stop and say so.Because the model reads the system prompt on every turn, it is the right place for rules that must always hold. It is the wrong place for the task itself.
The user prompt: the goal
The user prompt carries the task: 'Find me an event this weekend', 'Book the flight', 'Create the invoice'. A chat model answers it once. An agent treats it as a goal and keeps working until the goal is met, so a clear, checkable goal matters. 'Create an invoice for Acme for 10 hours of design work' gives the agent something it can finish; 'sort out Acme' does not.
Tools: how the model acts
A tool is a function your code runs on the model's behalf. You describe each one to the model with a name, a plain-language description and a schema for its inputs. The model never runs the tool itself: it replies with a request to call it, your code runs the function, and you hand the result back.
TOOLS = [
{
"name": "find_customer",
"description": "Look up a customer by name.",
"input_schema": {
"type": "object",
"properties": {
"name": {"type": "string"},
},
"required": ["name"],
},
},
]The description is not decoration. The model chooses tools by reading those descriptions, so a vague one ('does customer stuff') leads to wrong or missing calls. Most providers support this pattern under the name 'tool use' or 'function calling', with small differences in the exact format.
Typical tools include a web browser or search, calls to internal APIs, database reads and writes, and payments. The more a tool can change, the more care it needs: reading a customer record is harmless, taking a payment is not.
The agent loop
The loop is what makes an agent an agent. Each pass goes like this:
- Think: the model reads the goal, the rules and everything that has happened so far.
- Pick a tool: it decides which tool to call and with what inputs, or decides it is finished.
- Use it: your code runs the tool.
- Read the result: the output goes back into the conversation.
- Think again, with the new information.
It repeats until the goal is complete or something stops it: a step limit, an error it can't recover from, or a human saying stop. Here is the loop working through the invoice example, calling two different tools:
An agent loop creating an invoice
Step 1 of 7: You give the agent a goal: invoice Acme for 10 hours of work.
In code, the whole loop fits in a dozen lines. Here llm stands in for whichever model SDK you use; the shape is the same across providers:
def run_agent(goal, max_steps=10):
messages = [
{"role": "system", "content": SYSTEM},
{"role": "user", "content": goal},
]
for _ in range(max_steps):
reply = llm.chat(messages, tools=TOOLS)
messages.append(reply.message)
if not reply.tool_calls:
return reply.text # goal met
for call in reply.tool_calls:
result = run_tool(call.name, call.args)
messages.append(
tool_result(call.id, result)
)
raise RuntimeError("Too many steps")Two details matter. The conversation keeps growing, so the model sees every earlier tool call and result when it decides the next step. And max_steps is a hard stop: without it, a confused model can loop for ever, calling tools and spending money.
Why agents are hard to run for real
The loop above works on your laptop. In production, agents are messy:
- They crash. The process restarts, a deployment rolls out, a machine dies. An in-memory loop loses everything: which steps ran, what the tools returned, where it was up to.
- They retry. Model APIs rate-limit and time out, and tools fail. Retrying blindly can repeat an action that already happened, such as creating the same invoice twice.
- They wait. Some steps need a human to approve them, and that might take minutes or days. Holding a process open that long is fragile.
- They are unpredictable. Ask the same model the same question twice and it can choose a different tool or a different plan. You can't rebuild lost state by simply running the model again and hoping it makes the same choices.
You can handle each of these yourself with a database, queues, retry logic and a scheduler. That is a lot of infrastructure around a loop that looked simple.
Making the agent durable with Temporal
Temporal is an open-source platform for durable execution: code that keeps running correctly through crashes, restarts and long waits. You write the agent loop as a Temporal workflow, and every call that talks to the outside world (the model and each tool) as an activity.
Temporal records the result of each activity as the workflow runs. If the process crashes, another worker picks the workflow up and replays it, using the recorded results rather than calling the model again. The agent resumes exactly where it stopped, with the same decisions it had already made. That is how Temporal deals with the model's unpredictability: it doesn't make the model predictable, it remembers what the model actually said.
It also gives you, without extra infrastructure:
- Retries for failed activities, with back-off and limits you configure.
- Durable waits: a workflow can wait for a human's approval for days without holding a process open.
- State that survives restarts, so the conversation and every tool result are never lost.
A sketch of the invoice agent as a workflow, using Temporal's Python SDK, with an approval step before the invoice is created:
@workflow.defn
class InvoiceAgent:
def __init__(self):
self.approved = False
@workflow.signal
def approve(self):
self.approved = True
@workflow.run
async def run(self, goal: str) -> str:
messages = start_messages(goal)
while True:
reply = await workflow.execute_activity(
call_model, messages,
start_to_close_timeout=MODEL_TIMEOUT,
)
messages.append(reply.message)
if not reply.tool_calls:
return reply.text
for call in reply.tool_calls:
if call.name == "create_invoice":
await workflow.wait_condition(
lambda: self.approved
)
result = await workflow.execute_activity(
run_tool, call,
start_to_close_timeout=TOOL_TIMEOUT,
)
messages.append(result)The model call and the tools live in activities because they are unpredictable or touch the outside world. The workflow itself only makes decisions from their recorded results, which is what lets Temporal replay it safely.
When not to build an agent
An agent trades speed, cost and predictability for flexibility. If the steps are always the same (look up the customer, then create the invoice), a fixed piece of code that calls the model once or twice is cheaper, faster and easier to test. Reach for an agent when the path genuinely depends on what the tools return: which tool to use, what information is missing, or when to change the plan.
Common mistakes
- Vague tool descriptions. The model picks tools by reading them. If two tools sound alike, expect it to confuse them.
- No step limit. A loop with no cap can run for ever and run up a large bill.
- Risky tools with no approval. Payments, deletes and emails to customers should wait for a human, or at least be checked by code, before they run.
- Putting the task in the system prompt. The rules belong there; the goal belongs in the user prompt.
- Keeping agent state only in memory. Anything long-running or important needs to survive a crash, whether through Temporal or your own storage.
Key takeaways
- An agent is a model with a goal, tools and a loop that runs until the goal is done or something stops it.
- The system prompt holds the rules; the user prompt holds the goal.
- The model decides which tool to call; your code runs it and hands the result back.
- The loop is think, pick a tool, use it, read the result, think again.
- Temporal makes an agent durable: it records each step so the agent survives crashes, retries safely and waits for approval.