A background job is work your app does after it has already answered the user: converting a file, sending an email, building a report. It keeps requests fast, but it moves the hard part somewhere else, because a job that runs later can also fail later, halfway through, with nobody watching.
Why some work can't happen in the request
A web request should finish in well under a second. Some work simply takes longer than that. Take a user uploading a 300-page PDF. The app has to:
- Check the file is a valid PDF.
- Convert each page to an image.
- Generate small preview thumbnails.
- Store the results.
- Email the user when it's all ready.
Doing all of that before responding would leave the user staring at a spinner for minutes, and many servers and proxies cut off a request long before then. Worse, if the user closes the tab, the work may be lost with the connection.
So the app splits the job in two. The request does the quick part: it saves the upload, records that a job is waiting, and replies straight away with something like 'Processing, we'll email you'. The slow part runs somewhere else, in its own time.
How a job queue works
The usual shape has three parts:
- The producer: your web app. It puts a small message on a queue describing the job, such as the document's id. It doesn't send the PDF itself; the file stays in storage and the message just points at it.
- The queue: a store that holds jobs until something takes them. Redis, RabbitMQ, Amazon SQS and a plain database table are all common choices.
- The workers: separate processes that take a job off the queue, do the work, and mark it done.
Because workers are separate processes, you can run as many as the load needs, and they can be on different machines from the web app. A burst of uploads just makes the queue longer for a while; the web app stays responsive, and the workers catch up. This is often called queue-based load levelling.
Most job libraries look similar in code. Here is the idea with Celery in Python:
@app.task
def process_pdf(doc_id):
doc = load(doc_id)
validate(doc)
for page in doc.pages:
convert(page)
make_preview(page)
store_results(doc)
send_email(doc.owner)
# in the request handler
process_pdf.delay(doc.id)
return {"status": "processing"}delay puts a message on the queue and returns at once. A worker runs process_pdf later.
What goes wrong when a worker dies
The happy path is easy. The trouble starts when a job fails part-way through.
Picture one unusually large upload. The worker's memory climbs while it converts pages, the operating system kills the process, and the job stops somewhere around page 4. Most queues are built for exactly this: a job that was taken but never confirmed as finished goes back on the queue, and another worker picks it up. That is good, because the work isn't lost.
But the job above is one function. The new worker runs it from the top. Pages 1 to 3 are converted again, their previews are generated again, and any notifications the first attempt already sent, such as 'your upload is processing', go out a second time. For a PDF that's wasteful and a little embarrassing. For a job that charges a card or sends a payout, running a step twice is a real bug.
Make each step safe to repeat
The first defence is making steps idempotent: running one twice has the same effect as running it once. Two common ways to get there:
- Check before doing. Before converting page 4, see whether its image already exists, and skip it if so.
- Record what's done. Keep a flag or a unique key for side effects, so a second attempt sees 'email already sent for doc 42' and stops.
def send_ready_email(doc):
key = f"ready-email:{doc.id}"
if already_done(key):
return
send_email(doc.owner)
mark_done(key)This helps a lot, but notice the gap: if the worker dies after send_email and before mark_done, the email still goes twice. Idempotency makes repeats cheap; it doesn't make them impossible. And you now have hand-written bookkeeping around every step.
Split the job into smaller jobs
The second defence is breaking one big job into many small ones: a job per page, then a final job that sends the email once every page has finished. A crash now only repeats one page. The cost is coordination. Something has to know when all 300 page jobs are done, handle the ones that failed, and make sure the email job runs exactly when it should. You end up writing a small workflow engine on top of your queue.
Durable execution with Temporal
That coordination is what tools like Temporal take off your hands. Instead of a job function, you write a workflow: ordinary code that describes the steps in order. Each step that touches the outside world, such as converting a page or sending an email, is an activity.
The difference is that Temporal records every step's result in a durable history as it happens. The workflow's progress doesn't live in a worker's memory, so a worker dying doesn't lose it.
@workflow.defn
class ProcessPdf:
@workflow.run
async def run(self, doc_id: str, pages: int):
t = timedelta(minutes=5)
for page in range(1, pages + 1):
await workflow.execute_activity(
convert_page,
args=[doc_id, page],
start_to_close_timeout=t,
)
await workflow.execute_activity(
send_ready_email,
doc_id,
start_to_close_timeout=t,
)Each execute_activity call is a checkpoint. When page 3 finishes, Temporal writes that down. If the worker then dies on page 4, another worker picks the workflow up, Temporal replays the history to rebuild where it was, sees pages 1 to 3 already done, and carries on from page 4. Failed activities are retried automatically, with a retry policy you can tune.
Here is the crash and the resume, step by step:
A worker dies and another resumes at page 4
Step 1 of 7: The web app starts a workflow for the uploaded PDF and answers the user at once.
Two things are worth being clear about. First, activities still run 'at least once': if a worker dies in the middle of an activity, that one activity runs again. So each activity should still be idempotent, but now it's only one page or one email, not the whole job. Second, the workflow code itself must be deterministic, because Temporal replays it to rebuild state. Anything random, time-based or involving I/O belongs in an activity, not directly in the workflow.
Seeing what happened
Because the history is stored, you can look at it. The Temporal web UI shows each workflow that is running, which have finished, which steps were retried and why a failed one failed. With a plain queue you usually get a job id and a log line, and piecing together what a job did means searching logs across machines.
When to use which
A plain queue with idempotent jobs is the right choice for most background work: sending one email, resizing one image, clearing old sessions. The jobs are short, a repeat is harmless, and a queue library is simple to run.
Reach for durable execution when a job is long, has many steps, has side effects that must not repeat, or needs to wait on things like a payment confirmation or a human approval. Order processing, onboarding flows and multi-step file pipelines are typical examples. The cost is another system to run (or pay for) and a new way of writing code, with rules such as workflow determinism that take some getting used to.
Common mistakes
- Putting the whole payload on the queue. Send an id and keep the file in storage. Queues often have small message size limits, and a stale copy of data in a message can go out of date.
- Assuming a job runs exactly once. Almost every queue delivers at least once. Design every job and activity for a repeat.
- One giant job. A job that does everything repeats everything when it fails. Smaller steps fail smaller.
- No limits on input size. The crash in this story came from one huge file. Check sizes up front and give workers memory headroom.
- Retrying forever. A job that will never succeed, such as a corrupt file, should fail clearly after a few attempts and tell someone, not loop.
Key takeaways
- Background jobs keep requests fast by moving slow work to workers that pull from a queue.
- Workers can die part-way, and most queues then run the whole job again.
- Idempotent steps make a repeat harmless; smaller steps make a repeat cheap.
- Durable execution tools like Temporal record each step, so a new worker resumes where the old one stopped.
- Activities can still run more than once, so keep each one safe to repeat.