AWS Lambda runs your code only when something asks for it, and charges you only while it runs. You hand Amazon a function rather than a server, and it finds a machine, runs the function and tidies up afterwards.
What Lambda runs
Lambda is a 'function as a service'. Instead of renting a virtual machine and keeping it running, you upload a function: a handler plus whatever libraries it needs. You pick a runtime such as Python, Node.js, Java or .NET, and you set how much memory the function gets, from 128 MB up to 10,240 MB. CPU power scales with the memory you choose, so a slow function sometimes just needs more of it.
'Serverless' doesn't mean there are no servers. It means they aren't yours to look after. You never patch an operating system, choose an instance size or decide how many machines to run. AWS does all of that and gives you back one thing to think about: the code.
The trade is control. Each run can last at most 15 minutes, you can't log in to the machine, and anything you need to keep has to live somewhere else, such as a database or a storage bucket.
Events start everything
A Lambda function sits there doing nothing until an event invokes it. Events come from many places:
- an HTTP request, through API Gateway or a function URL;
- a file landing in an S3 bucket;
- a message arriving on an SQS queue;
- a schedule in EventBridge, like a cron job.
How the caller waits depends on the source, and it changes how you handle failure:
- Synchronous: the caller waits for the answer. An HTTP request works this way, so an error goes straight back to the user.
- Asynchronous: Lambda queues the event, replies 'accepted' at once and runs the function afterwards, retrying if it fails. S3 notifications work this way.
- Polling: for queues and streams, Lambda reads messages in batches and hands each batch to your function.
A worked example
Here is a small Python handler that saves an order sent to it over HTTP, through a function URL or API Gateway:
import json
import os
import boto3
# Runs once per environment, not per request
table = boto3.resource("dynamodb").Table(
os.environ["ORDERS_TABLE"]
)
def handler(event, context):
body = json.loads(event["body"])
order = {
"id": context.aws_request_id,
"item": body["item"],
}
table.put_item(Item=order)
return {
"statusCode": 201,
"body": json.dumps(order),
}Three things are worth noticing:
eventis the input. Its shape depends on the source: an HTTP event has abody, an S3 event lists the files that changed.contextcarries details about this run, such as a unique request id and how much time is left before the timeout.- The DynamoDB table is set up outside the handler. That code runs once when Lambda prepares an environment, and every later request that lands in the same environment reuses it.
That last point only makes sense once you know how Lambda runs your code.
Cold starts and warm starts
Every function runs inside an execution environment: a small, isolated micro virtual machine holding the runtime and your code. When a request arrives and no environment is free, Lambda creates one. It downloads your code, starts the runtime and runs your initialisation code. This is a cold start, and it adds latency to that one request: from tens of milliseconds to several seconds, depending on the runtime, the size of your package and how much your init code does.
After the handler returns, Lambda doesn't throw the environment away straight away. It freezes it and keeps it for a while. If another request arrives, Lambda thaws the same environment and runs only the handler. That is a warm start, and it is where the 'milliseconds' comes from. If no requests come for a while, Lambda shuts the environment down, and the next request pays for a cold start again.
Here is one cold request followed by a warm one:
A cold start, then a warm start
Step 1 of 8: A request arrives, and no environment is ready for this function.
You can shrink cold starts by keeping your package small and your init code lean. For latency-sensitive functions, provisioned concurrency keeps a set number of environments initialised and waiting, and you pay for them while they wait. Some runtimes, Java among them, also support SnapStart, which restores an environment from a saved snapshot instead of initialising it from scratch.
How it scales
By default, an environment handles one request at a time. If 50 requests arrive together, Lambda runs up to 50 environments side by side. The number of requests in flight at once is the function's concurrency.
That makes scaling automatic, but not unlimited. Each AWS account has a concurrency limit per region, shared by all its functions, and you can set reserved concurrency on a function to guarantee it a share or to cap it. Capping matters when the function talks to something that can't scale with it: 500 environments each opening a connection can overwhelm a relational database.
What you pay for
You pay per request and for how long your code runs, measured in milliseconds and weighted by the memory you gave the function. No requests means no compute bill: there is no idle machine to pay for.
The services around the function bill separately, though: API Gateway, CloudWatch logs, data transfer and provisioned concurrency all have their own prices. And the pay-per-use model has a crossover point. For spiky or low traffic it is usually far cheaper than a server left running. For heavy, steady traffic around the clock, a server or container you keep busy can cost less than paying for every millisecond of every request.
When to use it
Lambda fits work that comes in bursts and finishes quickly:
- APIs with uneven traffic, or a quiet internal tool;
- glue between AWS services, such as making a thumbnail whenever an image is uploaded;
- scheduled jobs, like a nightly clean-up;
- processing messages from a queue.
It fits badly when a job runs longer than 15 minutes, when traffic is heavy and constant, when every request must be fast and a cold start is unacceptable, or when you need hardware Lambda doesn't offer, such as a GPU. If you are still deciding where your code should run at all, Cloud vs. on-prem covers that bigger choice.
Common mistakes
- Keeping state in memory. A variable or a file in
/tmpmay survive to the next request if it lands in the same environment, but nothing guarantees that. Anything that must persist belongs in a database or storage bucket. - Building clients inside the handler. Creating a database client or loading a model on every request repeats work that should happen once per environment. Move it above the handler.
- Ignoring the timeout. The default is only 3 seconds. Set it to what the function really needs, and make sure it is shorter than whatever is waiting for the answer.
- Forgetting retries. Asynchronous events are retried after a failure, so the same event can run twice. Make handlers safe to repeat, for example by keying writes on an id.
- Triggering yourself. A function that runs on uploads to a bucket and writes its output to the same bucket can call itself in a loop, and bill you for every call. Write to a different bucket or prefix.
Key takeaways
- Lambda runs a function when an event arrives, and you never manage the servers underneath.
- You pay per request and per millisecond of run time, so an idle function costs nothing to run.
- A cold start creates and initialises a new environment; a warm start reuses one and runs only the handler.
- By default each environment serves one request at a time, so concurrency grows with traffic, up to your account's limit.
- It suits short, bursty work; long jobs and heavy, steady traffic are often better on a server or container.