Terraform lets you describe your servers, databases and networks in text files, then builds them for you. Because the setup lives in code, you can review it, version it in Git, and rebuild the same thing as often as you like, instead of clicking through a cloud console and hoping you remember every setting.
The problem with clicking
Every cloud provider has a web console where you can create a server by filling in a form. That works for one server. It stops working when you need three servers, a database, a private network, firewall rules and a load balancer, and then the same again for a test environment.
Clicking has three problems:
- Nobody can see what changed. A setting tweaked by hand last Tuesday leaves no record of who changed it or why.
- It isn't repeatable. Rebuilding an environment means repeating dozens of steps exactly, from memory or from a wiki page that is already out of date.
- Environments drift apart. Dev and prod were meant to match, but small manual changes pile up until a bug appears in one and not the other.
Infrastructure as code (IaC) fixes this by treating infrastructure like application code: written in files, reviewed in pull requests, stored in Git and applied by a tool. Terraform, made by HashiCorp, is one of the most widely used IaC tools.
Describe the result
Terraform is declarative: you write down what should exist, and Terraform works out how to get there. Instead of 'create a server, then wait, then attach a disk', you write 'there should be three servers like this', and Terraform decides what to create, change or delete.
Configuration is written in HCL (HashiCorp Configuration Language) in files ending in .tf. Each block describes one thing. Here are some web servers, their number set by a variable, and a database on AWS:
provider "aws" {
region = "eu-west-2"
}
resource "aws_instance" "web" {
count = var.web_count
ami = var.ami_id
instance_type = "t3.micro"
tags = {
Name = "tuna-web-${count.index}"
}
}
resource "aws_db_instance" "tuna" {
engine = "postgres"
instance_class = "db.t3.micro"
allocated_storage = 20
username = "kitty"
password = var.db_password
skip_final_snapshot = true
}The pieces to know:
- A provider is a plugin that knows how to talk to one platform's API. The
awsprovider turns your blocks into AWS API calls. - A resource is one thing to manage, such as a server, a database or a DNS record.
aws_instanceis its type andwebis the name you use to refer to it elsewhere. - Variables (
var.web_count,var.db_password) keep values out of the code, so the same files work in different places and secrets stay out of Git. Each one is declared in avariableblock, usually in avariables.tffile.
When one resource refers to another, for example a server that uses a security group's id, Terraform knows the security group has to exist first. It builds a dependency graph from those references and creates things in the right order, in parallel where it can.
Plan, then apply
Day to day, Terraform comes down to a few commands:
terraform init # download the providers
terraform plan # show what would change
terraform apply # make the changes
terraform destroy # remove everything it managesterraform plan is the important one. It reads your files, checks what already exists, and shows the difference, without changing anything. Here is that comparison, from your files to a built cloud:
How terraform plan and apply decide what to change
Step 1 of 9: You change the files to ask for three servers and a database. One server already exists.
A plan marks every resource with a symbol: + to create, ~ to update in place, - to destroy, and -/+ to destroy and create a replacement. The last line sums it up:
Plan: 2 to add, 0 to change, 0 to destroy.terraform apply shows the plan again and waits for you to type yes. In a pipeline, save the plan with terraform plan -out=tfplan and run terraform apply tfplan, so what gets applied is exactly what was reviewed.
To change one thing, edit the file and apply again. Change instance_type to t3.small and the plan shows a ~ update for each web server, and nothing for the database. Terraform only acts on the difference between your files and reality.
State: how Terraform remembers
The state file (terraform.tfstate) maps each resource in your code to the real object in the cloud, such as aws_instance.web[0] to a particular instance id. Without it, Terraform couldn't tell a server it created from one somebody else made, and it would try to create everything again.
By default the state is a local file, which is fine while learning and a problem on a team. If two people each have their own copy, they will overwrite each other's changes. Teams keep state remotely instead, in a backend such as an S3 bucket, Azure Storage or HCP Terraform:
terraform {
backend "s3" {
bucket = "kitty-terraform-state"
key = "prod/terraform.tfstate"
region = "eu-west-2"
}
}Most remote backends also lock the state while someone runs apply, so two runs can't change the same infrastructure at once.
The state can hold sensitive values in plain text, including database passwords. Treat it like a secret: never commit it to Git, and restrict who can read the bucket.
One codebase, many environments
The same code can build dev and prod. Put the differences in variable files:
# dev.tfvars
web_count = 1
# prod.tfvars
web_count = 3Then pass the right one: terraform apply -var-file=prod.tfvars.
Each environment also needs its own state, in a separate backend key, a separate directory or a Terraform workspace. With one shared state, applying the dev variables would shrink prod from three servers to one.
As a codebase grows, repeated blocks move into modules: reusable packages of resources with their own inputs. One web_app module can build the servers, database and network for any environment, so dev and prod differ only in the values passed in.
Many clouds, one workflow
Terraform has providers for AWS, Azure, Google Cloud and many other platforms, including Cloudflare, GitHub and Kubernetes. One project can manage a server on AWS, its DNS record on Cloudflare and the repository that deploys it on GitHub.
What carries over between clouds is the language, the plan and apply workflow, and the state model. The resource definitions stay specific to each cloud: an aws_instance and a google_compute_instance take different arguments, so moving clouds still means rewriting those blocks.
If you're new to the platforms themselves, the Cloud video covers what a cloud provider gives you.
Common mistakes
- Don't change things by hand in the console. A manual change is drift, and the next plan offers to put it back. Make every change in code.
- Read the plan before you apply. Some changes can't happen in place, so Terraform destroys the resource and creates a new one. On a database, that
-/+can mean losing the data. Theprevent_destroylifecycle setting makes Terraform refuse to plan a destroy for a resource you mark with it. - Keep secrets out of
.tffiles. Feed passwords in through variables, from a secrets manager or the pipeline. - Split a large setup into several states. One state for everything makes every plan slow and every mistake big, so divide it by environment and by area, such as networking apart from applications.
When to use it
Terraform suits infrastructure that lasts and has to be rebuilt reliably: environments, networks, databases, DNS, permissions. It fits well into a CI/CD pipeline, where a pull request shows the plan and merging runs the apply.
It isn't the only option. CloudFormation and Bicep do the same job for AWS and Azure alone. Pulumi takes the same approach in general-purpose languages like TypeScript and Python. Ansible is mostly used to configure what runs inside servers rather than to create them. OpenTofu is an open-source fork of Terraform, started after HashiCorp changed Terraform's licence in 2023, and it uses the same language.
For a quick experiment you will throw away tomorrow, clicking in the console is fine. Once anyone else has to understand, repeat or fix the setup, write it down as code.
Key takeaways
- Terraform is infrastructure as code: you describe what should exist in
.tffiles, and Terraform builds it. terraform planshows exactly what will be created, changed or destroyed before anything happens;terraform applymakes those changes.- The state file is how Terraform knows what it manages. Keep it in a locked remote backend, never in Git.
- The same code, with different variables and separate state, builds dev and prod alike.
- Providers let one workflow manage many clouds, but each cloud's resources are still written for that cloud.