GitOps is a way of running servers where a Git repository holds the full description of what should be running, and automation keeps the real systems matching it. Every change becomes a reviewed commit, so you can always see what is deployed and who changed it.
The problem: changes nobody can trace
Picture a team running a handful of services. Most changes go through a deploy script, but now and then someone logs into a server and edits a config file by hand, to raise a memory limit or flip a feature flag during an incident. It works, and nobody writes it down.
A few weeks later, things start to go wrong:
- No record: nobody can say what changed, when, or why.
- No rollback: there is no earlier version to go back to, only someone's memory of what the file used to say.
- Drift: two servers that should be identical behave differently, because one was edited and the other wasn't.
- Snowflake servers: if a server dies, rebuilding it exactly is guesswork, because its real configuration only ever existed on that machine.
Each manual edit is small. Together they leave you with servers whose real state nobody can describe.
What GitOps changes
GitOps takes the tools developers already trust for code, Git history and pull request review, and applies them to everything that describes a running system.
Git is the source of truth
One repository, or a small set of them, describes the whole system: the app's configuration, the infrastructure it runs on and the deployment settings, such as which image version to run and how many copies. If it isn't in Git, it shouldn't exist on the servers.
The files describe what you want, not how to get there
GitOps relies on declarative configuration. Rather than a script of steps ('stop the old container, pull the new image, start it again'), the files state the end result ('run three copies of version 1.5'). Kubernetes manifests, Helm charts and Terraform files all work this way. Declarative files are what make the next part possible: a program can compare 'what the file says' with 'what is running' and work out the difference itself.
Changes go through pull requests
Nobody edits a server to make a change. They edit a file in the repository and open a pull request. Git records the change, a teammate reviews it, and once it is approved and merged, it becomes the new plan.
An agent keeps the servers in line
A GitOps agent, such as Argo CD or Flux, runs next to the servers (usually inside a Kubernetes cluster). It watches the repository and keeps comparing it with what is running. When Git changes, it notices and updates the live system to match, exactly as written. This loop of compare and fix is called reconciliation.
Here is one change from start to finish, and what happens when someone edits the cluster by hand afterwards:
A GitOps change, then a manual edit undone
Step 1 of 8: Git says version 1.4 and the cluster runs 1.4. The developer opens a pull request to move to 1.5.
A change made straight on the cluster doesn't last, because the agent treats anything that differs from Git as a mistake to fix. To keep a change, put it in Git.
Pull, not push
Plenty of teams deployed from Git before GitOps had a name: a CI pipeline runs kubectl apply or an SSH script after every merge. That is a push model. GitOps usually means a pull model.
In the push model, the pipeline needs credentials that can change production, so anyone who can tamper with the pipeline can change production too. It also only acts when a pipeline runs. If someone edits the server afterwards, nothing notices.
In the pull model, the agent lives inside the environment and fetches the desired state itself. Production credentials never leave the cluster, and because the agent keeps checking, it catches drift whenever it happens, not only at deploy time.
A worked example: shipping version 1.5
Say a Kubernetes app's deployment lives in a repository called cat-shop-config. Part of it looks like this:
apiVersion: apps/v1
kind: Deployment
metadata:
name: cat-shop
spec:
replicas: 3
template:
spec:
containers:
- name: web
image: ghcr.io/kitty/cat-shop:1.4.2To release version 1.5.0, nobody touches the cluster. The developer changes one line on a branch and opens a pull request:
git switch -c release-1.5.0
# edit the image tag to cat-shop:1.5.0
git commit -am "Release cat-shop 1.5.0"
git push -u origin release-1.5.0A teammate reviews the one-line diff and merges it. Within a few minutes, or straight away if the repository sends the agent a webhook, the agent sees the new commit and rolls the deployment out to 1.5.0.
Now suppose 1.5.0 has a bug. Rolling back is another commit, reviewed the same way:
git switch -c rollback-1.5.0
git revert <release-commit>
git push -u origin rollback-1.5.0Once that merges, the agent sees Git say 1.4.2 again and rolls the deployment back. The release and the rollback are both in the history, with names and timestamps attached. That history is what makes GitOps auditable: 'what was running last Tuesday at 3pm?' is a git log away.
Where CI still fits
GitOps doesn't replace a CI/CD pipeline; it takes over the last step. A common split:
- CI builds the code, runs the tests and pushes a container image to a registry.
- CI, or a bot, opens a pull request that bumps the image tag in the config repository.
- The GitOps agent deploys whatever the config repository says.
Many teams keep application code and deployment config in separate repositories, so a config change doesn't trigger a full build, and access to production config can be locked down more tightly than access to the code. GitOps is one practice within the wider DevOps approach, where the people who build software also share responsibility for running it.
Common mistakes
- Don't commit secrets in plain text. Git history is permanent, and repositories get cloned widely. Store secrets encrypted, with tools such as Sealed Secrets or SOPS, or keep only a reference in Git and let something like the External Secrets Operator fetch the value from a secrets manager.
- Don't fix things by hand during an incident and leave it there. The fix either disappears at the next reconciliation or, if self-healing is off, quietly becomes drift. If you must act fast, pause syncing for that app, then put the fix in Git straight afterwards.
- Don't stop at a pipeline that runs
kubectl apply. It is a good start, but without an agent comparing Git with the live state, nothing catches drift. - Plan for one-off steps. Tasks such as database migrations don't map neatly onto 'make it look like this file'. Decide how they run, usually as a job the agent triggers, rather than doing them by hand.
When to use it
GitOps fits best when your platform is declarative and an agent can reconcile it: Kubernetes is the usual home, and Argo CD and Flux are built for it. It pays off when several people or teams deploy to shared environments and you need a clear record of who changed what.
It is less worth the setup for a single small server with one person deploying, where a simple deploy script is easier to understand. It also adds moving parts: an agent to run and upgrade, and a slower path for emergency changes.
Key takeaways
- In GitOps, a Git repository is the single source of truth for app config, infrastructure and deployment settings.
- Changes go through pull requests, so every change is recorded, reviewed and easy to revert.
- An agent such as Argo CD or Flux watches Git and reconciles the live system to match it, pulling changes rather than having them pushed in.
- Manual edits on the server are drift: the agent undoes them, so lasting changes belong in Git.
- Keep secrets out of plain-text Git, and keep CI for building and testing.