Grafana is an open-source dashboard tool that turns the numbers your systems produce into graphs you can read at a glance. Point it at the places your metrics and logs already live, and you get one screen that tells you whether your servers, apps and databases are healthy right now.
Why you need to see your systems
A running system is full of moving parts: web servers, background workers, databases, caches, queues. Each one can be 'working' in the sense that it hasn't crashed, while quietly getting worse. Memory creeps up after every deploy. One endpoint gets slower each day. A small share of requests start failing.
None of that shows up until someone looks, and by the time users complain, the problem has usually been building for a while. This is the job of monitoring: collecting numbers about your systems over time so you can see trends and catch trouble early. It is a core part of running software in a DevOps team, where the people who ship the code also keep it running.
The numbers themselves are easy to collect. The hard part is making sense of thousands of them. A column of CPU readings taken every 15 seconds tells you nothing until it is drawn as a line, and you can see the spike at 3 a.m.
What Grafana does
Grafana is the layer that draws those lines. It gives you:
- Dashboards: pages of graphs, tables and single big numbers, arranged on a grid.
- Queries: each graph is backed by a query that fetches the data it shows.
- Alerts: rules that watch a query and notify you when something crosses a line.
It is open source, maintained by Grafana Labs, and you can run it yourself or use the hosted version, Grafana Cloud.
It doesn't store your data
Grafana doesn't collect or keep your metrics. It keeps its own settings, users and dashboard definitions, but the data in every graph comes from somewhere else, fetched when the dashboard loads or refreshes.
Those somewhere-elses are called data sources. Common ones include:
- Prometheus, a time-series database that collects metrics such as CPU, memory and request counts.
- Elasticsearch or Loki, for searching and counting log lines.
- Relational databases such as PostgreSQL and MySQL, for business numbers like sign-ups or orders.
- Cloud monitoring services, such as AWS CloudWatch.
Here is where Grafana sits, between you and the tools that hold the data:
Because Grafana only reads, you can put metrics from Prometheus next to error logs from Elasticsearch and order numbers from PostgreSQL on the same dashboard. That is the main reason teams use it: one place to look, whatever stores the data underneath.
Panels, queries and dashboards
A dashboard is made of panels. Each panel runs one or more queries against a data source and draws the result as a visualisation: a time series (lines over time), a bar chart, a gauge, a table, or a single 'stat' number. A panel can show more than one line, and can even compare today's numbers with an hour ago.
Here are two panels from Grafana's public demo site: memory and CPU on one graph, and logins plotted against the same series shifted back by an hour:
At the top of every dashboard sits a time picker ('Last 6 hours', 'Last 7 days'). Change it and every panel re-runs its query for the new window, which makes it quick to zoom from 'this week looks odd' down to 'it started at 14:05'.
Dashboards can also have variables, shown as drop-downs at the top. Pick a server or environment once and every panel filters to it, so one dashboard can cover twenty servers instead of twenty dashboards covering one each.
A worked example: a web app dashboard
Say you run a web app, and Prometheus already collects two kinds of metrics: server metrics from node_exporter (Prometheus's agent for machine stats), and a counter your app exposes called http_requests_total, with a status label for the response code. A useful first dashboard has four panels, one for each question you'd ask during an incident.
The first shows requests per second, so you can tell whether traffic is normal. rate() turns an ever-growing counter into a per-second rate, averaged over the last five minutes:
sum(rate(http_requests_total[5m]))The second asks whether requests are failing. Divide the rate of 5xx responses by the rate of all responses:
sum(rate(http_requests_total{status=~"5.."}[5m]))
/
sum(rate(http_requests_total[5m]))Set the panel's unit to 'Percent (0.0-1.0)' and 0.02 shows as 2%.
The third asks whether the machine is busy. It takes the CPU time spent idle and turns it into a busy percentage per server:
100 - avg by (instance) (
rate(node_cpu_seconds_total{mode="idle"}[5m])
) * 100The fourth watches for leaks, with the share of memory that isn't available:
1 - node_memory_MemAvailable_bytes
/ node_memory_MemTotal_bytesThese queries are written in PromQL, Prometheus's query language, because that is the data source. Point a panel at PostgreSQL instead and you write SQL; Grafana passes each query to its data source and draws whatever comes back.
With those four panels side by side, a bad deploy is obvious: the error rate jumps while requests per second stay flat. Flat traffic rules out a surge, which points at the new code.
To try this without installing anything, Grafana's public demo site (in the useful links below) has dashboards like this already built. To run Grafana locally, the official Docker image starts it on port 3000:
docker run -d -p 3000:3000 grafana/grafanaAlerts: when nobody is watching
Dashboards only help when someone is looking at them. Alerts cover the rest of the time.
A Grafana alert rule is a query plus a condition, evaluated on a schedule: for example, 'every minute, check whether the error rate is above 5%'. Each rule can also have a pending period: how long the condition must stay true before Grafana acts on it. That stops a single bad scrape from waking someone up.
An alert moves between states like this:
When an alert fires, Grafana sends it to a contact point: an email address, a Slack channel, PagerDuty or a webhook. Larger teams use notification policies to route alerts by label, so database alerts reach the database team and the rest go to whoever is on call. When the condition clears, Grafana can send a 'resolved' message too.
Grafana and the tools around it
Grafana is usually one piece of a monitoring stack, and it helps to know the neighbours:
- Prometheus collects and stores metrics, and has its own basic graphing page. Teams pair it with Grafana for the dashboards.
- Loki stores logs, Tempo stores traces, and Mimir stores metrics at large scale. They are Grafana Labs projects designed to plug into Grafana.
- Kibana is the dashboard tool built for Elasticsearch. It goes deeper into log search; Grafana is broader, reading from many data sources at once.
A common small setup is Prometheus for metrics, Loki for logs and Grafana in front of both.
Common mistakes
- Building dashboards nobody reads: fifty panels on one page hide the three that matter. Start with the questions you'd ask during an outage (is traffic normal, are requests failing, is it slow, is a machine full) and build those first.
- Graphing everything and alerting on nothing: a dashboard can't page you at night. Every panel that would make you act should have an alert behind it.
- Alerting on everything: alerts that fire often and mean nothing teach people to ignore them. Alert on what users feel, such as errors and slowness, and give rules a pending period.
- Expecting Grafana to keep the data: if Prometheus only keeps 15 days of metrics, Grafana can only show 15 days. How long data is kept is set in the data source.
- Leaving the default login: a new install starts with the user admin and password admin. Change it straight away, and don't expose Grafana to the internet without proper sign-in in front of it.
Key takeaways
- Grafana turns metrics and logs into dashboards, so you can see your systems' health on one screen.
- It stores no metrics itself: it queries data sources such as Prometheus, Elasticsearch and SQL databases.
- Each panel is a query plus a visualisation, and the time picker and variables reuse one dashboard across time ranges and servers.
- Alert rules watch a query on a schedule and notify a contact point once a condition holds for the pending period.
- Start small: a few panels that answer the questions you'd ask in an incident beat a wall of graphs.