Skip to content

Grafana

On this page

Grafana as a universal visualization and analytics layer — not just for infrastructure. Full case studies for infra, project management, business intelligence, manufacturing, energy, and healthcare, plus the basics, cheatsheet, thumb rules, and best practices.

Basics

Grafana is an open-source platform for querying, visualizing, alerting on, and exploring data — wherever it lives. Its superpower is being data-source agnostic: one dashboard can blend Prometheus metrics, SQL tables, logs, traces, and CSV/JSON APIs.

Core concepts

Concept What it is
Data source A connection to a backend (Prometheus, MySQL/Postgres, Loki, InfluxDB, Elasticsearch, Google Sheets, JSON API, BigQuery, Snowflake…).
Dashboard A set of panels on a grid, saved as JSON.
Panel One visualization (time series, stat, gauge, bar, table, heatmap, geomap, state timeline…).
Query A data-source-specific request (PromQL, SQL, LogQL…) powering a panel.
Variable A dropdown/template that makes dashboards reusable ($datacenter, $patient, $line).
Transformations In-Grafana data shaping: join, group by, calculate fields, pivot — no code.
Alerting Unified alert rules that evaluate queries and route notifications.
Annotations Event markers overlaid on graphs (deploys, incidents, maintenance).

The Grafana "LGTM" stack

Grafana Labs pairs Grafana with Loki (logs), Tempo (traces), Mimir (long-term metrics), and Pyroscope (profiling). You don't need all of it — Grafana works fine on top of just Prometheus, or just a SQL database.

Why it goes "beyond infra"

Because Grafana speaks SQL and arbitrary APIs, any operational or business data in a database can become a live dashboard with alerting. That's what makes it usable for hospitals, factories, energy grids, finance, and project tracking — not only servers.

Cheatsheet

Run it

# Docker
docker run -d -p 3000:3000 --name grafana grafana/grafana-oss
# default login admin / admin (change immediately)

# As code: provision data sources & dashboards on startup
# /etc/grafana/provisioning/datasources/*.yaml
# /etc/grafana/provisioning/dashboards/*.yaml

Provision a data source (file-based)

apiVersion: 1
datasources:
  - name: Prometheus
    type: prometheus
    access: proxy
    url: http://prometheus:9090
    isDefault: true
  - name: AppDB
    type: postgres
    url: db:5432
    user: grafana_ro
    jsonData: { sslmode: "require", database: "appdb" }
    secureJsonData: { password: "$DB_PASS" }

Dashboard building blocks

Variables:   $env = label_values(up, env)        # Prometheus
             $region = SELECT DISTINCT region FROM sales   # SQL
Repeat:      repeat a panel/row per variable value
Thresholds:  color stat/gauge green→amber→red
Transforms:  Join by field · Group by · Add field from calculation · Filter
Links:       data links drill from a panel to another dashboard with context
Annotations: query a deploys/incidents table to overlay events

Alerting (unified)

Alert rule = query + condition (e.g. avg() > threshold) + "for" duration
Contact points: Slack, Email, PagerDuty, Webhook, Teams, Opsgenie...
Notification policies: route by labels (severity, team, sector)

Useful PromQL/SQL panel patterns

# Infra: node CPU %
100 - (avg by (instance)(rate(node_cpu_seconds_total{mode="idle"}[5m]))*100)
-- BI: revenue by month
SELECT date_trunc('month', created_at) AS time, sum(amount) AS revenue
FROM orders GROUP BY 1 ORDER BY 1;

Thumb Rules

Rules of thumb

  • Start from the question, not the panel. Decide what decision the viewer must make, then choose the visualization.
  • One dashboard, one audience. Exec, on-call, and factory-floor views have different needs — don't cram them together.
  • Top-left = most important. Eyes land there first; put the headline KPI/stat there.
  • Template with variables so one dashboard serves every host/line/region/department.
  • Use thresholds and color sparingly and consistently (green good, red bad) — color is information.
  • Push heavy aggregation into the data source (SQL GROUP BY, Prometheus recording rules); Grafana visualizes, it isn't a database.
  • Stat/gauge for "is it OK right now", time series for "how did we get here", table for detail.
  • Annotate events (deploys, shift changes, maintenance) so spikes have explanations.
  • Read-only DB users for SQL data sources. A dashboard should never be able to write.

Use Cases (Case Studies)

Each case study below shows a realistic data flow → dashboard design → alerting → value. Figures are illustrative of what you can build; validate specific metrics against your own systems.

1) Infrastructure & Reliability (the classic)

Goal: know the health of hosts, containers, and services at a glance, and get paged before users notice.

Data flow: node_exporter + cAdvisor + app metrics → Prometheus → Grafana; logs via Loki, traces via Tempo.

Dashboard design:

  • Top row of stat panels: cluster CPU %, memory %, disk %, error rate, p95 latency.
  • Time series for the four golden signals (latency, traffic, errors, saturation).
  • Table of top noisy pods/hosts; state timeline for service up/down.
  • Logs panel (Loki) linked from a failing service via data links; trace drill-down.
  • Variables: $cluster, $namespace, $instance.

Alerting: up == 0, error ratio > threshold for: 10m, disk predicted full (predict_linear). Routed to Slack/PagerDuty by severity.

Value: faster MTTR, capacity planning, SLO/error-budget tracking. This is the best-documented Grafana use case and the foundation everything else borrows from.

2) Project Management & Delivery

Goal: make agile delivery visible — throughput, cycle time, WIP, blockers, and SLA on support tickets — without living inside Jira/Zoho reports.

Data flow: project tool (Jira / Zoho Projects / GitHub Issues) → either a SQL mirror of the data, the Infinity plugin (REST/JSON/CSV), or the tool's API → Grafana. Git/CI metrics can come from Prometheus exporters.

Dashboard design:

  • Stat panels: open vs closed this sprint, WIP count, overdue items, average cycle time.
  • Bar chart: throughput (issues completed) per sprint/week.
  • Cumulative flow (stacked time series): backlog → in-progress → done, to spot bottlenecks.
  • Table: aging work items sorted by days-in-status, color-coded past SLA.
  • Cycle/lead time trend; burndown via a SQL query over status-change history.
  • Variables: $project, $team, $assignee, $sprint.

Alerting: ticket breaching SLA, WIP above the team's limit, zero throughput for N days, a release-blocker label appearing.

Value: real-time delivery health for standups and stakeholders, early warning on bottlenecks, and SLA accountability — refreshed automatically instead of hand-built slides.

3) Business Intelligence (BI)

Goal: executive and operational KPIs — revenue, conversion, churn, pipeline, marketing funnel — combining multiple sources in one live view.

Data flow: transactional DB (Postgres/MySQL) + warehouse (BigQuery/Snowflake/Redshift) + spreadsheets (Google Sheets plugin) + SaaS APIs (Infinity/JSON) → Grafana. Heavy modeling stays in the warehouse; Grafana queries the modeled tables.

Dashboard design:

  • KPI stat row: MRR/revenue, new customers, conversion %, churn %, each with a sparkline and % change vs previous period (Grafana's "value and area" + reduce options).
  • Time series: revenue and orders over time, with annotations for campaigns/launches.
  • Bar gauge / table: revenue by product, region, or channel (variable-driven).
  • Funnel (bar chart with transformations) for the marketing/sales funnel.
  • Geomap: sales or users by geography.
  • Variables: $region, $product, $channel, time range comparison.

Alerting: revenue/orders drop below a daily floor, conversion falls, a data pipeline goes stale (freshness check: max(updated_at) older than X).

Value: a single source of truth that refreshes continuously, with the same alerting engine the infra team uses — so "revenue dropped" can page someone, not just sit in a report.

4) Manufacturing & Industrial (IIoT / OEE)

Goal: monitor production lines in real time — OEE (Availability × Performance × Quality), throughput, downtime reasons, and machine health (predictive maintenance).

Data flow: PLCs/sensors → industrial protocol (OPC-UA/MQTT/Modbus) → a time-series DB (InfluxDB, TimescaleDB, or Prometheus via an IIoT gateway like Telegraf) → Grafana. SCADA/MES data can come via SQL.

Dashboard design:

  • OEE gauge per line (and a roll-up across the plant), with green/amber/red thresholds.
  • State timeline: machine running / idle / fault, per station, across the shift.
  • Bar chart (Pareto): downtime by reason code — the classic "what's hurting us most".
  • Time series: cycle time, units/hour, scrap/defect rate; heatmap of temperature/vibration.
  • Table: active alarms; annotations for shift changes and maintenance windows.
  • Variables: $plant, $line, $shift, $machine.

Alerting: OEE below target, defect rate spike, vibration/temperature trending toward a maintenance threshold (predict_linear-style), a line stopped for > N minutes.

Value: floor and management see the same live numbers; downtime Pareto drives continuous improvement; sensor-trend alerts enable condition-based maintenance instead of run-to-failure. Grafana is widely used as the visualization layer on IIoT/InfluxDB stacks.

5) Energy Sector

Goal: monitor generation, consumption, grid metrics, and renewables (solar/wind/battery) in real time; track demand vs supply and efficiency.

Data flow: smart meters, inverters, SCADA, BMS → MQTT/Modbus → InfluxDB/TimescaleDB → Grafana. Weather APIs (for solar/wind forecasting context) via the JSON/Infinity plugin.

Dashboard design:

  • Stat/gauge: current load (kW), today's generation (kWh), battery state-of-charge, grid frequency, power factor.
  • Time series: generation vs consumption (stacked), per source (solar/wind/grid/battery).
  • Geomap: site/asset map colored by output or fault state for distributed assets.
  • Heatmap: load by hour-of-day vs day to reveal demand patterns.
  • Table: assets under-performing vs expected; annotations for outages/maintenance.
  • Variables: $site, $feeder, $asset, $source.

Alerting: voltage/frequency out of band, inverter fault, battery SoC critically low, consumption exceeding contracted demand (peak-shaving trigger), generation far below forecast.

Value: real-time situational awareness for operators, efficiency and yield tracking, demand-response and peak-shaving triggers, and faster fault isolation across distributed assets.

6) Healthcare

Compliance first

Healthcare data is highly regulated (e.g. HIPAA/GDPR). Dashboards must use de-identified or access-controlled data, encryption in transit/at rest, audited access, least-privilege read-only DB users, and SSO/RBAC. Treat patient identifiers as high-cardinality and sensitive — usually keep them out of metrics entirely.

Goal: operational and clinical visibility — bed/OR utilization, ED wait times, device fleet uptime, lab turnaround — and, with proper governance, near-real-time patient vitals monitoring.

Data flow: EHR/HIS and operational DBs (SQL) + medical-device/IoT telemetry (HL7/FHIR ingested into a database, or MQTT for device streams) → time-series/SQL store → Grafana, behind SSO + RBAC.

Dashboard design:

  • Operational (BI-style): ED wait time, bed occupancy %, OR utilization, admissions vs discharges, lab/imaging turnaround — stat panels + trends + tables by department.
  • Biomedical engineering: uptime/fault state of infusion pumps, ventilators, imaging machines (a fleet/IoT dashboard, structurally like the infra case).
  • Clinical (governed): ward-level vitals overview (HR, SpO₂, BP) as state timelines and thresholds, with alerting to clinical staff — only within a compliant, validated setup.
  • Variables: $facility, $ward, $device_type, $department.

Alerting: ED wait exceeding target, bed capacity threshold, a critical device offline, a vital crossing a clinically defined threshold (with clinician-defined rules and escalation).

Value: smoother patient flow and resource utilization, proactive device maintenance, and faster clinical response — all on one platform, provided governance and validation are done properly. (Vital-sign alerting is safety-critical: it must be clinically validated, not a hobby dashboard.)

Common Issues

Panel shows “No data”

Wrong data source selected, the query returns nothing for the time range, or a variable is empty. Open the panel's Query inspector to see the raw request/response, and widen the time range.

“Too many data points” / slow dashboards

Querying raw high-resolution data over long ranges. Push aggregation into the source (SQL GROUP BY time, Prometheus recording rules), use $__interval, and limit series.

Wrong timezone / shifted timestamps

Mismatch between the database timezone, the time column, and Grafana's display timezone. Store/return UTC, use $__timeFilter() macros, and set the dashboard timezone explicitly.

Variable dropdown is empty or huge

The variable query is wrong, or it returns thousands of values (cardinality). Constrain it with a query filter and consider a regex on the variable.

SQL panel error / injection risk

Use Grafana's macros ($__timeFilter, $__timeGroup) and built-in variable quoting; never string-concatenate untrusted input. Use a read-only DB user.

Alerts noisy or never fire

Missing for: duration, threshold off, or No Data/Error states mishandled. Tune the condition, set the no-data behavior deliberately, and test the contact point.

Dashboards drift / lost after restart

They were edited in the UI but not saved as code. Provision dashboards/data sources from files or Git and treat them as version-controlled artifacts.

Best Practices

  • Design for the decision and the audience; keep the most important number top-left.
  • Templatize with variables and reuse one dashboard across many entities.
  • Aggregate in the data source, visualize in Grafana; use recording rules / SQL views.
  • Provision as code (data sources + dashboards in Git) for reproducibility and review.
  • Lock down access: SSO, RBAC/folders, read-only data-source users, TLS, audit logs.
  • Use unified alerting with clear severities, for: durations, and routed contact points.
  • Annotate events (deploys, shifts, maintenance, campaigns) for context on anomalies.
  • Consistent color/threshold semantics across all dashboards (green good → red bad).
  • For regulated data (healthcare/finance), bake in compliance — de-identification, encryption, access control, and validation — from day one.

Official Sources