Prometheus is an open-source monitoring and alerting toolkit built for reliability and scalability.
It collects metrics from systems and services, stores them in a time-series database, and allows querying via the PromQL language.
Grafana, on the other hand, is an open-source visualization platform that connects to Prometheus (and many other data sources) to build dashboards and alerts.
Together, they form the de facto standard for observability in DevOps and Kubernetes environments.
Answer:
Prometheus is an open-source monitoring system that collects metrics from servers, containers, and applications.
It helps DevOps engineers detect issues proactively and analyze performance trends.
Why it’s used:
Example use case:
Monitor CPU and memory usage across EC2 instances and trigger alerts when thresholds are breached.
Answer:
Grafana is a data visualization and analytics platform that turns raw metrics from Prometheus into interactive dashboards.
In DevOps workflows:
Example:
Grafana can display Kubernetes cluster health (CPU, memory, pod restarts) using data pulled from Prometheus.
Answer:
Metrics are numerical data points that describe system performance over time.
Each metric has:
node_cpu_seconds_total)instance="web1")Example:
http_requests_total{method="GET", handler="/"} 25493
Metrics are usually exposed by an /metrics endpoint on monitored services.
Answer:
Prometheus pulls metrics from targets by scraping HTTP endpoints.
Each monitored target exposes data at /metrics.
Example:
Prometheus fetches data like this:
GET http://node-exporter:9100/metrics
Why pull-based?
It simplifies configuration, works well with dynamic environments (e.g., Kubernetes), and doesn’t rely on agents pushing data.
Answer:
Exporters expose metrics from third-party systems in a format Prometheus can understand.
Examples:
In practice:
You deploy exporters alongside services to monitor their performance.
Answer:
PromQL (Prometheus Query Language) is used to query and aggregate metrics stored in Prometheus.
Example:
rate(http_requests_total[5m])
This query returns the per-second rate of HTTP requests over the last 5 minutes.
Why important:
PromQL powers alerts and dashboards — understanding it is critical for real-world monitoring.
Answer:
A target is any system Prometheus scrapes for metrics.
Targets can be:
prometheus.yml)Example:
scrape_configs:
- job_name: 'webapp'
static_configs:
- targets: ['192.168.1.10:9100']
Answer:
Data flow:
Targets → Prometheus Server → Alertmanager/Grafana → Notifications.
Answer:
Alertmanager receives alerts from Prometheus and routes them to:
Example configuration:
receivers:
- name: 'slack-alerts'
slack_configs:
- channel: '#alerts'
send_resolved: true
Answer:
Example Query:
avg(rate(container_cpu_usage_seconds_total[5m])) by (pod)
Answer:
Add multiple scrape jobs in prometheus.yml:
scrape_configs:
- job_name: 'node'
static_configs:
- targets: ['server1:9100', 'server2:9100']
- job_name: 'nginx'
static_configs:
- targets: ['web1:9113']
This lets you monitor infrastructure, applications, and databases concurrently.
Answer:
| Component | Purpose |
|---|---|
| Exporter | Exposes metrics for long-lived services (e.g., node-exporter). |
| Pushgateway | Allows short-lived jobs (e.g., batch jobs) to push metrics before they exit. |
Example:
A nightly data processing job can push its metrics to the Pushgateway since it terminates quickly.
Answer:
High cardinality (too many unique label combinations) can bloat storage.
To mitigate:
sum or avg) instead of detailed metrics.Answer:
Deploy Prometheus Operator or kube-prometheus stack in the cluster.
It monitors:
kubelet and cAdvisor.Helm installation:
helm install prometheus prometheus-community/kube-prometheus-stack
Answer:
A Grafana Dashboard displays visualizations like graphs, heatmaps, and gauges.
Steps:
Answer:
Alert setup in YAML:
alert:
name: High CPU Usage
expr: avg(node_cpu_seconds_total) > 0.85
Answer:
Answer:
Grafana supports:
This allows building end-to-end observability dashboards that combine metrics, logs, and traces.
Answer:
By default, Prometheus stores data locally (~15 days).
To extend retention:
--storage.tsdb.retention.time=90d).These tools enable long-term and scalable storage.
Answer:
Prometheus exposes internal metrics at /metrics too.
You can create alerts on:
prometheus_tsdb_head_samples_appended_totalprometheus_notifications_queue_capacitySelf-monitoring ensures Prometheus is scraping targets efficiently and not running out of memory.
Answer:
Prometheus is designed for simplicity but can scale using:
| Type | Description | Example Tool |
|---|---|---|
| Metrics | Quantitative data over time | Prometheus |
| Logs | Event details or errors | Loki, ELK |
| Traces | Distributed request flow | Jaeger, Tempo |
A strong DevOps observability stack integrates all three.
Answer:
All tools together = full observability pipeline.
Answer:
Example:
Instead of “CPU > 80%”, use “CPU > 80% for 10 minutes”.
Answer:
Example metric:
jenkins_builds_success_total{job="DeployApp"} 45
Answer:
rate() over long intervals).Answer:
Use Prometheus CloudWatch Exporter or remote_write config:
remote_write:
- url: https://cloudwatch-agent:port/receive
This allows integrating existing AWS dashboards and alarms.
Answer:
Use Grafana organizations or folders per team/environment.
Assign RBAC roles (Viewer, Editor, Admin).
Combine with multi-data-source Prometheus setups for isolation.
Q: You were tasked with diagnosing high CPU usage across production servers. How did Prometheus and Grafana help?
Answer (STAR):
rate(node_cpu_seconds_total[5m]) by instance.Q: Describe a time you implemented observability using Prometheus and Grafana for a Kubernetes cluster.
Answer (STAR):
kube-prometheus-stack via Helm.