Server downtime at 3 AM. No alerts, no centralized logs, no historical metrics. The support team faced a complete unknown. The post-mortem diagnosis revealed the server’s disk had been full for six hours before the incident—a problem easily predictable and, more importantly, preventable. This scenario, common in many IT environments, highlights the critical need for a robust and proactive monitoring system. It’s not just about collecting data; it’s about transforming it into actionable insights that anticipate issues, optimize resources, and ensure operational continuity.
In this guide, I’ll explore how to build a complete and high-performing monitoring stack using Prometheus and Grafana on a Linux system. Through clear, practical steps, I will show you how to install and configure these tools to monitor vital server metrics, create intuitive dashboards, and set up alerting systems that notify you before a minor issue escalates into a disaster. The goal is to equip you with the tools to reduce downtime, improve performance, and transition from reactive to proactive IT infrastructure management, all within approximately 60 minutes.
Prerequisites / Test Environment
To follow this guide, you will need a Linux server (preferably Ubuntu Server 22.04 LTS or higher, or an RHEL-based distribution) with sudo access. Ensure the server has internet access to download necessary packages. The test environment I will use is an Ubuntu 22.04 VM with 2 vCPUs and 4GB of RAM, sufficient for a basic monitoring setup. All operations will be performed from the command line.
Stack Architecture: Prometheus, Exporters, Grafana
The core of our monitoring system consists of three main components working in synergy:
- Prometheus: This is the metric collection engine. It operates on a pull model, periodically querying HTTP endpoints exposed by the services we want to monitor. Prometheus stores metrics in a time-series database, ideal for temporal analysis.
- Exporters: These are small agents installed on the targets to be monitored (servers, databases, applications, etc.). They translate specific metrics from that target into a format Prometheus understands. For example, the
Node Exportercollects hardware metrics from a Linux server. - Grafana: This is the visualization platform that interfaces with Prometheus (and other data sources) to create interactive dashboards, graphs, and reports. Grafana makes complex data immediately understandable, allowing for rapid analysis of performance and anomalies.
This modular architecture provides great flexibility and scalability, adapting to infrastructures of any size, from a single server to a complex data center. In an enterprise environment with 2,000 workstations and 300+ VMs, this scalability is crucial for maintaining control over the entire infrastructure.
Installing Prometheus on Linux (systemd service)
Let’s start by downloading and configuring Prometheus. We will create a dedicated user and a systemd service for optimal management.
- Create user and directories:
sudo useradd --no-create-home --shell /bin/false prometheus
sudo mkdir /etc/prometheus
sudo mkdir /var/lib/prometheus
sudo chown prometheus:prometheus /var/lib/prometheus
- Download and extract Prometheus:
Download the latest stable version of Prometheus (always check the GitHub page for the most recent version; here, I use v2.51.0).
wget https://github.com/prometheus/prometheus/releases/download/v2.51.0/prometheus-2.51.0.linux-amd64.tar.gz
tar -xvf prometheus-2.51.0.linux-amd64.tar.gz
sudo mv prometheus-2.51.0.linux-amd64 /usr/local/bin/prometheus-dist
sudo cp /usr/local/bin/prometheus-dist/prometheus /usr/local/bin/
sudo cp /usr/local/bin/prometheus-dist/promtool /usr/local/bin/
sudo chown prometheus:prometheus /usr/local/bin/prometheus
sudo chown prometheus:prometheus /usr/local/bin/promtool
- Configure Prometheus (
prometheus.yml):
Create a basic configuration file. This file will tell Prometheus which targets to monitor.
sudo nano /etc/prometheus/prometheus.yml
Add the following content:
global:
scrape_interval: 15s
scrape_configs:
- job_name: 'prometheus'
static_configs:
- targets: ['localhost:9090']
Assign correct permissions:
sudo chown prometheus:prometheus /etc/prometheus/prometheus.yml
- Create the systemd service:
sudo nano /etc/systemd/system/prometheus.service
Add the following content:
[Unit]
Description=Prometheus
Wants=network-online.target
After=network-online.target
[Service]
User=prometheus
Group=prometheus
Type=simple
ExecStart=/usr/local/bin/prometheus \n --config.file /etc/prometheus/prometheus.yml \n --storage.tsdb.path /var/lib/prometheus \n --web.external-url=http://localhost:9090 \n --web.listen-address=:9090
Restart=on-failure
[Install]
WantedBy=multi-user.target
- Start and enable the service:
sudo systemctl daemon-reload
sudo systemctl start prometheus
sudo systemctl enable prometheus
sudo systemctl status prometheus
Verify Prometheus is running by accessing http:// from your browser.
Node Exporter: CPU, Memory, Disk, Network Metrics
Node Exporter is essential for collecting Linux operating system metrics. Follow these steps to install and configure it.
- Create user and directory:
sudo useradd --no-create-home --shell /bin/false node_exporter
- Download and extract Node Exporter:
wget https://github.com/prometheus/node_exporter/releases/download/v1.7.0/node_exporter-1.7.0.linux-amd64.tar.gz
tar -xvf node_exporter-1.7.0.linux-amd64.tar.gz
sudo mv node_exporter-1.7.0.linux-amd64 /usr/local/bin/node_exporter-dist
sudo cp /usr/local/bin/node_exporter-dist/node_exporter /usr/local/bin/
sudo chown node_exporter:node_exporter /usr/local/bin/node_exporter
- Create the systemd service for Node Exporter:
sudo nano /etc/systemd/system/node_exporter.service
Add the following content:
[Unit]
Description=Node Exporter
Wants=network-online.target
After=network-online.target
[Service]
User=node_exporter
Group=node_exporter
Type=simple
ExecStart=/usr/local/bin/node_exporter \n --web.listen-address=:9100
Restart=on-failure
[Install]
WantedBy=multi-user.target
- Start and enable the service:
sudo systemctl daemon-reload
sudo systemctl start node_exporter
sudo systemctl enable node_exporter
sudo systemctl status node_exporter
Verify Node Exporter is running by accessing http://.
Configure Scrape Targets in prometheus.yml
Now we need to tell Prometheus to start collecting metrics from Node Exporter. Modify the /etc/prometheus/prometheus.yml file to add a new job.
sudo nano /etc/prometheus/prometheus.yml
Update the file by adding the section for node_exporter:
global:
scrape_interval: 15s
scrape_configs:
- job_name: 'prometheus'
static_configs:
- targets: ['localhost:9090']
- job_name: 'node_exporter'
static_configs:
- targets: ['localhost:9100'] # If Node Exporter is on another server, use <NODE_EXPORTER_IP>:9100
After saving, restart Prometheus to apply the changes:
sudo systemctl restart prometheus
You can verify that Prometheus is collecting metrics from Node Exporter by visiting the Status -> Targets page (http://). You should see a node_exporter target with an UP status. Another quick check for active targets is:
curl localhost:9090/api/v1/targets | jq '.data.activeTargets[].health'
Install Grafana and Connect to Prometheus
Grafana is our interface for visualizing metrics. Installation is straightforward and well-documented.
- Install Grafana (on Ubuntu):
sudo apt-get install -y apt-transport-https software-properties-common wget
wget -q -O - https://packages.grafana.com/gpg.key | sudo gpg --dearmor -o /usr/share/keyrings/grafana.gpg
echo "deb [signed-by=/usr/share/keyrings/grafana.gpg] https://packages.grafana.com/oss/deb stable main" | sudo tee /etc/apt/sources.list.d/grafana.list
sudo apt-get update
sudo apt-get install grafana
- Start and enable the Grafana service:
sudo systemctl daemon-reload
sudo systemctl start grafana-server
sudo systemctl enable grafana-server
sudo systemctl status grafana-server
Access Grafana from your browser at http://. Default credentials are admin/admin. You will be prompted to change the password on first login.
- Add Prometheus as a Data Source:
Once logged into Grafana, navigate to Connections -> Data sources -> Add new data source. Select Prometheus.
- Name:
Prometheus(or a name of your choice) - URL:
http://localhost:9090(or the IP of your Prometheus server) - Leave other settings as default and click
Save & Test. You should see aData source is workingmessage.
Import Pre-built Dashboard (ID 1860 — Node Exporter Full)
One of Grafana’s strengths is its vast library of community-built dashboards. To monitor our Linux server, we will use the Node Exporter Full dashboard (ID 1860), which provides a comprehensive overview of system metrics.
- In Grafana, navigate to
Dashboards->New Dashboard->Import. - In the
Import via grafana.comfield, enter1860and clickLoad. - On the next screen, select the
Prometheusdata source you just configured and clickImport.
Now you will have a dashboard rich with information on your server’s CPU, memory, disk, network, and processes. You can explore metrics, such as CPU utilization, with PromQL queries like:
100 - (avg by(instance)(rate(node_cpu_seconds_total{mode='idle'}[5m])) * 100)
This query calculates the average CPU utilization (non-idle) over the last 5 minutes, a key indicator of system performance.
Create Alert Rules in Prometheus
Monitoring is incomplete without an alerting system. Prometheus allows you to define rules to detect anomalies and trigger notifications. We will create an alerting rules file.
- Create the
alert.rules.ymlfile:
sudo nano /etc/prometheus/alert.rules.yml
Add a simple rule for disk usage:
groups:
- name: node_exporter_alerts
rules:
- alert: HighDiskUsage
expr: node_filesystem_avail_bytes{mountpoint="/"} / node_filesystem_size_bytes{mountpoint="/"} * 100 < 20
for: 5m
labels:
severity: critical
annotations:
summary: "Disk usage on {{ $labels.instance }} is low ({{ $value | printf "%.2f" }}% available)"
description: "Filesystem {{ $labels.mountpoint }} on {{ $labels.instance }} has less than 20% space left. Please check."
- Update
prometheus.yml:
We need to tell Prometheus to load these rules. Modify prometheus.yml:
sudo nano /etc/prometheus/prometheus.yml
Add the rule_files section:
# ... (other configurations)
rule_files:
- "/etc/prometheus/alert.rules.yml"
- Restart Prometheus:
sudo systemctl restart prometheus
Now Prometheus will evaluate this rule and generate an alert if the conditions are met.
Alertmanager: Email and Telegram Notifications
Prometheus generates alerts, but Alertmanager handles routing these notifications to the correct channels (email, Telegram, Slack, PagerDuty, etc.), managing deduplication, grouping, and suppression of alerts.
- Download and extract Alertmanager:
wget https://github.com/prometheus/alertmanager/releases/download/v0.27.0/alertmanager-0.27.0.linux-amd64.tar.gz
tar -xvf alertmanager-0.27.0.linux-amd64.tar.gz
sudo mv alertmanager-0.27.0.linux-amd64 /usr/local/bin/alertmanager-dist
sudo cp /usr/local/bin/alertmanager-dist/alertmanager /usr/local/bin/
sudo cp /usr/local/bin/alertmanager-dist/amtool /usr/local/bin/
sudo chown prometheus:prometheus /usr/local/bin/alertmanager
sudo chown prometheus:prometheus /usr/local/bin/amtool
sudo mkdir /etc/alertmanager
- Configure Alertmanager (
alertmanager.yml):
sudo nano /etc/alertmanager/alertmanager.yml
Example basic configuration for email:
global:
resolve_timeout: 5m
route:
group_by: ['alertname']
group_wait: 30s
group_interval: 5m
repeat_interval: 1h
receiver: 'email-receiver'
receivers:
- name: 'email-receiver'
email_configs:
- to: 'your_email@example.com'
from: 'alertmanager@example.com'
smarthost: 'smtp.example.com:587'
auth_username: 'your_smtp_username'
auth_password: 'your_smtp_password'
Note: For Telegram or other services, the configuration differs and requires specific tokens. Consult the official Alertmanager documentation for details.
- Create the systemd service for Alertmanager:
sudo nano /etc/systemd/system/alertmanager.service
[Unit]
Description=Alertmanager
Wants=network-online.target
After=network-online.target
[Service]
User=prometheus
Group=prometheus
Type=simple
ExecStart=/usr/local/bin/alertmanager \n --config.file=/etc/alertmanager/alertmanager.yml \n --web.listen-address=:9093 \n --storage.path=/var/lib/alertmanager
Restart=on-failure
[Install]
WantedBy=multi-user.target
- Start and enable the service:
sudo systemctl daemon-reload
sudo systemctl start alertmanager
sudo systemctl enable alertmanager
sudo systemctl status alertmanager
- Connect Prometheus to Alertmanager:
Finally, we need to tell Prometheus to send alerts to Alertmanager. Modify /etc/prometheus/prometheus.yml to add the alerting section:
# ... (other configurations)
alerting:
alertmanagers:
- static_configs:
- targets: ['localhost:9093']
Restart Prometheus to apply the change:
sudo systemctl restart prometheus
Now your monitoring stack is complete: Prometheus collects, Grafana visualizes, and Alertmanager notifies. This configuration allows you to detect an ongoing attack before it spreads to 2,000 endpoints, or a capacity issue before it impacts 300+ VMs, ensuring a higher SLA.
Common Errors and Troubleshooting
- Firewall: Ensure that ports 9090 (Prometheus), 9100 (Node Exporter), and 3000 (Grafana) are open on the server’s firewall (
ufw allow 9090,ufw allow 9100,ufw allow 3000). - Incorrect Permissions: If services fail to start, check directory and configuration file permissions (
chown,chmod). - YAML Configuration Error: An indentation error in a YAML file can cause service startup to fail. Use a YAML linter or an editor with YAML support to verify syntax.
- Prometheus Not Seeing Targets: Check the
http://page and Prometheus logs (:9090/targets journalctl -u prometheus).
Conclusions with Operational Takeaways
Configuring a monitoring stack with Prometheus and Grafana is a fundamental step for any IT infrastructure. It not only provides real-time visibility into your system’s performance but also equips you with the tools to anticipate problems, optimize resources, and significantly reduce the risk of downtime. Remember that an effective monitoring system is an investment that pays off in terms of stability, efficiency, and operational peace of mind. Start with essential metrics, expand gradually, and never underestimate the value of a timely alert. By implementing this solution in your environment, you can transform raw data into a true protective shield for your infrastructure.
Read also: SSH Hardening Linux: Complete Guide 2026 (10 Critical Settings)