Collecting Windows System and Hardware Metrics with Prometheus
Collecting Windows System and Hardware Metrics with Prometheus
windows_exporter supplies Windows CPU, memory, and disk metrics, while temperature, GPU power, and VRAM need a verified sensor source. On a Ryzen 5600H / RTX 3050 Ti laptop, I used two exporters and one Prometheus server.
OS metrics are based on windows_exporter 0.30.7; hardware comes from HardwareExporterWindows / LibreHardwareMonitor. The August deployment record was rechecked against endpoints on September 15. The configurations below are examples, not the complete live dashboard, and other hardware may expose different sensors.
Two sources for one machine
Windows
├─ windows_exporter :9182 → CPU/memory/disk/network/services
└─ HardwareExporter :9888 → temperature/GPU/power/VRAM
↓
Prometheus → Grafana
└─ alert rules → AlertmanagerTwo processes are the choice here, not a universal requirement. WMI thermalzone must not be assumed to represent CPU core temperature. Verify sensor names, units, and behavior.
Read real metrics after installation
An elevated cmd.exe MSI example, using an official downloaded installer and a documentation address for the monitoring server:
msiexec /i windows_exporter-0.30.7-amd64.msi /qn /norestart ENABLED_COLLECTORS=cpu,memory,logical_disk,net,os,service,system LISTEN_PORT=9182 REMOTE_ADDR=192.0.2.30Inspect all existing firewall allows: adding a narrow rule does not override a broad old allow. Both endpoints should be restricted to the monitoring path.
Get-Service windows_exporter,HardwareExporter,PawnIO
Invoke-WebRequest http://localhost:9182/metrics -UseBasicParsing
Invoke-WebRequest http://localhost:9888/metrics -UseBasicParsingHardware runtime and driver requirements depend on the chosen release. This deployment used .NET 10 and PawnIO. A running driver does not guarantee every chipset is supported.HardwareExporterWindows
| OS metric, version 0.30.7 | Meaning |
|---|---|
windows_cpu_time_total{mode="idle"} | Accumulated idle CPU time |
windows_memory_available_bytes | Available memory |
windows_memory_physical_total_bytes | Total physical memory |
windows_logical_disk_free_bytes / size_bytes | Volume capacity |
windows_system_boot_time_timestamp | Boot Unix timestamp |
windows_service_state | 0/1 state selected by the state label |
Some deprecated aliases remain in 0.30.7. Deprecated does not mean absent; use installed-version output as the authority.Version README
Observed hardware metrics included hardware_gpu_temperature_gpu_core, hardware_gpu_memory_used_bytes, and hardware_cpu_power_package. Use °C, bytes, and watts appropriately. Preserve distinguishing sensor labels instead of summing unrelated readings.
Distinguish source with job and machine with instance
Replace documentation addresses before deployment. This example retains service states so alert data remains available:
global:
scrape_interval: 60s
scrape_configs:
- job_name: windows
static_configs:
- targets: ['192.0.2.20:9182']
labels:
instance: legion
- job_name: hardware-lhm
static_configs:
- targets: ['192.0.2.20:9888']
labels:
instance: legionA shared instance supports device-oriented panels; job distinguishes the endpoints, including separate up series. Validate the configuration and actual target scrapes, not only Windows service status.
Queries that preserve device identity
# Per-instance CPU utilization (%)
100 * (1 - avg by (instance) (
rate(windows_cpu_time_total{job="windows",mode="idle"}[5m])
))
# Memory utilization (%)
100 * (1 - windows_memory_available_bytes{job="windows"}
/ windows_memory_physical_total_bytes{job="windows"})
# Per-volume usage (%)
100 * (1 - windows_logical_disk_free_bytes{job="windows"}
/ windows_logical_disk_size_bytes{job="windows"})
# One measured GPU sensor family
hardware_gpu_temperature_gpu_core{job="hardware-lhm",instance="legion"}Selectors belong on metric selectors, not after expressions such as (a+b)/c. Preserve instance grouping so one busy machine is not hidden by averaging all hosts. Disk free/size series need matching volume labels; inspect raw series when panels show no data.
Alert on zero and missing services separately
A stopped service generally has running-state value zero. An uninstalled service, disabled collector, or dropped series may instead be absent. A comparison to zero does not catch absence.Service collector documentation
This rule targets one explicitly required example service; label case must match actual output. Load it through Prometheus rule_files:
groups:
- name: windows-example
rules:
- alert: WindowsExporterScrapeFailed
expr: up{job="windows"} == 0
for: 2m
labels:
severity: warning
- alert: RequiredWindowsServiceUnavailable
expr: |
windows_service_state{job="windows",instance="legion",name="mssqlserver",state="running"} == 0
or absent(windows_service_state{job="windows",instance="legion",name="mssqlserver",state="running"})
for: 2m
labels:
severity: warningThe SQL service is illustrative, not installed on this laptop. up=0 means a failed scrape, not necessarily a powered-off host. Removed discovery targets, stopped Prometheus, and failed remote writing need separate checks.
promtool check config prometheus.yml
promtool check rules windows-rules.ymlThese are deployment checks, not claims of executing against production during article preparation.Rule validation documentation
Investigate No data along the pipeline
Check scrape health, collector/metric availability, labels, and finally dashboard variables. up=1 does not prove every sensor exists; up=0 does not identify shutdown versus network failure, exporter crash, or timeout.
| Layer | Evidence | Common mistake |
|---|---|---|
| Windows | Raw metrics, services, devices | Running service implies every metric |
| Scrape network | Target errors/timeouts/reachability | Localhost success implies remote access |
| Prometheus | Raw names and job/instance/sensor labels | Applying rate to a nonexistent metric |
| Grafana | Resolved variables, query preview, units | Replacing missing temperatures with zero |
| Notification | Pending/firing, receiver, inhibition | Rule presence implies delivered notification |
Archive real metrics as a version contract. Sensor renames and upgrades can break old queries. Missing temperatures are not 0°C, and a zero reading does not establish a powered-off device.
Turn example rules into accepted operation
A per-instance absent() checks one expected series; it does not generate missing-host alerts for an entire fleet. Maintain an expected inventory or rule-generation mechanism. Removed discovery targets do not keep producing up=0.
Stopping a service, uninstalling it, and stopping the exporter test state, existence, and scrape failure separately. In a test environment, wait for scrape and for windows, inspect labels, recovery notifications, and deduplication, then expand deployment.
Associate thresholds with duration and workload. Disk percentage needs capacity/growth context. Preserve interface and host dimensions for network counters; service state is a gauge, not a rate input. Back up configuration, rules, dashboard source, version samples, and receivers together rather than relying on screenshots.
Panels, cardinality, and acceptance
Group Grafana panels by machine. CPU/GPU temperatures can share a chart; bytes and percentages need separate charts or explicit axes. Verify semantics before setting thresholds and duration. Group Alertmanager notifications by instance/alertname and limit repeats.
Process and service collectors increase cardinality. Enable only useful collections. Remote writing does not eliminate scrape-side costs. Maintain dashboard source; a sanitized public-dashboard API response is not a complete backup to write back.
Acceptance includes readable endpoints, UP scrapes, useful queries, correct units, stopped and missing-service alert tests, and restricted firewall access. Installing exporters is only the beginning.
The article date is the main KB’s first Git commit date, 2026-08-25 (UTC+8), commit c2ae3ba. Experiment dates are stated separately. No production systems were accessed or changed while preparing this article.
