Linux Fan Control and PVE Quiet-Tuning: From hwmon to EC Registers
Linux Fan Control and PVE Quiet-Tuning: From hwmon to EC Registers
There is no universal Linux fan-control command. Desktop boards usually expose PWM through hwmon, laptops use vendor interfaces, servers have BMC/IPMI policies, and some mini PCs leave all control inside the embedded controller (EC).
This article first chooses the safest control path, then shows how a noisy single-temperature PVE script became a multi-sensor controller with hysteresis, read-back verification, and automatic firmware fallback.
Identify the control path
for h in /sys/class/hwmon/hwmon*; do
printf '%s: ' "$h"
cat "$h/name"
find "$h" -maxdepth 1 \( -name 'temp*_input' -o -name 'pwm*' -o -name 'fan*_input' \) -print
done| File | Meaning |
|---|---|
temp1_input | Millidegrees Celsius |
fan1_input | Measured RPM |
pwm1 | PWM value, commonly 0–255 |
pwm1_enable | Commonly 1 for manual and 2 for automatic |
Prefer the highest-level supported option:
- Writable
pwm*:lm-sensorsandfancontrol. - Supported laptop:
thinkfan,nbfc-linux, or another model-specific tool. - Server BMC: retain the vendor automatic policy unless there is a validated reason not to.
- Temperatures but no PWM: consider EC reverse engineering only as a last resort.
Standard PWM with fancontrol
sudo apt install lm-sensors fancontrol
sudo sensors-detect
sensors
sudo pwmconfig
sudo systemctl enable --now fancontrolReview the generated configuration:
INTERVAL=10
FCTEMPS=hwmon0/pwm1=hwmon0/temp1_input
FCFANS=hwmon0/pwm1=hwmon0/fan1_input
MINTEMP=hwmon0/pwm1=40
MAXTEMP=hwmon0/pwm1=75
MINSTART=hwmon0/pwm1=50
MINSTOP=hwmon0/pwm1=30
MINPWM=hwmon0/pwm1=30
MAXPWM=hwmon0/pwm1=255PWM duty is not RPM percentage. A motor has a start threshold and a lower sustain threshold, which is why MINSTART is often higher than MINSTOP. Fan stop may also be inappropriate when NVMe, memory, and NICs keep generating heat.
Why the first PVE script was noisy
The original script mapped instantaneous Ryzen Tctl values directly to fan levels:
12% → 80% → 60% → 20% → 12% → 35% → 60% → 12%Short boost spikes were amplified, downshifts had no independent exit threshold, “stable” samples were not necessarily consecutive, hwmon numbers were hard-coded, sensor failure was treated as a low temperature, and process exit left the EC in its last manual state. It also ignored GPU, memory, NVMe, and NIC temperatures.
Multi-sensor design
Give each component its own curve and select the highest requested level:
final level = max(CPU curve, GPU guard, RAM guard, NVMe guard, NIC guard)Example CPU curve:
| Filtered temperature | Fan duty |
|---|---|
| <68°C | 12% |
| 68–74°C | 18% |
| 74–79°C | 25% |
| 79–84°C | 35% |
| 84–88°C | 50% |
| 88–92°C | 70% |
| ≥92°C | 100% |
Smooth short spikes with an exponential moving average:
filtered = (filtered × 7 + raw) / 8Require several consecutive samples to move up, a much longer stable period to move down, and only drop one level at a time. Keep a raw-temperature emergency path: a 92°C reading should trigger full speed immediately rather than wait for the filter.
EC control is the last resort
This mini PC exposed no standard PWM node. The fan mode and duty registers were confirmed from its own ACPI DSDT. Register maps vary by model; copying offsets from another machine can damage hardware.
sudo apt install acpica-tools
sudo sh -c 'cat /sys/firmware/acpi/tables/DSDT > /tmp/dsdt.dat'
iasl -d /tmp/dsdt.dat
rg -n 'FAN|FCMI|WMAA|WMAB' /tmp/dsdt.dslA production controller should write only confirmed registers, read back every value, restore firmware automatic mode after repeated sensor failure, trap EXIT/INT/TERM, add a second systemd ExecStopPost fallback, and force maximum cooling on emergency temperature.
[Unit]
Description=Safe EC fan controller
After=multi-user.target
[Service]
ExecStart=/usr/local/sbin/fan-daemon
ExecStopPost=/usr/local/sbin/fan-control auto
Restart=on-failure
RestartSec=3
[Install]
WantedBy=multi-user.targetBefore enabling boot startup, stop the service manually and verify that the EC really returns to firmware-controlled automatic mode.
Validate with the real workload
The final controller was tested with concurrent Vulkan speech-recognition requests, which heated the integrated GPU and shared memory as well as the CPU. The measured peaks were roughly 66.5°C CPU, 65°C GPU, and 56°C memory.
The fan moved only from 12% to 18% and 25%, never oscillated, and stepped down conservatively after the load ended. Memory—not CPU—became the sustained airflow owner, showing why a CPU-only test would have missed the important heat source.
Log structured telemetry and transitions:
telemetry cpu=55.4/54.8C gpu=50.0/49.2C ram=52.9/52.2C \
nvme=38.3C nic=57.1C duty=25% owner=ramjournalctl -u fan-daemon.service --since today -o cat | grep change
journalctl -u fan-daemon.service --since today -o cat | grep -E 'WARNING|ERROR'Summary
- Use standard hwmon and vendor-supported control before touching EC registers.
- A safe controller needs smoothing, per-level hysteresis, multiple sensors, and a raw-temperature emergency path.
- Any read or write failure should return control to firmware, not preserve the last low manual level.
- Re-run the same real workload after every curve change and retain telemetry.
Quiet tuning is not the lowest possible PWM value. It is the removal of thermally meaningless oscillation while preserving a measurable, tested, and reversible safety boundary.
