Running a homelab on repurposed laptop hardware is a great idea right up until the point where the fan starts screaming and the machine vanishes offline in the middle of the night. I’ve been chasing two hardware gremlins on my HP EliteBook Proxmox node for a while now, and I want to document exactly what I tried – because the honest answer right now is: I still don’t know.
The Setup #
The node is an older HP laptop repurposed as a headless Proxmox VE host – display and lid removed, sitting in a Thermaltake case for airflow. It’s currently running three k3s VMs (nodes 101, 102, 103) and has been reasonably stable, except for two nagging issues.
Issue 1: The fan is loud. Not “laptop fan under load” loud – louder than expected for the temperatures I’m seeing.
Issue 2: The machine has shut down abruptly at least once, with no clean poweroff in the journal – just a hard stop, mid-log-entry.
Both could be hardware. Both could be software. I wanted to know which.
The Fan Problem #
First Instinct: OS Fan Control #
The natural place to start on Linux is hp-wmi, the kernel module that handles HP fan and thermal management. It loads at boot and exits with -22 (EINVAL). That’s not a transient glitch – it’s a firmware/driver mismatch. The driver is trying to talk to something the BIOS isn’t exposing.
Digging into the ACPI tables revealed why.
The Root Cause (Probably) #
The SSDTs reference an HP embedded controller device named \_SB.PCI0.LPCB.H_EC. Every method that touches the fan – CPUP, PCAP, PECC, thermal zones TZ00/TZ01, the whole lot – routes through H_EC. But H_EC is declared External in every table and never actually defined. The OS has no path to the fan controller at all.
This is why hp-wmi fails. It’s also why the kernel logs show AE_NOT_FOUND aborts on _TZ.TZ00._TMP and _TZ.TZ01._TMP – the thermal zone methods can’t find H_EC either.
The EC that is reachable – \_SB.PCI0.LPCB.EC0 at the standard ports 0x62/0x66 – is a generic helper with raw byte access (FANG(n) / FANW(n,v)) and no fan curve. Dead end.
What I Tried and Ruled Out #
| # | What I Tried | Result |
|---|---|---|
| 1 | hp_wmi reload |
Same -22 (EINVAL). Confirmed, not a glitch. |
| 2 | /sys/firmware/acpi/platform_profile |
Does not exist. BIOS predates ACPI platform profiles. |
| 3 | hp_bioscfg BIOS attributes |
Only Sure_Start exposed. No fan/thermal toggle. |
| 4 | nbfc (NoteBook FanControl) | No config exists for this board. Checked all 313 configs in nbfc-linux/configs. No prior art. |
| 5 | EC register 0x2E/0x2F |
Used by every Broadwell-era HP ProBook/EliteBook nbfc config. 0x2F silently discards writes. Not the fan register here. |
| 6 | EC register 0x58 |
Writable, holds a value of 103. Proven NOT the fan register – see thermal test below. |
| 7 | Extended EC space (ERIB/ERBD), indices 0x0000–0x07FF |
No live registers, no plausible fan/RPM values. Effectively unimplemented. |
| 8 | HP EC shared-memory window at 0xFF000000 |
Readable via /dev/mem, but returns a repeating 1C D4 1E FA pattern – not the EC block on this platform. |
| 9 | DPTF Fan cooling devices (cooling_device3–7) |
Registered, but their AML calls the undefined H_EC.ECMD(0x1A). Non-functional. Writing cur_state=0 vs 1 moved CPU temp by 1 °C – noise. |
The EC Register Thermal Test #
Register 0x58 looked promising – writable and holding a non-trivial value. To test it, I ran a sustained single-core load and forced the register to each extreme, with a watchdog ready to restore max at 82 °C:
PHASE 1 0x58 = 103 (BIOS max) + load → plateau 64 °C
PHASE 2 0x58 = 0 (min/off) + load → plateau 64 °C
delta = 0 °CA register that genuinely commanded the fan could not produce a 0 °C delta under sustained load. 0x58 is an unrelated config/telemetry byte.
So… Is the Fan Actually Broken? #
Here’s the reframe that changed my thinking: the temperatures aren’t alarming. Under sustained single-core load the CPU plateaus at 63–64 °C. At the normal working load of three k3s VMs it sits at 54–57 °C. For a 15 W Broadwell-U part, those are reasonable numbers – not what you’d see if the fan were genuinely pinned at maximum.
The BIOS-side fan controller is almost certainly working. The OS simply has no window into it.
So the noise is more likely mechanical – a dried-out bearing, or airflow/chassis resonance from the lidless laptop sitting in a Thermaltake case – than a control problem.
The critical missing measurement is a fan tach reading, and there is no OS interface on this board that can provide one.
The Phantom Shutdown Problem #
What the Journal Says #
Re-examining journalctl -b -1 is unambiguous: this is a hard power loss, not a software shutdown.
- The log ends abruptly at
17:17:01on acron.hourlyentry. - No
systemd-shutdown, noReached target poweroff.target, noSyncing filesystems, no kernel panic – nothing. - Boot
-3ends with a clean textbook poweroff sequence, proving clean shutdowns are recorded when they occur. Their absence here is meaningful. - The machine stayed offline for ~29 minutes before the next boot – it didn’t auto-recover, ruling out a kernel panic with auto-reboot.
Proxmox Watchdog Ruled Out #
pve-guests.service (PVE 9’s watchdog) is enabled and active. But a watchdog trip panics and reboots with a log entry. There is no such record. The watchdog didn’t cause this.
Leading Suspect: Power Delivery #
The abrupt-loss signature is identical whether the cause is:
- The wall losing power
- The charger brick failing
- The laptop’s charging/battery circuit misbehaving when run lidless long-term
With three k3s VMs running, sustained power draw is meaningfully higher than when I first noticed this issue. That makes power-delivery failure more likely, not less.
The only way to distinguish “wall lost power” from “machine dropped power internally” is AC-side monitoring – which I don’t have yet.
Where This Stands #
Neither issue is resolved. I have a much clearer picture of what is happening and why the obvious fixes don’t work, but the actual answers require physical access or hardware I don’t have on hand.
Fan – Next Steps #
- Get a real tach reading. A thermal camera on the exhaust under load, or a clip-on RPM probe on the fan connector. This gates everything else.
- Check for a newer BIOS. The only real software fix candidate.
F.11from 2015-07 may be missing theH_ECdefinition; HP shipped BIOS updates for 2015 platforms as late as 2023. Board ID080C1selects the image. Do not attempt remotely – a failed flash on a headless machine with no IPMI is a permanent brick. - Mechanical inspection. Clean the heatsink and fan, check bearing spin freely by hand, confirm the intake isn’t obstructed. Given the temperature data, this is now the most likely actual cause.
Shutdown – Next Steps #
- Fit a UPS or a logged AC-side power meter. The only way to determine whether the wall or the machine is at fault.
- Try a different charger brick and outlet. Reseat the DC connector while at it.
- Capture
journalctl -b -1 -eimmediately after any recurrence. The end of the previous boot’s log is the most useful artifact.
Final Thought #
This is the part of homelab work that doesn’t make it into the tutorials: sometimes you dig deep, rule out everything you can reach, and the answer is still “I need a multimeter, physical access, and probably a spare part.” The investigation wasn’t wasted – I know the ACPI tables are broken, I know the BIOS-side controller is the only thing regulating the fan, and I know power delivery is the most likely culprit for the shutdowns.
I’ll update this post when I have hands on the machine. For now it stays on the unsolved list.