The Faulted Disk: harbor Replacement Writeup
The sequel to panicking-led-to-losing-my-desktop — this time the monitoring actually caught the disk dying, and nothing was lost. What happened # is my replica pool — a 10.9T mirror (2x 12TB) that receives syncoid snapshots from. One side of the mirror, a Seagate Exos (serial,), went FAULTED with 14 read + 22 checksum errors. The IronWolf mirror side carried the pool —. ZFS redundancy did exactly its job. The difference from last time # Last failure: no monitoring, found out by accident months later, desktop died. This failure: SigNoz + node-exporter’s ZFS collector → → alert rule → Gotify → my phone. The gotify notification fired before I knew anything was wrong. Diagnosis # Before or replace — check SMART. already scrapes all disks into SigNoz, so I didn’t even need sudo: raw = 3024 and counting value/worst still 100 but raw errors climbing SMART overall: still PASS (SMART’s overall bit is conservative until threshold) 3k+ remapped sectors is a platter going bad — not a cable blip. V…