<?xml version="1.0" encoding="UTF-8"?>
<?xml-stylesheet href="/atom.xsl" type="text/xsl"?>
<feed xmlns="http://www.w3.org/2005/Atom">
  <title>Posts tagged: backup</title>
  <id>https://pype.dev/tags/backup/atom.xml</id>
  <updated>2026-10-06T21:30:00Z</updated>
  <subtitle>All posts with the tag &#34;backup&#34;</subtitle>
  <link href="https://pype.dev/tags/backup/" rel="alternate" type="text/html"></link>
  <link href="https://pype.dev/tags/backup/atom.xml" rel="self" type="application/atom+xml"></link>
  <author>
    <name>Nic Payne</name>
  </author>
  <generator uri="https://github.com/WaylonWalker/markata-go">markata-go</generator>
  <entry>
    <title>The Faulted Disk: harbor Replacement Writeup</title>
    <id>https://pype.dev/harbor-faulted-disk-replacement/</id>
    <updated>2026-10-06T21:30:00Z</updated>
    <published>2026-10-06T21:30:00Z</published>
    <link href="https://pype.dev/harbor-faulted-disk-replacement/" rel="alternate" type="text/html"></link>
    <summary type="text">The sequel to panicking-led-to-losing-my-desktop — this time the monitoring actually caught the disk dying, and nothing was lost.</summary>
    <content type="html">&lt;p&gt;The sequel to &lt;a href=&#34;/panicking-led-to-losing-my-desktop&#34;&gt;panicking-led-to-losing-my-desktop&lt;/a&gt; — this time the monitoring actually caught the disk dying, and nothing was lost.&lt;/p&gt;&#xA;&lt;h2 id=&#34;what-happened&#34;&gt;&lt;span class=&#34;heading-wear-glyph&#34;&gt;What happened&lt;/span&gt; &lt;a href=&#34;#what-happened&#34; class=&#34;heading-anchor&#34;&gt;#&lt;/a&gt;&lt;/h2&gt;&#xA;&lt;p&gt;&lt;code&gt;harbor&lt;/code&gt; is my replica pool — a 10.9T mirror (2x 12TB) that receives syncoid snapshots from &lt;code&gt;tank&lt;/code&gt;. One side of the mirror, a Seagate Exos &lt;code&gt;ST12000NM0127&lt;/code&gt; (serial &lt;code&gt;ZJV4QFLB&lt;/code&gt;, &lt;code&gt;/dev/sdb&lt;/code&gt;), went &lt;strong&gt;FAULTED&lt;/strong&gt; with 14 read + 22 checksum errors.&lt;/p&gt;&#xA;&lt;pre class=&#34;chroma&#34;&gt;&lt;code&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;mirror-0                              DEGRADED     0     0     0&#xA;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;  ata-ST12000VN0008-2PH103_ZTM0NFDW  ONLINE       0     0     0&#xA;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;  ata-ST12000NM0127_ZJV4QFLB         FAULTED     14     0    22  too many errors&#xA;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;p&gt;The IronWolf mirror side carried the pool — &lt;code&gt;No known data errors&lt;/code&gt;. ZFS redundancy did exactly its job.&lt;/p&gt;&#xA;&lt;h2 id=&#34;the-difference-from-last-time&#34;&gt;&lt;span class=&#34;heading-wear-glyph&#34;&gt;The difference from last time&lt;/span&gt; &lt;a href=&#34;#the-difference-from-last-time&#34; class=&#34;heading-anchor&#34;&gt;#&lt;/a&gt;&lt;/h2&gt;&#xA;&lt;p&gt;Last failure: no monitoring, found out by accident months later, desktop died.&lt;/p&gt;&#xA;&lt;p&gt;This failure: SigNoz + node-exporter&amp;rsquo;s ZFS collector → &lt;code&gt;node_zfs_zpool_state{state=&amp;quot;degraded&amp;quot;}&lt;/code&gt; → alert rule → Gotify → my phone. The gotify notification fired &lt;em&gt;before&lt;/em&gt; I knew anything was wrong.&lt;/p&gt;&#xA;&lt;h2 id=&#34;diagnosis&#34;&gt;&lt;span class=&#34;heading-wear-glyph&#34;&gt;Diagnosis&lt;/span&gt; &lt;a href=&#34;#diagnosis&#34; class=&#34;heading-anchor&#34;&gt;#&lt;/a&gt;&lt;/h2&gt;&#xA;&lt;p&gt;Before &lt;code&gt;zpool clear&lt;/code&gt; or replace — check SMART. &lt;code&gt;smartctl-exporter&lt;/code&gt; already scrapes all disks into SigNoz, so I didn&amp;rsquo;t even need sudo:&lt;/p&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;code&gt;Reallocated_Sector_Ct&lt;/code&gt; raw = &lt;strong&gt;3024&lt;/strong&gt; and counting&lt;/li&gt;&#xA;&lt;li&gt;&lt;code&gt;Offline_Uncorrectable&lt;/code&gt; value/worst still 100 but raw errors climbing&lt;/li&gt;&#xA;&lt;li&gt;SMART overall: still PASS (SMART&amp;rsquo;s overall bit is conservative until threshold)&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&lt;p&gt;3k+ remapped sectors is a platter going bad — not a cable blip. Verdict: replace, don&amp;rsquo;t clear.&lt;/p&gt;&#xA;&lt;h2 id=&#34;the-swap-the-annoying-part&#34;&gt;&lt;span class=&#34;heading-wear-glyph&#34;&gt;The swap (the annoying part)&lt;/span&gt; &lt;a href=&#34;#the-swap-the-annoying-part&#34; class=&#34;heading-anchor&#34;&gt;#&lt;/a&gt;&lt;/h2&gt;&#xA;&lt;p&gt;Hot-plug is never as smooth as it should be:&lt;/p&gt;&#xA;&lt;ol&gt;&#xA;&lt;li&gt;&lt;code&gt;zpool offline harbor &amp;lt;old&amp;gt;&lt;/code&gt; to stop writes to the dead drive — actually skipped; drive fell off the bus on its own&lt;/li&gt;&#xA;&lt;li&gt;New IronWolf &lt;code&gt;ZTN1CQ07&lt;/code&gt; in hand → plugged into ghost&amp;rsquo;s SATA → &lt;strong&gt;didn&amp;rsquo;t enumerate&lt;/strong&gt;&lt;/li&gt;&#xA;&lt;li&gt;Forced rescan (&lt;code&gt;echo &amp;quot;- - -&amp;quot; | sudo tee /sys/class/scsi_host/host*/scan&lt;/code&gt;) → nothing&lt;/li&gt;&#xA;&lt;li&gt;Moved it to the &lt;em&gt;old drive&amp;rsquo;s port&lt;/em&gt; → nothing&lt;/li&gt;&#xA;&lt;li&gt;Sabrent USB dock on aurora → dock enumerated as &lt;code&gt;0B&lt;/code&gt; device, no disk behind it → reseated + replugged → &lt;strong&gt;drive spun up and appeared&lt;/strong&gt;: &lt;code&gt;/dev/sdd&lt;/code&gt;, 10.9T, old ZFS partition table on it (used drive — &lt;code&gt;part1&lt;/code&gt;/&lt;code&gt;part9&lt;/code&gt; layout)&lt;/li&gt;&#xA;&lt;li&gt;Back to ghost, direct SATA → enumerated as &lt;code&gt;sdb&lt;/code&gt;&lt;/li&gt;&#xA;&lt;/ol&gt;&#xA;&lt;p&gt;Lesson: &amp;ldquo;spins up but doesn&amp;rsquo;t enumerate&amp;rdquo; on direct SATA + dock-shows-0B = seating/power problem, not DOA. The drive was fine all along.&lt;/p&gt;&#xA;&lt;h2 id=&#34;the-replace&#34;&gt;&lt;span class=&#34;heading-wear-glyph&#34;&gt;The replace&lt;/span&gt; &lt;a href=&#34;#the-replace&#34; class=&#34;heading-anchor&#34;&gt;#&lt;/a&gt;&lt;/h2&gt;&#xA;&lt;pre class=&#34;chroma&#34;&gt;&lt;code&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;sudo zpool replace harbor &lt;span class=&#34;se&#34;&gt;\&#xA;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;  ata-ST12000NM0127_ZJV4QFLB &lt;span class=&#34;se&#34;&gt;\&#xA;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;  ata-ST12000VN0008-2PH103_ZTN1CQ07&#xA;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;p&gt;Resilver: &lt;strong&gt;8.80T in 18h38m, 0 errors&lt;/strong&gt; (~154M/s). Pool stayed online and usable the whole time.&lt;/p&gt;&#xA;&lt;p&gt;One gotcha: after resilver, &lt;code&gt;zpool status&lt;/code&gt; showed &lt;code&gt;errors: 1 data errors&lt;/code&gt; — but &lt;code&gt;zpool status -v&lt;/code&gt; showed an &lt;strong&gt;empty error list&lt;/strong&gt;. The corrupted data was already repaired; the counter was just stale. &lt;code&gt;sudo zpool clear harbor&lt;/code&gt; → clean.&lt;/p&gt;&#xA;&lt;h2 id=&#34;the-monitoring-that-made-this-possible&#34;&gt;&lt;span class=&#34;heading-wear-glyph&#34;&gt;The monitoring that made this possible&lt;/span&gt; &lt;a href=&#34;#the-monitoring-that-made-this-possible&#34; class=&#34;heading-anchor&#34;&gt;#&lt;/a&gt;&lt;/h2&gt;&#xA;&lt;p&gt;Wired during this session (all now in SigNoz):&lt;/p&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;code&gt;node_zfs_zpool_state&lt;/code&gt; — pool health (node-exporter zfs collector)&lt;/li&gt;&#xA;&lt;li&gt;&lt;code&gt;smartctl-exporter&lt;/code&gt; — SMART attributes incl. reallocated sectors (the early-death signal)&lt;/li&gt;&#xA;&lt;li&gt;&lt;code&gt;zfs-metrics.sh&lt;/code&gt; textfile bridge — sanoid &lt;code&gt;--monitor-*&lt;/code&gt; exit codes, zpool error counters, scrub ages, resilver %, syncoid last-success&lt;/li&gt;&#xA;&lt;li&gt;Alert rules → Gotify → phone&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&lt;p&gt;The replication alerting even proved itself live: syncoid failed twice during the resilver window and I got paged on both. Turned out to be send-time contention — self-healed once resilver finished — but the notification path works end to end.&lt;/p&gt;&#xA;&lt;h2 id=&#34;remaining-homework&#34;&gt;&lt;span class=&#34;heading-wear-glyph&#34;&gt;Remaining homework&lt;/span&gt; &lt;a href=&#34;#remaining-homework&#34; class=&#34;heading-anchor&#34;&gt;#&lt;/a&gt;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;Buy the replacement spare (the shelf&amp;rsquo;s empty now)&lt;/li&gt;&#xA;&lt;li&gt;&lt;code&gt;10Fold&lt;/code&gt; datasets are garbage + have zero snapshots — destroy&lt;/li&gt;&#xA;&lt;li&gt;Every scheduled job emits &lt;code&gt;cron_last_success_epoch&lt;/code&gt; now — if a job dies silently again, I&amp;rsquo;ll know&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&lt;p&gt;Same failure as April, opposite outcome. Redundancy did the protection; monitoring did the &lt;em&gt;detection&lt;/em&gt;. You need both.&lt;/p&gt;&#xA;&lt;hr&gt;&#xA;&lt;p&gt;&lt;em&gt;Co-authored with Devin, who ran the monitoring stack, the SMART diagnosis, and the alerting loop.&lt;/em&gt;&lt;/p&gt;&#xA;</content>
    <author>
      <name>Nic Payne</name>
      <uri>https://pype.dev</uri>
    </author>
  </entry>
  <entry>
    <title>Panicking Led to Losing My Desktop</title>
    <id>https://pype.dev/panicking-led-to-losing-my-desktop/</id>
    <updated>2026-05-13T08:24:00Z</updated>
    <published>2026-05-13T08:24:00Z</published>
    <link href="https://pype.dev/panicking-led-to-losing-my-desktop/" rel="alternate" type="text/html"></link>
    <summary type="text">I thought I had backups handled... can you imagine how the rest of this post is going to go with that intro?</summary>
    <content type="text">&#xA;## False Sense of Security&#xA;&#xA;I thought I had backups handled... can you imagine how the rest of this post is&#xA;going to go with that intro?&#xA;&#xA;To be fair, I do have backups figured out on my NAS - simple ZFS +&#xA;sanoid/syncoid + replica pool + off-site backup with simple restore pathways.&#xA;However, my desktop has been another story entirely. My desktop OS didn&#39;t&#xA;support ZFS when I started checking it out, and I spent weeks thinking through&#xA;how I would backup my HOME directory and projects mostly. I landed on a&#xA;solution that I did validate once, but it fell off my radar and lo&#39; and behold&#xA;that was problematic...&#xA;&#xA;So that backup was based on restic for my home directory, but it was lazy. I&#xA;verified it one time but I had built it with ai, thought I understood the&#xA;restic repo part, and then promptly moved on with my life never buttoning it&#xA;all up. That home directory backup got too big for where I was going to end up&#xA;restoring it. My desktop system was installed on a 4 TB NVMe drive and due to&#xA;the circumstances spawning this blog post I was gonna have to drop to a 500 GB&#xA;boot drive with some extra disks as the storage layer. Overall it looked like:&#xA;&#xA;- A 4 TB SSD that was going bad - old OS&#xA;- A 500 GB SSD, that was going to be my new operating system boot disk&#xA;- A 2 TB SSD that was originally going to be this external storage volume&#xA;  anyways but I never set it up because the version of Aurora I was running&#xA;  didn&#39;t have ZFS, I was married to the idea of using ZFS, so I never ended up&#xA;  taking advantage of the space. However it was moot to me because my boot drive&#xA;  was 4 TB, high quality drive, so I was &#34;just sure&#34; I didn&#39;t need it.&#xA;- AND a 4 TB rust disk as well, which was already a ZFS pool, left over from a&#xA;  previous desktop configuration, and admittedly I had forgotten it was even in the system.&#xA;&#xA;## The Storm&#xA;&#xA;If it wasn&#39;t clear the problem is that my super-nice high-speed 4TB NVMe drive&#xA;was going bad, like really bad. Eventually my OS stopped booting, it was even&#xA;difficult to live-boot from any other ISO due to, I think ultimately, that disk&#xA;causing such extreme latency in the start-up processes that they just failed.&#xA;So I quickly found myself with little-to-no access to my primary desktop&#39;s&#xA;data...&#xA;&#xA;## Where It Went Wrong&#xA;&#xA;What I did is I live booted into an Ubuntu server environment (which took blood&#xA;sweat and tears to successfully get into), mounted my home directory from the 4 TB SSD, and&#xA;tried to continue my restic backup to my NAS, like an idiot. But at the same&#xA;time I also tried to prune it by only backing up a few projects because I&#xA;was getting worried about time. This was the first primary mistake - trying to&#xA;muck with my backup script under duress.&#xA;&#xA;Then over the course of the whole thing it ended up taking over a week to solve&#xA;this when it could&#39;ve been 2-3 days. So say it with me kids - &#34;Don&#39;t make&#xA;decisions under duress&#34;&#xA;&#xA;## Climbing Out&#xA;&#xA;I downloaded opencode and had it help me write the right excludes syntax in my&#xA;restic backup script and got it back up going. That went ok but opencode agents&#xA;had no historical context for why anything was the way it was, and frankly an&#xA;agent would&#39;ve been misled thinking the backup solution was much more solid&#xA;than it was due to how I documented it.&#xA;&#xA;Agents also miss things... in my chat sessions it knew about the other 2&#xA;available disks on the desktop system, I could have done a fresh backup to the 4 TB&#xA;spinning rust disk no problem: install zfs, mount the pool, change target of&#xA;restic, run full... that would&#39;ve been beautifully simple. But instead I&#xA;trimmed it down and backed not-everything up to the NAS over the network, and&#xA;to a different backup target nonetheless... SMH.&#xA;&#xA;As I started to consider which OS I was going to go with next I failed to&#xA;install Pop_OS! or Ubuntu onto the new disc... Then I tried Omarchy and the&#xA;install script just looped. So, I reinstalled Aurora onto the new 500 GB disk&#xA;and then quickly realized I don&#39;t have Firefox tabs, my SSH keys are in that&#xA;restic backup, my ssh config, api keys in hidden files.... Everything is in&#xA;that restic backup... The backup that&#39;s too big to restore to my new boot drive.&#xA;&#xA;But you know what I have? That 2 terabyte disk mounted just fine as a&#xA;ZFS dataset. And I could mount the 4 TB rust disk with zfs as well because this&#xA;version of Aurora has zfs working flawlessly!&#xA;&#xA;## Hindsight&#xA;&#xA;What I should&#39;ve done is so simple... While in that ubuntu live environment I&#xA;should&#39;ve just either updated restic to be a local backup to the 4 TB rust&#xA;disk, or rsync&#39;d my home directory to it plain and simple... I got all in my&#xA;head about not backing up python venvs, node_modules, etc. that I didn&#39;t think&#xA;to just basically carbon copy it all to a healthy disk and then prune it later.&#xA;Then I could&#39;ve synced everything back over that I needed to the new Desktop&#39;s&#xA;$HOME and then scheduled the rsync or restic again to that locally mounted disk.&#xA;&#xA;## The Detail I Left Out&#xA;&#xA;The keen reader might stop to think... why not just mount the old 4TB disk and&#xA;copy what you need to your new desktop? And that&#39;s a prudent question...&#xA;However, in order to get anything installed I had to physically remove the 4TB&#xA;SSD from the motherboard, which was basically a full PC tear-down. From there I&#xA;was able to at least boot in and out of iso&#39;s like you&#39;d otherwise expect, and&#xA;I have a USB/NVMe adapter so I planned to mount the old drive and copy things over from&#xA;there... But sadly... it won&#39;t mount. it&#39;s dead-dead and it appears that&#xA;anything I didn&#39;t save in my days-long-panicked-state is just. gone.&#xA;&#xA;I feel pretty stupid to have not taken advantage of the 2 available disks local&#xA;to the machine, to have naively copied stuff over and dealt with the&#xA;organization later once my OS was back up. I tried to be smart and efficient&#xA;and ended up wasting so much time and losing quite a lot of &#34;stuff&#34;... ideas,&#xA;blog posts that I never committed, etc.&#xA;&#xA;## Current Status&#xA;&#xA;So a few lessons...&#xA;&#xA;1. untested backups are not backups&#xA;2. false backups might be worse than none, although I did at least save a few things so maybe the jury is out here&#xA;3. making decisions while stressed out will lead to missing obviously better pathways... slow down, talk it out&#xA;&#xA;As for my current status - I&#39;m working on [[desktop-setup-2026]] and recovering what I can from my haphazard&#39;d rsyncs in the live ubuntu env I got into. I&#39;m also setting up a new Linux laptop at work at the same time so maybe I&#39;ll hve some workflow changes to write about in the future. For now, it&#39;s nice to be forced to accept that not every idea was that important, the good stuff will come back around, and ultimately computers and shit are just things, they&#39;re not life.&#xA;</content>
    <author>
      <name>Nic Payne</name>
      <uri>https://pype.dev</uri>
    </author>
  </entry>
</feed>