Thoughts on tech

186 posts latest post 2026-10-06
Publishing rhythm
May 2026 | 3 posts

Backups interrupted by full disk usage

I just got a message from HCIO that my primary backup script is late… This happens every now and then but I decided to check on it… Quickly in and I notice it takes a long time to connect, but nothing errors out… I’ve been in the game long enough to think this kind of latency is often disk IO. After waiting a bit and haulting the zshrc sourcing, I hit a and bam… There’s a container logfile that’s about 80GB sitting in … Easy enough to remove, but I need to setup some alerts on disk usage, and identify the container to limit the log file size Useful Notifications I use healthchecks.io to monitor several scripts. You can self-host it but I use the hosted service since he has simple Signal integration

Windows Update Broke Wifi

Windows Update Behind My Back After an unapproved windows update on a machine I help administer for my church, the wifi became super finnicky. I had installed a pretty standard AX5400 WiFi 6 PCIe adapter and things were great for a while. ISP and Top-Tier Laziness # The ISP installed the modem and router/AP combo all in one room encased in CMU block in the basement… That is in the worst possible spot in the entire building and the Windows computer of note is as far away from it as possible - but we were cooking just fine with that PCIe card. Suddenly though, after the update and reboot the machine quit receiving an IPv4 address via that interface… Many applications broke in odd ways, and the browser connected to some websites but not others. State Matters # It took a lot of troubleshooting because of 2 things in the mix… Tailscale A second Wi-Fi adapter via USB dongle that was installed while I was out of town (due to the WiFi issues) So because of these 2 things, and mostly the USB ad…

homelab-computer-vision-pipelines

Done in 11 seconds! Subtitle and audio files are in the outputs folder. I wanted to talk through an idea I have for some computer vision pipelines at home. I heard of a tool called Roboflow, which is like a no-code computer vision pipeline training and inference tool. It seems very cool. I also use Frigate at home, and Frigate supports lots of detection on the NVR side. And so deploying a model in Frigate or using Roboflow, depending on what kind of customization you would need, you can have these detections potentially be as accurate as you want, at least with Roboflow. The thing is, I don’t know about Frigate, but my ideas are if I have a model at home that knows me and my wife and my children by name, then when monitored areas issue detections, for example, say I’m in the backyard and the temperature is above 80 degrees, have a home assistant kick on a sprinkler system or something like that. And if my wife is back there and it’s 80 degrees, don’t do that because she wouldn’t want t…

I keep running out of space with my swap getting maxxed out… I don’t know why but U-Blue uses Zram already and apparently I can easily override the defaults:

in usr/lib/systemd/zram-generator.conf


# This config file enables a /dev/zram0 device with the default settings:
# — size — same as available RAM or 8GB, whichever is less
# — compression — most likely lzo-rle
#
# To disable, uninstall zram-generator-defaults or create empty
# /etc/systemd/zram-generator.conf file.
[zram0]
zram-size = min(ram, 8192)

And so I just made my own file in /etc/systemd/zram-generator.conf:


[zram0]
zram-size = min(ram, 16384)

TIL

Today I learned that the .info() of a pandas.DataFrame will always give you an answer, but it is wildly difficult to know how accurate it is, because it depends on the underlying data types. If everything is a NumPy Dtype, then the result you get is the amount of memory NumPy is using under the hood, which is very accurate. But as soon as you have something like a Python string in the data frame, now pandas will tell you how much memory the NumPy data types are using, and it will give you some quick estimate of how much memory space the Python string types are using because numpy is just storing a pointer to that data in memory, but this is far less accurate because it doesn’t actually know what the strings are.

Memory Usage Deep or True #

To get an honest answer about a DataFrame’s memory usage, you need to pass memory_usage='deep' to the .info() method.

# This is the default
In [50]: df.info(memory_usage=True)
<class 'pandas.core.frame.DataFrame'>
RangeIndex: 10 entries, 0 to 9
Data columns (total 3 columns):
 #   Column  Non-Null Count  Dtype 
---  ------  --------------  ----- 
 0   s1      10 non-null     int64 
 1   s2      10 non-null     int64 
 2   s3      10 non-null     object
dtypes: int64(2), object(1)
memory usage: 372.0+ bytes

Notice that the memory usage has a + in it? You can force a deeper analysis by passing memory_usage='deep'

In [51]: df.info(memory_usage='deep')
<class 'pandas.core.frame.DataFrame'>
RangeIndex: 10 entries, 0 to 9
Data columns (total 3 columns):
 #   Column  Non-Null Count  Dtype 
---  ------  --------------  ----- 
 0   s1      10 non-null     int64 
 1   s2      10 non-null     int64 
 2   s3      10 non-null     object
dtypes: int64(2), object(1)
memory usage: 792.0 bytes

I heard about SearXNG on a couple podcasts and saw it trending on GitHub several times before I finally decided to stand it up. I used it transparently when trying out khoj, a self-hosted AI LLM agent playground kind of a thing, but I’ve also been messing around with Open-WebUI to have a self-hosted ChatGPT-like experience. I don’t know if SearchXNG used to be harder to set up, but it was pretty simple with a Docker Compose up and a couple configuration options given my personal homelab setup. Using it feels very nice.

services:
  redis:
    container_name: redis
    image: docker.io/valkey/valkey:8-alpine
    command: valkey-server --save 30 1 --loglevel warning
    restart: unless-stopped
    volumes:
      - ./data/redis:/data
    logging:
      driver: "json-file"
      options:
        max-size: "1m"
        max-file: "1"

  searxng:
    container_name: searxng
    image: docker.io/searxng/searxng:latest
    restart: unless-stopped
    ports:
      - "8080:8080"
    volumes:
      - ./data/searxng:/etc/searxng:rw
    env_file: .env
    environment:
      # - SEARXNG_BASE_URL=https://${SEARXNG_HOSTNAME:-localhost}/
      - UWSGI_WORKERS=${SEARXNG_UWSGI_WORKERS:-4}
      - UWSGI_THREADS=${SEARXNG_UWSGI_THREADS:-4}
    logging:
      driver: "json-file"
      options:
        max-size: "1m"
        max-file: "1"

I love the self-hosted aspect. I love not seeing any ads on my search results. I did a few searches where I know what results to expect and it did okay. It is good right now at filtering out garbage in my results.

SearXNG

So I look forward to tweaking it and using it as a search backend with open web UI. Next on my list is having enough resources to run Ollama and a stable diffusion generator at the same time and have image generation working through open web UI.

My MCP Configuration

My MCP The Tools # Docker # RAGDocs # Sequential Thinking # Git # Not Tried Yet # https://github.com/modelcontextprotocol/servers/tree/main/src/sqlite The Config #

I was wrecked by a weird combo of » and -e

TL;DR # If state matters then check it in the beginning or handle it on a failure… Let me explain I ran into some trouble recently almost losing some encrypted data… Now this would be the second time I’ve had that happen, so I’m going to write a little bit about it now that I figured it out The Stage # Here’s the scenario - I use to keep some sensitive data in a few public repos, and I keep the keys I use in Bitwarden Secrets Manager. I use as a command runner and have common and recipes throughout many justfiles. It might look something like this Spoiler # Above is a good example of the 2 recipes, but prior to this morning they looked like this Problem # So here’s the story - in real life I started using distrobox when I ran into this and I thought it was an issue with different versions of ansible installed in distrobox and on my desktop. After a few more odd deployment workflows I think I’ve determined that the real problem is that appends to a file (see the issue yet?), but if my r…