Posts tagged: tech
All posts with the tag "tech"
Backups interrupted by full disk usage
Windows Update Broke Wifi
homelab-computer-vision-pipelines
I keep running out of space with my swap getting maxxed out… I don’t know why but U-Blue uses Zram already and apparently I can easily override the defaults:
in usr/lib/systemd/zram-generator.conf
# This config file enables a /dev/zram0 device with the default settings:
# — size — same as available RAM or 8GB, whichever is less
# — compression — most likely lzo-rle
#
# To disable, uninstall zram-generator-defaults or create empty
# /etc/systemd/zram-generator.conf file.
[zram0]
zram-size = min(ram, 8192)
And so I just made my own file in /etc/systemd/zram-generator.conf:
[zram0]
zram-size = min(ram, 16384)
TIL
Today I learned that the .info() of a pandas.DataFrame will always
give you an answer, but it is wildly difficult to know how accurate it is, because it depends
on the underlying data types.
If everything is a NumPy Dtype, then the result you get is the amount of memory NumPy is using
under the hood, which is very accurate.
But as soon as you have something like a Python string in the data frame, now pandas will
tell you how much memory the NumPy data types are using, and it will give you some quick
estimate of how much memory space the Python string types are using because numpy is just storing
a pointer to that data in memory, but this is far less
accurate because it doesn’t actually know what the strings are.
Memory Usage Deep or True #
To get an honest answer about a DataFrame’s memory usage, you need to pass memory_usage='deep' to the .info() method.
# This is the default
In [50]: df.info(memory_usage=True)
<class 'pandas.core.frame.DataFrame'>
RangeIndex: 10 entries, 0 to 9
Data columns (total 3 columns):
# Column Non-Null Count Dtype
--- ------ -------------- -----
0 s1 10 non-null int64
1 s2 10 non-null int64
2 s3 10 non-null object
dtypes: int64(2), object(1)
memory usage: 372.0+ bytes
Notice that the memory usage has a + in it? You can force a deeper analysis by passing memory_usage='deep'
In [51]: df.info(memory_usage='deep')
<class 'pandas.core.frame.DataFrame'>
RangeIndex: 10 entries, 0 to 9
Data columns (total 3 columns):
# Column Non-Null Count Dtype
--- ------ -------------- -----
0 s1 10 non-null int64
1 s2 10 non-null int64
2 s3 10 non-null object
dtypes: int64(2), object(1)
memory usage: 792.0 bytes
I heard about SearXNG on a couple podcasts and saw it trending on GitHub several times before I finally decided to stand it up. I used it transparently when trying out khoj, a self-hosted AI LLM agent playground kind of a thing, but I’ve also been messing around with Open-WebUI to have a self-hosted ChatGPT-like experience. I don’t know if SearchXNG used to be harder to set up, but it was pretty simple with a Docker Compose up and a couple configuration options given my personal homelab setup. Using it feels very nice.
services:
redis:
container_name: redis
image: docker.io/valkey/valkey:8-alpine
command: valkey-server --save 30 1 --loglevel warning
restart: unless-stopped
volumes:
- ./data/redis:/data
logging:
driver: "json-file"
options:
max-size: "1m"
max-file: "1"
searxng:
container_name: searxng
image: docker.io/searxng/searxng:latest
restart: unless-stopped
ports:
- "8080:8080"
volumes:
- ./data/searxng:/etc/searxng:rw
env_file: .env
environment:
# - SEARXNG_BASE_URL=https://${SEARXNG_HOSTNAME:-localhost}/
- UWSGI_WORKERS=${SEARXNG_UWSGI_WORKERS:-4}
- UWSGI_THREADS=${SEARXNG_UWSGI_THREADS:-4}
logging:
driver: "json-file"
options:
max-size: "1m"
max-file: "1"
I love the self-hosted aspect. I love not seeing any ads on my search results. I did a few searches where I know what results to expect and it did okay. It is good right now at filtering out garbage in my results.
So I look forward to tweaking it and using it as a search backend with open web UI. Next on my list is having enough resources to run Ollama and a stable diffusion generator at the same time and have image generation working through open web UI.