Quick tips and learnings

118 posts latest post 2026-04-03
Publishing rhythm
Dec 2025 | 2 posts

I have a specific need for counting the number of lines in a file quickly. At work we use S3 for data storage during our Kedro pipeline development, and in the development process we may end up orphaning several datasets. In order to keep our workspace clean I have a short utility that compares the datasets in a Kedro DataCatalog with the files in the relevant S3 location.

To get that list I run an internal tool like this:

kedro our-liter | grep s3 >> orphaned_datasets.txt

This simply parses our internal linter for the lines releated to my s3 linter utility and pipes those lines to a file.

To get a quick idea of how out of wack a pipeline is I could open the text file in vim, git it with the G and see what line number I’m on but I’m way too lazy for that…

AWK #

awk 'END {print NR}' orphaned_datasets.txt gives me the number of lines and I can alias this to whatever feels appropriate in my zshrc!

built-ins for the win!

Did you know you can spell check in Vim?!

Vim Spell check

Without... #

Here is a missspelled word.

  <h3>With!</h3>
  <p>Here is a <u>missspelled</u> word.</p>

What is this magic??? #

set: spell spelllang=en_us

Custom words? #

Sometimes there’s things that are words to you but not the default spell checker…

Common example: package names!

plotly, streamlit, psutil, etc etc…

You can easily add these to your vim config by hitting zw ontop of the word!

I am working on a project to create a small system monitoring dashboard using the python psutil library.

The repo is here (if you want actual system monitoring please use netdata).

I’m using streamlit and plotly for the webserver, design, and plotting at the moment.

My Use Case #

I needed a way to refresh my plotly charts with a fixed window of time so that I’m able to just see relevant recent data instead of cramming all data for all time into one plot that’s 500 pixels wide…

Checking the length of arrays or lists every time I get a new piece of data feels kind of dumb and I thought “python must have a way to do this”…

“This” meaning, update values in a fixed length array without reallocating memory or recreating a copy of the list

Deques #

Enter the deque. It means “double ended queue” and is in general an Iterable that you can append values to either side or pop values from either side.

The init signature is straightforward enough and I’m sure there’s more to them than I know yet but here’s how I use it…

from collections import deque

my_deque = deque([1,2,3])

This gives us my_deque, created from an iterable, with several familiar methods like index, extend, append, etc. However there’s some new ones too such as appendleft and popleft.

my_deque.appendleft('a')
print(my_dequqe)
>>> deque(['a', 1, 2, 3])

my_deque.popleft()
>>> 'a'

<!--markata-attribution-->
print(my_deque)
>>> deque([1, 2, 3])

These are handy ways to manipulate the iterable that I needed for the arrays I plot with plotly!

See my follow-up to this on using Deques with plotly and streamlit to create a quick “dashboard” with live streaming data!

follow-up

EDA #

I work with data a lot, but the nature of my job isn’t to dive super deep into a small amount of datasets, I’m often jumping between several projects every day and need to just get a super quick glance at some tables to get a high level view.

When I’m doing more interactive exploration I’ve graduated from Jupyter cells with df_N.head() to using an amazing tool called visidata

However, Visidata is a terminal based application and I’m often in an iPython console… so is there a way to move even faster for my super quick summary views?

yes!

Skimpy #

First thing to do is pip install skimpy and then it’s as easy to get some summary stats with skimpy <data>

Skimpy ZSH

This is super nice for seeing missing values in particular as well as the distribution shape of the data.

iPython #

But wait… I just said I’m normally in an iPython session but that was called from zsh.. If I’m hoping back into zsh I might as well use visidata to have more powerful exploration at my fingertips. So… can I see this table quickly without breaking my iPython workflow?

Of course you can with magic!

Skimpy iPython

The above assumes you’re looking at a file, like you would in the terminal. skimpy works even better in iPython with from skimpy import skim then pass any DataFrame to skim!

Skimpy iPython2

I like to keep my workspace clean and one thing that I don’t personally love looking at is the __pycache__ directory that pops up after running some code. The *.pyc files that show up there are python bytecode and they are cached to make subsequent runs a tad faster. My stuff never really needs this bonus speed boost and so I came across a neat tool called pyclean!

Pyclean #

The easiest way (in my opinion) to run pyclean is to just use pipx run.

sandbox/src  🌱 main 🗑️  ×3🛤️  ×2via 🐍 v3.8.11 (sandbox)  took 9s
❯ ls
abcmeta.py  __pycache__  python-print-align.py  system-monitor-psutils.py

sandbox/src  🌱 main 🗑️  ×3🛤️  ×2via 🐍 v3.8.11 (sandbox)
❯ pipx run pyclean .
⚠️  pyclean is already on your PATH and installed at /usr/bin/pyclean. Downloading and running anyway.
Cleaning directory .
Total 1 files, 1 directories removed.

sandbox/src  🌱 main 🗑️  ×3🛤️  ×2via 🐍 v3.8.11 (sandbox)
❯ ls
abcmeta.py  python-print-align.py  system-monitor-psutils.py

Why not bash? #

You could accomplish something similar with rm **/*.pyc or find -n '*.py?' -delete but there’s a chance you’ll find something you don’t love gone. Also this won’t help our poor Windows friends out there! pyclean is fully python so it’s OS independent.

Credits! #

repo

Mike Driscoll has been posting some awesome posts about psutil lately. I’m interested in making my own system monitoring dashboard now using this library. I don’t expect it to compete with Netdata or Glances but it’ll just be for fun to see how Python can solve this problem!

Repo coming soon

Example code: #

Here’s a short snippit to get used/available/total RAM and disk space (on partitions that you probably care about)


import psutil
import socket

print(f"System Memory used: {psutil.virtual_memory().used // (1024 ** 3)} GB")
print(f"System Memory available: {psutil.virtual_memory().available // (1024 ** 3)} GB")
print(f"System Memory total: {psutil.virtual_memory().total // (1024 ** 3)} GB")


print(f"Hostname: {socket.gethostname()}")

partitions = psutil.disk_partitions()

for part in partitions:
    mnt = part.mountpoint
    if "snap" in mnt or "boot" in mnt:
        continue
    disk = psutil.disk_usage(mnt)
    print(f"Usage at {mnt} on {part.device}: {disk.used // (1024 ** 3)} GB")
    print(f"Free at {mnt} on {part.device}: {disk.free // (1024 ** 3)}GB")
    print(f"Total at {mnt} on {part.device}: {disk.total // (1024 ** 3)}GB")

Bonus Ipython tip! Save this to a script called my_script.py and in Ipython you can %run -m my_script to run it!

project ↪ main v3.8.11 ipython
❯ %run -m system-monitor-psutils
System Memory used: 25 GB
System Memory available: 5 GB
System Memory total: 31 GB
Hostname: ryzen-3600x
Usage at / on /dev/nvme1n1p2: 81 GB
Free at / on /dev/nvme1n1p2: 351 GB
Total at / on /dev/nvme1n1p2: 456 GB

If you work with a template for several projects then you might sometimes need to do the same action across all repos. A good example of this is updating a package in requirements.txt in every project, or refactoring a common module. If you have several repos to do this across then it can be time consuming… enter mu-repo

Mu #

mu-repo is an awesome cli tool for working with multiple git repositories at the same time. There are several things you can do:

  1. mu status will give you the git status of every registered repo (see below)
  2. mu sh will let you execute system level commands in every repo
  3. mu stash will stash all changes across all registered repos
  4. There’s literally a ton more but these are some handy ones

Registration #

mu tracks its own groups, and there is a default group when no particular one is active. It’s as simple as mu register proj1 prog2 ... to get repos registered


❯ mu register proj1 proj2
Repository: proj1 registered
Repository: proj2 registered

❯ mu status

  proj1 : git status
    On branch main

    No commits yet

    Untracked files:
    (use "git add <file>..." to include in what will be committed)
    requirements.txt

    nothing added to commit but untracked files present (use "git add" to track)

  proj2 : git status
    On branch main

    No commits yet

    Changes to be committed:
    (use "git rm --cached <file>..." to unstage)
    new file:   requirements.txt


Working with mu #

As you can see above I have two projects each with a requirements.txt added but not committed yet. Using mu I can stage this change across both repos at once.


❯ mu add requirements.txt

  proj1 : git add requirements.txt

  proj2 : git add requirements.txt

Then as you might imagine, I can make the commit in each repo


❯ mu commit -m "Add requirements.txts"

  proj1 : git commit -m Add requirements.txts
    [main (root-commit) 18376d7] Add requirements.txts
    1 file changed, 1 insertion(+)
    create mode 100644 requirements.txt

  proj2 : git commit -m Add requirements.txts
    [main (root-commit) 18376d7] Add requirements.txts
    1 file changed, 1 insertion(+)
    create mode 100644 requirements.txt

mu groups #

The other thing I got a lot of use out of recently was mu’s groups. At work I have about 40 repos cloned that are all based on the same kedro pipeline template. Some of these projects have been deprecated. I also have several more repos that are not kedro template - custom libraries or something. group let me utilize mu across different groups of repos.

Say proj2 is a deprecated project that I don’t need to worry about making changes to anymore. I don’t just have to unregister it, instead I can make a group called “active” and register proj1 in that group


❯ mu group add active --empty

~/personal
❯ mu group add deprecated --empty

~/personal
❯ mu group
  active
* deprecated

The * tells me which group is active. The --empty flag tells mu to not add all registered repos to that group. If I don’t want to use any groups then mu group reset will go back to the default group with all registered repos.

With groups I can register only the repos that I want to be working across in their own group and not worry about affecting other repos with my batch changes!

ABCMeta #

I don’t do a lot of OOP currently, but I have been on a few heavy OOP projects and this ABCMeta and abstractmethod from abc would’ve been super nice to know about!

If you are creating a library with classes that you expect your users to extend, but you want to ensure that any extension has explicit methods defined then this is for you!.

from abc import ABCMeta, abstractmethod
class Family(metaclass=ABCMeta):
    @abstractmethod
    def get_dad(self):
        """Any extension of the Family class must implement a `get_dad` method"""

class MyFamily(Family):
    pass

If I try to instantiate MyFamily I will not be allowed:


❯ my_fam = MyFamily()
╭─────────────────────────────── Traceback (most recent call last) ────────────────────────────────╮
│ <ipython-input-8-ecb8e21ce815>:1 in <module>                                                     │
╰──────────────────────────────────────────────────────────────────────────────────────────────────╯
TypeError: Can't instantiate abstract class MyFamily with abstract methods get_dad

abcmetadata

In order for me to extend Family I have to implement the method get_dad

class MyFamily(Family):
    def get_dad(self):
        return "Me"

Now everything works as expected and I can sleep well knowing no one can extend my base class without creating methods I know they need.


my_fam = MyFamily()

my_fam.get_dad()
'Me'

I am personally trying to use logger instead of print in all of my code, however I learned from [@Python-Hub] that you can align printouts using print with f-strings!.

This little python script shows how options in the f-string can format the printout.


import random

variables = "Foo Bar Baz Bing".split()
scores = random.sample(range(1, 11), len(variables))

print("*" * 30)
print("\n")
print("With 'varable' left aligned")
for varable, score in zip(variables, scores):
    print(f"{varable:<10} | {score}")

print("*" * 30)
print("\n")
print("With 'varable' right aligned")
for varable, score in zip(variables, scores):
    print(f"{varable:>15} | {score}")

print("*" * 30)
print("\n")
print("With 'varable' center aligned")
for varable, score in zip(variables, scores):
    print(f"{varable:^5} | {score}")

pyprintalign

Being lazy #

I almost exclusively use Python for my job and have been eye-balls deep in it for almost 5 years but I really lack in-depth knowledge of builtins. I recently learned of an awesome builtin called calendar that has way more than I know about for sure but I’m glad I know it’s here now!

I only needed it because I was too lazy to hard code the 7 weekdays into my module but it turns out there’s a lot of useful things like calendar.isleap()!

builtin ## Future use

I’m not exactly sure what will come my way where calendar will be super relevant but like anything, I’m just glad to know it exists for when the time arises!