TODO
title = "my Title"
eval('"my" in title')
>>> True
print("hello, world"); print("formatting")
No themes match.
title = "my Title"
eval('"my" in title')
>>> True
print("hello, world"); print("formatting")
I wrote up a little on exporting DataFrames to markdown and html here
But I’ve been playing with a web app for with lists and while I’m toying around I learned you can actually give your tables some style with some simple css classes!
Reminder that if you have a dataframe, df, you can df.to_html() to get an HTML table of your dataframe.
Well you can pass some classes to make it look super nice!
I don’t know anything really about CSS so I won’t pretend otherwise, but as I was learning about bootstrap that’s where I stumbled upon this…
There are several classes you can pass but I found really good luck with table-bordered and table-dark for my use case
df.to_html(classes=["table table-bordered table-dark"])
| Unnamed: 0 | mpg | cyl | disp | hp | drat | wt | qsec | vs | am | gear | carb |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Mazda RX4 | 21.0 | 6 | 160.0 | 110 | 3.90 | 2.620 | 16.46 | 0 | 1 | 4 | 4 |
| Mazda RX4 Wag | 21.0 | 6 | 160.0 | 110 | 3.90 | 2.875 | 17.02 | 0 | 1 | 4 | 4 |
| Datsun 710 | 22.8 | 4 | 108.0 | 93 | 3.85 | 2.320 | 18.61 | 1 | 1 | 4 | 1 |
| Hornet 4 Drive | 21.4 | 6 | 258.0 | 110 | 3.08 | 3.215 | 19.44 | 1 | 0 | 3 | 1 |
| Hornet Sportabout | 18.7 | 8 | 360.0 | 175 | 3.15 | 3.440 | 17.02 | 0 | 0 | 3 | 2 |
Crack open ipython and make a dataframe, then df.to_html(classes=["table table-bordered table-dark"]), copy the output (minus the quote marks ipython uses to denote the string type) that into my-file.html, open that up in a browser and be amazed!
For added effeciency try using pyperclip to copy the output right to your clipboard!
pip install pyperclip and then pyperclip.copy(df.to_html(classes=["table table-bordered table-dark"]))
pandas.DataFrames are pretty sweet data structures in Python.
I do a lot of work with tabular data and one thing I have incorporated into some of that work is automatic data summary reports by throwing the first few, or several relevant, rows of a dataframe at a point in a pipeline into a markdown file.
Pandas has a method on DataFrames that makes this 100% trivial!
Say we have a dataframe, df… then it’s literally just: df.to_markdown()
❯ df.head()
Unnamed: 0 mpg cyl disp hp drat wt qsec vs am gear carb
0 Mazda RX4 21.0 6 160.0 110 3.90 2.620 16.46 0 1 4 4
1 Mazda RX4 Wag 21.0 6 160.0 110 3.90 2.875 17.02 0 1 4 4
2 Datsun 710 22.8 4 108.0 93 3.85 2.320 18.61 1 1 4 1
3 Hornet 4 Drive 21.4 6 258.0 110 3.08 3.215 19.44 1 0 3 1
4 Hornet Sportabout 18.7 8 360.0 175 3.15 3.440 17.02 0 0 3 2
In ipython I can call the method and get a markdown table back as a string
mental-data-lake new-posts via 3.8.11(mental-data-lake) ipython
❯ df.head().to_markdown()
'| | Unnamed: 0 | mpg | cyl | disp | hp | drat | wt | qsec | vs | am | gear | carb |\n|---:|:------------------|------:|------:|-------:|-----:|-------:|------:|-------:|-----:|-----:|-------:|-------:|\n| 0 | Mazda RX4 | 21 | 6 | 160 | 110 | 3.9 | 2.62 | 16.46 | 0 | 1 | 4 | 4 |\n| 1 | Mazda RX4 Wag | 21 | 6 | 160 | 110 | 3.9 | 2.875 | 17.02 | 0 | 1 | 4 | 4 |\n| 2 | Datsun 710 | 22.8 | 4 | 108 | 93 | 3.85 | 2.32 | 18.61 | 1 | 1 | 4 | 1 |\n| 3 | Hornet 4 Drive | 21.4 | 6 | 258 | 110 | 3.08 | 3.215 | 19.44 | 1 | 0 | 3 | 1 |\n| 4 | Hornet Sportabout | 18.7 | 8 | 360 | 175 | 3.15 | 3.44 | 17.02 | 0 | 0 | 3 | 2 |'
You can drop that string into a markdown file and using any reader that supports the rendering you’ll have a nicely formated table of example data in whatever report you’re making!
Just like markdown, you can export a dataframe to html with df.to_html() and use that if it’s more appropriate for your use case:
'<table border="1" class="dataframe">\n <thead>\n <tr style="text-align: right;">\n <th></th>\n <th>Unnamed: 0</th>\n <th>mpg</th>\n <th>cyl</th>\n <th>disp</th>\n <th>hp</th>\n <th>drat</th>\n <th>wt</th>\n <th>qsec</th>\n <th>vs</th>\n <th>am</th>\n <th>gear</th>\n <th>carb</th>\n </tr>\n </thead>\n <tbody>\n <tr>\n <th>0</th>\n <td>Mazda RX4</td>\n <td>21.0</td>\n <td>6</td>\n <td>160.0</td>\n <td>110</td>\n <td>3.90</td>\n <td>2.620</td>\n <td>16.46</td>\n <td>0</td>\n <td>1</td>\n <td>4</td>\n <td>4</td>\n </tr>\n <tr>\n <th>1</th>\n <td>Mazda RX4 Wag</td>\n <td>21.0</td>\n <td>6</td>\n <td>160.0</td>\n <td>110</td>\n <td>3.90</td>\n <td>2.875</td>\n <td>17.02</td>\n <td>0</td>\n <td>1</td>\n <td>4</td>\n <td>4</td>\n </tr>\n <tr>\n <th>2</th>\n <td>Datsun 710</td>\n <td>22.8</td>\n <td>4</td>\n <td>108.0</td>\n <td>93</td>\n <td>3.85</td>\n <td>2.320</td>\n <td>18.61</td>\n <td>1</td>\n <td>1</td>\n <td>4</td>\n <td>1</td>\n </tr>\n <tr>\n <th>3</th>\n <td>Hornet 4 Drive</td>\n <td>21.4</td>\n <td>6</td>\n <td>258.0</td>\n <td>110</td>\n <td>3.08</td>\n <td>3.215</td>\n <td>19.44</td>\n <td>1</td>\n <td>0</td>\n <td>3</td>\n <td>1</td>\n </tr>\n <tr>\n <th>4</th>\n <td>Hornet Sportabout</td>\n <td>18.7</td>\n <td>8</td>\n <td>360.0</td>\n <td>175</td>\n <td>3.15</td>\n <td>3.440</td>\n <td>17.02</td>\n <td>0</td>\n <td>0</td>\n <td>3</td>\n <td>2</td>\n </tr>\n </tbody>\n</table>'
My blog will render that html into a nice table! (After removing new line characters)
| Unnamed: 0 | mpg | cyl | disp | hp | drat | wt | qsec | vs | am | gear | carb |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Mazda RX4 | 21.0 | 6 | 160.0 | 110 | 3.90 | 2.620 | 16.46 | 0 | 1 | 4 | 4 |
| Mazda RX4 Wag | 21.0 | 6 | 160.0 | 110 | 3.90 | 2.875 | 17.02 | 0 | 1 | 4 | 4 |
| Datsun 710 | 22.8 | 4 | 108.0 | 93 | 3.85 | 2.320 | 18.61 | 1 | 1 | 4 | 1 |
| Hornet 4 Drive | 21.4 | 6 | 258.0 | 110 | 3.08 | 3.215 | 19.44 | 1 | 0 | 3 | 1 |
| Hornet Sportabout | 18.7 | 8 | 360.0 | 175 | 3.15 | 3.440 | 17.02 | 0 | 0 | 3 | 2 |
I try to commit a lot, and I also try to write useful tests appropriate for the scope of work I’m focusing on, but sometimes I drop the ball…
Whether by laziness, ignorance, or accepted tech debt I don’t always code perfectly and recently I was dozens of commits into a new feature before realizing I broke something along the way that none of my tests caught…
Before today I would’ve manually reviewed every commit to see if something obvious slipped by me (talk about a time suck 😩)
There must be a better way
git bisect is the magic sauce for this exact problem…
You essentially create a range of commits to consider and let git bisect guide you through them in a manner akin to Newton’s method for finding the root of a continuous function.
Start with git bisect start and then choose the first good commit (ie. a commit you know the bug isn’t present in)
sandbox bisect-post ×1 via v3.8.11(sandbox) on (us-east-1)
❯ git bisect start
sandbox bisect-post (BISECTING) ×1 via v3.8.11(sandbox) on (us-east-1)
❯ git bisect good 655332b
bisect-post HEAD main ORIG_HEAD
5b31e1e -- [HEAD] add successful print (52 seconds ago)
308247b -- [HEAD^] init another loop (77 seconds ago)
4555c59 -- [HEAD^^] introduce bug (2 minutes ago)
9cf6d55 -- [HEAD~3] add successful loop (3 minutes ago)
bcb41c3 -- [HEAD~4] change x to 10 (4 minutes ago)
3c34aac -- [HEAD~5] init x to 1 (4 minutes ago)
12e53bd -- [HEAD~6] print cwd (4 minutes ago)
655332b -- [HEAD~7] add example.py (10 minutes ago) # <- I want to start at this commit
59e0048 -- [HEAD~8] gitignore (23 hours ago)
fb9e1fb -- [HEAD~9] add reqs (23 hours ago)
sandbox bisect-post (BISECTING) ×1 via v3.8.11(sandbox) on (us-east-1)
❯ git bisect bad 5b31e1e
bisect-post ORIG_HEAD
HEAD refs/bisect/good-655332b6c384934c2c00c3d4aba3011ccc1e5b57
main
5b31e1e -- [HEAD] add successful print (5 minutes ago) # <- I start here with the "bad" commit
308247b -- [HEAD^] init another loop (6 minutes ago)
4555c59 -- [HEAD^^] introduce bug (6 minutes ago)
9cf6d55 -- [HEAD~3] add successful loop (7 minutes ago)
bcb41c3 -- [HEAD~4] change x to 10 (8 minutes ago)
3c34aac -- [HEAD~5] init x to 1 (9 minutes ago)
12e53bd -- [HEAD~6] print cwd (9 minutes ago)
655332b -- [HEAD~7] add example.py (14 minutes ago)
59e0048 -- [HEAD~8] gitignore (23 hours ago)
fb9e1fb -- [HEAD~9] add reqs (23 hours ago)
After starting bisect with a “good” start commit and a “bad” ending commit we can let git to it’s thing!
Git checksout a commit somewhere about halfway between the good and bad commit so you can see if your bug is there or not.
sandbox bisect-post (BISECTING) ×1 via v3.8.11(sandbox) on (us-east-1)
❯ git bisect bad 5b31e1e
Bisecting: 3 revisions left to test after this (roughly 2 steps)
[bcb41c3854e343eade85353683f2c1c4ddde4e04] change x to 10
sandbox HEAD (bcb41c38) (BISECTING) ×1 via v3.8.11(sandbox) on (us-east-1)
❯
In my example here I have a python script with some loops and print statements - they aren’t really relevant, I just wanted an easy to follow git history.
So I check to see if the bug is present or not either by running/writing tests or replicating the bug somehow.
In this session commit bcb41c38 is actually just fine, so I do git bisect good
sandbox HEAD (bcb41c38) (BISECTING) ×1 via v3.8.11(sandbox) on (us-east-1)
❯ git bisect good
Bisecting: 1 revision left to test after this (roughly 1 step)
[4555c5979268dff6c475365fdc5ce1d4a12bd820] introduce bug
And we see that git moves on to checkout another commit…
In this case the next commit is the one where I introduced a bug
git bisect bad then gives me:
sandbox HEAD (4555c597) (BISECTING) ×1 via v3.8.11(sandbox) on (us-east-1)
❯ git bisect bad
Bisecting: 0 revisions left to test after this (roughly 0 steps)
[9cf6d55301560c51e2f55404d0d80b1f1e22a33d] add successful loop
At 4555c597 the script works as expected so one more git bisect good yields…
sandbox HEAD (9cf6d553) (BISECTING) ×1 via v3.8.11(sandbox) on (us-east-1)
❯ git bisect good
4555c5979268dff6c475365fdc5ce1d4a12bd820 is the first bad commit
commit 4555c5979268dff6c475365fdc5ce1d4a12bd820
Author: ###########################
Date: Tue May 3 09:00:00 2022 -0500
introduce bug
example.py | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
Git sliced up a range of commits based on me saying of the next one was good or bad and localized the commit that introduced a bug into my workflow!
I didn’t have to manually review commits, click through logs, etc… I just let git checkout relevant commits and I ran whatever was appropriate for reproducing the bug to learn when it was comitted!
pandas.Series.str.contains accepts regular expressions and this is turned on by default!
We often need to filter pandas DataFrames based on several string values in a Series.
Notice that sweet pyflyby import 😁!
sandbox main via 3.8.11(sandbox) ipython
❯ df = pd.DataFrame({"A": ["string1", "string2", "string3"]})
[PYFLYBY] import pandas as pd
sandbox main via 3.8.11(sandbox) ipython
❯ df
A
0 string1
1 string2
2 string3
sandbox main via 3.8.11(sandbox) ipython
❯ df[df.A.str.contains('1') | df.A.str.contains('2')]
A
0 string1
1 string2
And this isn’t the worst thing in the world, especially for such a tiny example…
But what if we had dozens or more values to filter on?
Then it looks so much nicer to create an iterable of the values we want to filter on and join them with an apropriate regex operator (in this case | for inclusive or)
sandbox main via 3.8.11(sandbox) ipython
❯ vals = ["1", "2"] # iterable with whatever is appropriate for your use case
sandbox main via 3.8.11(sandbox) ipython
❯ df[df.A.str.contains("|".join(vals), regex=True)]
A
0 string1
1 string2
This is a super nice and concise way to do the kind of filtering my team does on a daily basis!
Unpacking iterables in python with * is a pretty handy trick for writing code that is just a tiny bit more pythonic than not.
arr: Tuple[Union[int, str]] = (1, 2, 3, 'a', 'b', 'c')
print(arr)
>>> (1, 2, 3, 'a', 'b', 'c')
# the * unpacks the tuple into the individual elements
print(*arr)
>>> 1, 2, 3, 'a', 'b', 'c'
x, y, z, *alphas = arr
# x = 1, y = 2, z = 3
# alphas = [ 'a', 'b', 'c' ]
But @Ned Batchelder showed me via Twitter than you can arbitrarily unpack arguments based on position - it doesn’t have to be done at the beginning or the end!
x, y, *mixed, alpha = arr
# x = 1, y = 2
# mixed = [3, 'a', 'b']
# alpha = 'c'
I’m not entirely sure when I’ll need this but it definitley shows me another example of how flexible python is!
htop is a common command line tool for seeing interactive output of your system resource utilization, running processes, etc.
I’ve always been super confused about htop showing seemingly the same process several times though…
Just hit H…. makes the view a lot nicer 😀
fx is an interactaive JSON viewer for the terminal.
It’s a simple tool built with Charmcli’s Bubble Tea.
The installation with go was broken for me - both via the link and direct from the repo. Now I’m not a gopher so I don’t really know how to fix that.
Luckily npm install fx also works and got me what I needed!
Usage is simple… fx <json file>.
The Github has a few other ways such as curl ... | fx etc.
Type hinting has helped me write code almost as much, if not more, than unit testing.
One thing I love is that with complete type hinting you get a lot more out of your LSP.
Typing dictionaries can be tricky and I recently learned about TypedDict to do exactly what I needed!
It might not be straight up obvious what the problem is, especially if you don’t utilize tools like mypy or flake8 in your development.
My handy-dandy nvim-lsp gives me a lot of feedback when I’m coding and it’s immensely helpful.
So with the LSP giving me constant feedback here’s the issue:
from typing import Dict, List, Union
my_dict: Dict[str, Union[List[str], str]] = {
"key_1": "val_1",
"key_2": ["ls_1", "ls_2"],
}
my_dict["key_2"].pop()
With the above script you’ll get an annoying warning about using pop on key_2.
Maybe you can stomach getting yelled at by your LSP but I like complete silence if at all possible.
TypedDict was the saving grace.
from typing import TypedDict
MyDict = TypedDict("MyDict", {"key_1": str, "key_2": List[str]})
my_typed_dict: MyDict = {
"key_1": "val_1",
"key_2": ["ls_1", "ls_2"],
}
my_typed_dict["key_2"].pop()
I was able to import TypedDict from typing, mypy_extensions, and typing_extensions
With TypedDict you define your custom type, match the first argument to TypedDict with the name of the variable (idk why), then type hint each key you expect in the dict!
It’s super easy and I think puts you into a position of being extremely explicit with your dictionary variables.
This isn’t always desired or appropriate but in most of my use cases it is.
There’s other implementation of TypedDict and while writing this I saw that most of the docs define a class for the type like this:
from typing import TypedDict
class MyDict(TypedDict):
key_1: str
key_2: List[str]
my_dict : MyDict = {'key_1': 'val_1', 'key_2': ["ls_1", "ls_2"]}