<?xml version="1.0" encoding="UTF-8"?>
<?xml-stylesheet href="/atom.xsl" type="text/xsl"?>
<feed xmlns="http://www.w3.org/2005/Atom">
  <title>Posts tagged: rada</title>
  <id>https://pype.dev/tags/rada/atom.xml</id>
  <updated>2026-04-08T07:40:14Z</updated>
  <subtitle>All posts with the tag &#34;rada&#34;</subtitle>
  <link href="https://pype.dev/tags/rada/" rel="alternate" type="text/html"></link>
  <link href="https://pype.dev/tags/rada/atom.xml" rel="self" type="application/atom+xml"></link>
  <author>
    <name>Nic Payne</name>
  </author>
  <generator uri="https://github.com/WaylonWalker/markata-go">markata-go</generator>
  <entry>
    <title>Data Loading is a Huge Deal</title>
    <id>https://pype.dev/data-loading-is-a-huge-deal/</id>
    <updated>2026-04-08T07:40:14Z</updated>
    <published>2026-04-08T07:40:14Z</published>
    <link href="https://pype.dev/data-loading-is-a-huge-deal/" rel="alternate" type="text/html"></link>
    <summary type="text">I&#39;ve been thinking about the work I am doing and have to do in my role at Cat, in Cat Autonomy, building Forge (see forge-ahead). I feel like I have little...</summary>
    <content type="html">&lt;p&gt;I&amp;rsquo;ve been thinking about the work I am doing and have to do in my role at Cat,&#xA;in Cat Autonomy, building Forge (see &lt;a href=&#34;/forge-ahead/&#34; class=&#34;wikilink&#34; data-title=&#34;Forge Ahead&#34; data-description=&#34;Yesterday&amp;#39;s reflection-contentment-and-work has a second-part this morning. As I was wrapping up a project I didn&amp;#39;t realize the closed-off-ness of leaving......&#34; data-date=&#34;2026-02-17&#34; data-preview=&#34;Yesterday&amp;#39;s reflection-contentment-and-work has a second-part this morning. As I was wrapping up a project I didn&amp;#39;t realize the closed-off-ness of leaving......&#34;&gt;Forge Ahead&lt;/a&gt;). I feel like I have&#xA;little revelations almost every day now, not that it means I&amp;rsquo;m writing&#xA;something amazing and producing it really fast but there&amp;rsquo;s just a whole suite&#xA;of problems that different technologies solve at different levels and the more&#xA;I become aware of the problems that exist, the more the existence of some&#xA;solutions makes sense.&lt;/p&gt;&#xA;&lt;div class=&#34;admonition note&#34;&gt;&#xA;&lt;p class=&#34;admonition-title&#34;&gt;The Problem Perspective&lt;/p&gt;&#xA;&lt;p&gt;I don&amp;rsquo;t know if this is a real thinking technique or if I&amp;rsquo;m onto something&#xA;novel(doubt) but I think a lot in terms of problems - &amp;ldquo;what problem needs&#xA;solving?&amp;rdquo; and that&amp;rsquo;s how I&amp;rsquo;ve come to prioritize my work, it&amp;rsquo;s only been very&#xA;recently that I&amp;rsquo;ve realized I do this and I think I should highlight it&amp;rsquo;s very&#xA;important to document the problem, otherwise every day you might try to solve a&#xA;different problem but be working on the same code&lt;/p&gt;&#xA;&lt;/div&gt;&#xA;&lt;div class=&#34;admonition note&#34;&gt;&#xA;&lt;p class=&#34;admonition-title&#34;&gt;Problem Space&lt;/p&gt;&#xA;&lt;p&gt;Kedro solves a lot of these problems, so when making rada in Reman the problem&#xA;space was already more contained and narrow, Forge&amp;rsquo;s problem-space is much more&#xA;vast&lt;/p&gt;&#xA;&lt;/div&gt;&#xA;&lt;p&gt;The one problem I&amp;rsquo;m really fixated on right now is data loading, the problem I&#xA;need to solve is accessing data from a wide-variety of scripts/tools, without a&#xA;framework or standard method/library of accessing data in the first place. I&amp;rsquo;ve&#xA;seen lots of projects on GitHub claim to make data loading easier and I didn&amp;rsquo;t&#xA;quite understand what problem they were solving&amp;hellip; one example to name is &lt;a href=&#34;https://dlthub.com/&#34;&gt;Data&#xA;Load Tool&lt;/a&gt;. I&amp;rsquo;ve seen similar ones I don&amp;rsquo;t have the name&#xA;for right now that claim to make it easy/fast to load data from s3, or ways to&#xA;make s3 and a database both abstract in the user-experience for loading data.&#xA;But I hadn&amp;rsquo;t really understood why these tools existed. I have a lot of&#xA;experience with &lt;a href=&#34;https://dlthub.com/&#34;&gt;Kedro&lt;/a&gt; and their&#xA;&lt;a href=&#34;https://docs.kedro.org/en/latest/catalog-data/introduction/&#34;&gt;DataCatalog&lt;/a&gt;&#xA;which provides a python object over a set of yamlfiles that makes it pretty&#xA;simple to load and save data in a way that isolates I/O from the business&#xA;logic. But what I didn&amp;rsquo;t realize at the time how powerful that catalog was, the&#xA;power of standard patterns and shared libraries. Now that I don&amp;rsquo;t have it&#xA;available to me, I&amp;rsquo;m quite aware of the absence.&lt;/p&gt;&#xA;&lt;p&gt;In my new role something I&amp;rsquo;m realizing is that for all the developers my team&#xA;now supports, there isn&amp;rsquo;t a canonical way to access data. When I was in Reman&#xA;and working with kedro, the DataCatalog was the access pattern and so when I&#xA;was developing a platform I never really had to think about it - it was an&#xA;established pattern that I treated as a constraint and then built processes&#xA;around it. I&amp;rsquo;ve been battling some mental block for weeks on Forge because of&#xA;the lack of that canonical pattern, and as I&amp;rsquo;ve talked with other engineers it&#xA;seems like the baseline assumption is that data is just available on a&#xA;filesystem, but everyone&amp;rsquo;s code loads data in different ways. On my small Reman&#xA;team, with common patterns to build on, it was easy to make things cloud-native&#xA;or shim in some devops to improve people&amp;rsquo;s lives. But when everyone&amp;rsquo;s doing&#xA;their own thing, and everyone&amp;rsquo;s &amp;ldquo;own thing&amp;rdquo; is very much built-on some rigid tribal&#xA;patterns then it&amp;rsquo;s hard to really move fast cause everyone isn&amp;rsquo;t already moving&#xA;in the same direction.&lt;/p&gt;&#xA;&lt;p&gt;That made me realize that the first problem Forge needed to solve was in&#xA;providing a way for engineers to have filesystem-native data access in the&#xA;Cloud, where we are S3-first in our storage philosophy. I didn&amp;rsquo;t need to figure&#xA;out a way for everyone to name a dataset, define the dataset in the first&#xA;place, and give a nice &lt;code&gt;my_dataset.load&lt;/code&gt; that worked in python, bash, cpp, and&#xA;who knows what else&amp;hellip;. I reframed the problem from &amp;ldquo;how do engineers load up&#xA;the data&amp;rdquo; to &amp;ldquo;how do engineers have access to the data&amp;rdquo;. The requirements of&#xA;the Cat Autonomy group was pretty simple: POSIX-compliant storage.&lt;/p&gt;&#xA;&lt;p&gt;My pathway to solving this problem is initially underway, I can&amp;rsquo;t imagine it&amp;rsquo;ll&#xA;be too difficult to setup for FSx instances for teams and give them an api to&#xA;run a Batch Job with the FSx mounted. From there, their code can load data from&#xA;&lt;code&gt;/mnt/fsx/&amp;lt;whatever&amp;gt;&lt;/code&gt; just like they otherwise could be doing locally. Or maybe&#xA;FSx will let us setup mounts to very flexible mount points and their local&#xA;scripts will &amp;ldquo;just work&amp;rdquo; :shrugs:. I don&amp;rsquo;t know the exact shape, but after&#xA;realizing the loading data is a big deal, I&amp;rsquo;m thankful I have a narrower&#xA;problem to solve first.&lt;/p&gt;&#xA;&lt;div class=&#34;admonition note&#34;&gt;&#xA;&lt;p class=&#34;admonition-title&#34;&gt;S3 Files&lt;/p&gt;&#xA;&lt;p&gt;Literally yesterday, AWS launched &amp;ldquo;S3 Files&amp;rdquo; offering an NFS filesystem service over buckets. I&amp;rsquo;m not sure if NFS is going to be a viable filesystem protocol for all of our use cases, but looks like we&amp;rsquo;re not the only people who need the filesystem access patterns over S3.&lt;/p&gt;&#xA;&lt;/div&gt;&#xA;&lt;div class=&#34;admonition warning&#34;&gt;&#xA;&lt;p class=&#34;admonition-title&#34;&gt;A Future Problem - Canonical Reference&lt;/p&gt;&#xA;&lt;p&gt;Another high-value thing Forge needs to solve is &amp;ldquo;what is data&amp;rdquo;. The data&#xA;formats we have are not super simple, it&amp;rsquo;s not just a set of SQL tables. We&#xA;have files that relate to each other based on hard-filepath patterns, and those&#xA;patterns are full of tribal knowledge and distributed processes. So a simple&#xA;question like &amp;ldquo;How do I use forge to access my data&amp;rdquo; is hard. In Reman a data&#xA;scientist would ask &amp;ldquo;how do I use rada to access my data?&amp;rdquo; and the answer is&#xA;&amp;ldquo;We use Kedro, and Kedro solved that problem for us via the Catalog&amp;rdquo; but&#xA;without Kedro, without 100% being in python (devs are also in embedded systems,&#xA;cpp code, and more), without even consistent practices in the existing &amp;ldquo;how do&#xA;I access my data&amp;rdquo; workflows, it&amp;rsquo;s really impossible to systemetize and codify&#xA;it. It is my next challenge to tackle though&amp;hellip;&lt;/p&gt;&#xA;&lt;/div&gt;&#xA;</content>
    <author>
      <name>Nic Payne</name>
      <uri>https://pype.dev</uri>
    </author>
  </entry>
</feed>