<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Data Structures on</title><link>/tags/data-structures/</link><description>Recent content in Data Structures on</description><generator>Hugo</generator><language>en</language><lastBuildDate>Mon, 16 Sep 2024 00:00:00 +0000</lastBuildDate><atom:link href="/tags/data-structures/index.xml" rel="self" type="application/rss+xml"/><item><title>Lesson 12: Ring Buffers — Fixed-size queues for real-time systems</title><link>/post/fundamentals/ds-ring-buffers/</link><pubDate>Mon, 16 Sep 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/ds-ring-buffers/</guid><description>&lt;p&gt;The ring buffer is the data structure that makes real-time systems possible. Audio processing, network packet capture, kernel I/O buffers, metrics collection — anywhere you have a producer and a consumer that need to exchange data with zero allocation and bounded latency, you&amp;rsquo;ll find a ring buffer.&lt;/p&gt;
&lt;p&gt;It&amp;rsquo;s also one of the most elegant structures in systems programming: a fixed-size array, two indices, and one invariant. Let me show you how it works and why it appears everywhere from Linux kernel drivers to Disruptor (the LMAX exchange&amp;rsquo;s million-transactions-per-second queue).&lt;/p&gt;</description></item><item><title>Lesson 11: Skip Lists — How Redis sorted sets work</title><link>/post/fundamentals/ds-skip-lists/</link><pubDate>Tue, 03 Sep 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/ds-skip-lists/</guid><description>&lt;p&gt;Redis chose skip lists for its sorted set implementation, and that choice is more interesting than it first appears. When you have &lt;code&gt;ZADD&lt;/code&gt;, &lt;code&gt;ZRANGE&lt;/code&gt;, and &lt;code&gt;ZRANK&lt;/code&gt; all needing to run at O(log n), you might reach for a balanced BST. But Redis chose a probabilistic alternative that&amp;rsquo;s simpler to implement, easier to reason about in concurrent contexts, and performs comparably in practice.&lt;/p&gt;
&lt;p&gt;Understanding skip lists taught me something important about engineering tradeoffs: sometimes &amp;ldquo;good enough with simpler code&amp;rdquo; beats &amp;ldquo;optimal but complex.&amp;rdquo;&lt;/p&gt;</description></item><item><title>Lesson 10: Bloom Filters — Probably yes, definitely no</title><link>/post/fundamentals/ds-bloom-filters/</link><pubDate>Mon, 19 Aug 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/ds-bloom-filters/</guid><description>&lt;p&gt;There&amp;rsquo;s a beautiful data structure that will tell you one of two things: &amp;ldquo;definitely not in the set&amp;rdquo; or &amp;ldquo;probably in the set.&amp;rdquo; That asymmetry — where false negatives are impossible but false positives are allowed — turns out to be useful in an enormous number of production scenarios.&lt;/p&gt;
&lt;p&gt;Bloom filters use a fraction of the memory of a hash set, and the math behind their false positive rate is surprisingly elegant. Once you understand them, you&amp;rsquo;ll see why databases, CDNs, and distributed caches reach for them constantly.&lt;/p&gt;</description></item><item><title>Lesson 9: Tries — Prefix matching, autocomplete, routing tables</title><link>/post/fundamentals/ds-tries/</link><pubDate>Tue, 06 Aug 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/ds-tries/</guid><description>&lt;p&gt;When you type &amp;ldquo;ath&amp;rdquo; into a search box and it suggests &amp;ldquo;atharva,&amp;rdquo; &amp;ldquo;athens,&amp;rdquo; and &amp;ldquo;athletics,&amp;rdquo; that&amp;rsquo;s a trie. When an IP packet arrives at a router and the router decides which interface to forward it to, that&amp;rsquo;s a trie (specifically a Patricia trie). When your web framework matches &lt;code&gt;/api/users/:id&lt;/code&gt; against an incoming URL, the fast implementations use a trie.&lt;/p&gt;
&lt;p&gt;Tries (pronounced &amp;ldquo;try,&amp;rdquo; from &amp;ldquo;re&lt;em&gt;trie&lt;/em&gt;val&amp;rdquo;) are specialized trees for string keys. They trade memory for speed in prefix-matching scenarios, and they make certain string operations fundamentally faster than any other structure.&lt;/p&gt;</description></item><item><title>Lesson 8: Graphs — You're solving graph problems without knowing it</title><link>/post/fundamentals/ds-graphs/</link><pubDate>Tue, 23 Jul 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/ds-graphs/</guid><description>&lt;p&gt;Most engineers don&amp;rsquo;t think of themselves as solving graph problems. They think they&amp;rsquo;re deploying microservices, resolving package dependencies, routing network traffic, or modeling social connections. But the underlying structure in all of these is a graph, and the algorithms that make those systems work — Dijkstra&amp;rsquo;s, topological sort, BFS, DFS — are graph algorithms.&lt;/p&gt;
&lt;p&gt;Once I started seeing graphs everywhere, I started solving systems problems better. Let me show you the representation, the key algorithms, and the production contexts where this thinking pays off.&lt;/p&gt;</description></item><item><title>Lesson 7: Heaps and Priority Queues — Scheduling, top-K, and rate limiters</title><link>/post/fundamentals/ds-heaps/</link><pubDate>Tue, 09 Jul 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/ds-heaps/</guid><description>&lt;p&gt;Every time your operating system picks the next process to run, it&amp;rsquo;s using a heap. Every time Kubernetes re-schedules a pod, it&amp;rsquo;s using a priority queue. Every time you&amp;rsquo;ve written a &amp;ldquo;get top 10 most frequent items from a billion-row stream,&amp;rdquo; the efficient solution involves a heap. These are workhorses of systems programming that look simple on paper and have genuinely tricky implementation details.&lt;/p&gt;
&lt;p&gt;The heap is also one of my favorite data structures to explain because it demonstrates a beautiful property: you can implement a tree efficiently inside an array, using only arithmetic to find parent and child nodes.&lt;/p&gt;</description></item><item><title>Lesson 6: B-Trees and B+ Trees — How every database index actually works</title><link>/post/fundamentals/ds-btrees/</link><pubDate>Wed, 26 Jun 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/ds-btrees/</guid><description>&lt;p&gt;Every time PostgreSQL, MySQL, SQLite, or MongoDB uses an index, it&amp;rsquo;s almost certainly a B-tree or B+ tree underneath. Not a binary search tree — a B-tree. The difference matters more than most engineers realize, and it comes down to one thing: disk read costs.&lt;/p&gt;
&lt;p&gt;I&amp;rsquo;ve seen engineers add indexes blindly to fix slow queries without understanding what an index actually is. When you understand B-trees, you understand why some indexes help more than others, why certain query patterns can&amp;rsquo;t use indexes, and why index-heavy tables slow down on writes.&lt;/p&gt;</description></item><item><title>Lesson 5: Trees and BSTs — Why your database is a tree</title><link>/post/fundamentals/ds-trees-bst/</link><pubDate>Thu, 13 Jun 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/ds-trees-bst/</guid><description>&lt;p&gt;When you run &lt;code&gt;SELECT * FROM orders WHERE user_id = 42&lt;/code&gt;, your database doesn&amp;rsquo;t scan every row. It walks a tree. Understanding why databases chose trees over hash maps — and which kind of tree, and why — is one of those &amp;ldquo;oh, everything makes sense now&amp;rdquo; moments that changes how you design systems.&lt;/p&gt;
&lt;p&gt;Let me start with binary search trees, get honest about their weaknesses, and set up the foundation for B-trees in the next lesson.&lt;/p&gt;</description></item><item><title>Lesson 4: Stacks and Queues — The structures hiding in every system</title><link>/post/fundamentals/ds-stacks-queues/</link><pubDate>Tue, 28 May 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/ds-stacks-queues/</guid><description>&lt;p&gt;Every time you make a function call, your program uses a stack. Every HTTP request in a web server sits in a queue. Every undo operation in your IDE is a stack. These structures are so fundamental they&amp;rsquo;re baked into the hardware itself — the CPU has dedicated stack instructions.&lt;/p&gt;
&lt;p&gt;Yet I consistently see engineers reach for generic slices or channels when a properly implemented stack or queue would be both clearer and faster. Let me show you what these structures actually are, why they&amp;rsquo;re the shape they are, and where they show up in the systems you work on every day.&lt;/p&gt;</description></item><item><title>Lesson 3: Hash Maps — O(1) with asterisks</title><link>/post/fundamentals/ds-hash-maps/</link><pubDate>Wed, 15 May 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/ds-hash-maps/</guid><description>&lt;p&gt;Hash maps are the workhorse of production software. Nearly every caching layer, session store, deduplication system, and configuration lookup you&amp;rsquo;ve ever written relies on one. They&amp;rsquo;re fast, they&amp;rsquo;re flexible, and they&amp;rsquo;re genuinely O(1) — with some asterisks that matter enormously in production.&lt;/p&gt;
&lt;p&gt;I want to walk through how they actually work, because the &amp;ldquo;O(1) lookup&amp;rdquo; claim hides a lot of complexity that shows up in the worst possible moments: under load, with adversarial input, or when your hash function is subtly wrong.&lt;/p&gt;</description></item><item><title>Lesson 2: Linked Lists — Almost never the right choice</title><link>/post/fundamentals/ds-linked-lists/</link><pubDate>Sun, 28 Apr 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/ds-linked-lists/</guid><description>&lt;p&gt;Linked lists are the most over-taught data structure in computer science and the most under-used structure in production systems. I&amp;rsquo;ve reviewed hundreds of pull requests across distributed systems codebases, and I can count on one hand the times a linked list was genuinely the right call.&lt;/p&gt;
&lt;p&gt;That said, understanding why linked lists are usually wrong teaches you something profound about how computers actually work. And there are a handful of situations where they&amp;rsquo;re exactly right — and when those situations come up, you need to recognize them fast.&lt;/p&gt;</description></item><item><title>Lesson 1: Arrays and Memory Layout — Cache lines decide your performance</title><link>/post/fundamentals/ds-arrays-memory/</link><pubDate>Mon, 15 Apr 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/ds-arrays-memory/</guid><description>&lt;p&gt;I spent two years writing Go services before I genuinely understood why iterating over a two-dimensional slice in the wrong order could tank my throughput by 5x. It wasn&amp;rsquo;t a bug. It wasn&amp;rsquo;t a bad algorithm. It was cache lines.&lt;/p&gt;
&lt;p&gt;Arrays are the first data structure everyone learns and the last one most engineers actually understand. This is my attempt to fix that — not with theory, but with the reasoning that makes you a better systems engineer.&lt;/p&gt;</description></item></channel></rss>