<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Fundamentals on</title><link>/tags/fundamentals/</link><description>Recent content in Fundamentals on</description><generator>Hugo</generator><language>en</language><lastBuildDate>Sun, 01 Mar 2026 00:00:00 +0000</lastBuildDate><atom:link href="/tags/fundamentals/index.xml" rel="self" type="application/rss+xml"/><item><title>Lesson 40: Mock Interview Strategy — The 45-minute framework for any problem</title><link>/post/fundamentals/interview-strategy/</link><pubDate>Sun, 01 Mar 2026 00:00:00 +0000</pubDate><guid>/post/fundamentals/interview-strategy/</guid><description>&lt;p&gt;I have had interviews where I solved the problem correctly and still got rejected. I have also had interviews where I struggled significantly with the solution but received strong positive feedback. The difference was not the code — it was everything around the code. How I communicated, how I managed time, how I responded when I got stuck, and whether the interviewer felt like they had seen how I actually think.&lt;/p&gt;</description></item><item><title>Lesson 39: Hard Composites — When one pattern isn't enough</title><link>/post/fundamentals/interview-hard-composites/</link><pubDate>Sun, 15 Feb 2026 00:00:00 +0000</pubDate><guid>/post/fundamentals/interview-hard-composites/</guid><description>&lt;p&gt;There is a class of interview problem designed specifically to differentiate senior candidates. These are not problems where knowing one pattern is enough — they require you to recognize that two or three patterns need to compose, figure out the seam between them, and implement the composition cleanly under pressure. I call them hard composites.&lt;/p&gt;
&lt;p&gt;The candidates who struggle here usually know the individual patterns. The gap is the synthesis. They apply binary search but miss that the search space itself requires a merge step. They build the trie but miss that the relationships between words encode a graph that needs topological sort. Practice the composites explicitly, not just their constituent patterns.&lt;/p&gt;</description></item><item><title>Lesson 38: Design Problems — Build it from scratch in 30 minutes</title><link>/post/fundamentals/interview-design/</link><pubDate>Tue, 03 Feb 2026 00:00:00 +0000</pubDate><guid>/post/fundamentals/interview-design/</guid><description>&lt;p&gt;Design problems in coding interviews are different from system design rounds. You are not sketching architecture at a whiteboard — you are implementing a concrete data structure from scratch, live, in 30 to 45 minutes. The interviewer cares about both correctness and your choices of underlying data structures. &amp;ldquo;Just use a map&amp;rdquo; is never a complete answer.&lt;/p&gt;
&lt;p&gt;I find these problems particularly satisfying because the solutions are compact. Once you see the underlying pattern — that almost every caching and feed problem requires a hash map layered on top of an ordered structure — the implementations become variations on a theme.&lt;/p&gt;</description></item><item><title>Lesson 37: Concurrency Problems — The questions Google loves</title><link>/post/fundamentals/interview-concurrency/</link><pubDate>Sun, 18 Jan 2026 00:00:00 +0000</pubDate><guid>/post/fundamentals/interview-concurrency/</guid><description>&lt;p&gt;Concurrency problems are the ones where an interviewer can tell immediately whether you actually understand concurrency or have just memorized solutions. They are also the problems where Go shines brightest — goroutines and channels are so expressive for synchronization that solutions which require dense mutex orchestration in Java become almost self-documenting in Go.&lt;/p&gt;
&lt;p&gt;Google, in particular, loves these. I have heard this pattern described by multiple engineers who have been through their interview loops: &amp;ldquo;expect at least one problem where you need to coordinate goroutines.&amp;rdquo; The underlying skill being tested is not just &amp;ldquo;can you prevent a race condition&amp;rdquo; — it is &amp;ldquo;do you understand which primitives to reach for, and can you reason about your solution&amp;rsquo;s correctness?&amp;rdquo;&lt;/p&gt;</description></item><item><title>Lesson 36: Intervals — Sort by start, merge by end</title><link>/post/fundamentals/interview-intervals/</link><pubDate>Mon, 05 Jan 2026 00:00:00 +0000</pubDate><guid>/post/fundamentals/interview-intervals/</guid><description>&lt;p&gt;Interval problems show up everywhere — in calendar APIs, in resource scheduling, in genomics, in timeline visualizations. In interviews they appear in a small number of canonical forms, and every form reduces to the same underlying operation: sort by start time, then sweep left to right making decisions based on where the current interval&amp;rsquo;s end overlaps with the next interval&amp;rsquo;s start.&lt;/p&gt;
&lt;p&gt;I have seen engineers panic at interval problems because the cases feel fiddly. Overlapping but not containing. Contained entirely. Adjacent but not touching. Once you drill the sort-and-sweep template into muscle memory, you handle all the cases in a single pass without tracking them explicitly.&lt;/p&gt;</description></item><item><title>Lesson 35: Greedy — Local optimal leads to global optimal (sometimes)</title><link>/post/fundamentals/interview-greedy/</link><pubDate>Tue, 23 Dec 2025 00:00:00 +0000</pubDate><guid>/post/fundamentals/interview-greedy/</guid><description>&lt;p&gt;Greedy algorithms have a brutal failure mode in interviews: the solution looks obvious, you implement it in fifteen minutes, and then the interviewer asks &amp;ldquo;does this always work?&amp;rdquo; and you have no good answer. I have been on both sides of that question. Greedy is powerful when the problem has the right structure, and dangerously wrong when it does not.&lt;/p&gt;
&lt;p&gt;The discipline is learning to tell the difference. Most greedy interview problems are structured so that a correct greedy choice exists — the challenge is identifying what that choice is and, if asked, arguing why locally optimal decisions accumulate to a globally optimal result.&lt;/p&gt;</description></item><item><title>Lesson 34: Backtracking — Try everything, undo what doesn't work</title><link>/post/fundamentals/interview-backtracking/</link><pubDate>Thu, 11 Dec 2025 00:00:00 +0000</pubDate><guid>/post/fundamentals/interview-backtracking/</guid><description>&lt;p&gt;Backtracking scared me for a long time. The problems looked like they needed some clever mathematical insight — some observation that would magically reduce the search space. Then I realized the actual technique is almost mechanical: build a candidate solution incrementally, check constraints at each step, and undo your last choice if the current path cannot lead anywhere valid. That&amp;rsquo;s it. The art is in recognizing when to prune.&lt;/p&gt;
&lt;p&gt;Every backtracking solution I have ever written follows the same skeleton. Once that skeleton is internalized, the remaining work is problem-specific constraint checking. The code almost writes itself.&lt;/p&gt;</description></item><item><title>Lesson 33: Trie — When you need prefix matching</title><link>/post/fundamentals/interview-trie/</link><pubDate>Wed, 26 Nov 2025 00:00:00 +0000</pubDate><guid>/post/fundamentals/interview-trie/</guid><description>&lt;p&gt;Tries showed up in a Google interview I did early in my career. The problem was autocomplete. I started sketching a hash map of prefixes to word lists and immediately knew something was wrong — the interviewer was watching too patiently. The correct structure, the one that makes the solution feel inevitable, is a trie. It organizes words so that every prefix lookup is just a traversal of shared nodes, and I had been fighting to reconstruct that structure from scratch.&lt;/p&gt;</description></item><item><title>Lesson 32: Heap Patterns — Keep the top K without sorting everything</title><link>/post/fundamentals/interview-heap/</link><pubDate>Tue, 11 Nov 2025 00:00:00 +0000</pubDate><guid>/post/fundamentals/interview-heap/</guid><description>&lt;p&gt;The first time I saw a &amp;ldquo;find the K largest elements&amp;rdquo; problem in an interview, I sorted the array and returned the last K. Correct answer, wrong approach. Sorting costs O(n log n). A heap does it in O(n log K). When n is a billion and K is ten, that difference matters enormously — and the interviewer knows it.&lt;/p&gt;
&lt;p&gt;Heaps feel mystical until you internalize one thing: a heap is not a sorted array. It is a partially ordered tree that guarantees one thing — you can get the minimum (or maximum) element in O(1) and remove it in O(log n). That partial ordering is enough to solve an entire class of problems that would otherwise require full sorting.&lt;/p&gt;</description></item><item><title>Lesson 31: Monotonic Stack — The next greater element trick</title><link>/post/fundamentals/interview-monotonic-stack/</link><pubDate>Thu, 30 Oct 2025 00:00:00 +0000</pubDate><guid>/post/fundamentals/interview-monotonic-stack/</guid><description>&lt;p&gt;I used to brute-force &amp;ldquo;next greater element&amp;rdquo; problems. Nested loops, O(n²), and a silent prayer that the input was small. Then a senior engineer at a Google mock interview drew me a picture of a stack where elements got popped the moment something bigger walked in, and the pattern clicked instantly. The monotonic stack is one of those techniques that, once you see it, you wonder how you ever missed it.&lt;/p&gt;</description></item><item><title>Lesson 30: Bitmask DP — When the state is a set</title><link>/post/fundamentals/interview-dp-bitmask/</link><pubDate>Fri, 17 Oct 2025 00:00:00 +0000</pubDate><guid>/post/fundamentals/interview-dp-bitmask/</guid><description>&lt;p&gt;Every DP pattern we&amp;rsquo;ve covered has kept the state manageable: an index, a remaining budget, a mode. Bitmask DP enters the picture when the state is a &lt;em&gt;subset&lt;/em&gt; of elements — specifically, which elements from a small set have been included or visited so far.&lt;/p&gt;
&lt;p&gt;The representation is elegant: a bitmask of n bits, where bit i is 1 if element i is in the current subset and 0 otherwise. With n = 20, there are 2²⁰ ≈ 1 million possible subsets. That&amp;rsquo;s the practical ceiling for bitmask DP — you&amp;rsquo;ll see n ≤ 20 in problem constraints, and that&amp;rsquo;s the signal.&lt;/p&gt;</description></item><item><title>Lesson 29: State Machine DP — Track what state you're in</title><link>/post/fundamentals/interview-dp-state-machine/</link><pubDate>Sat, 04 Oct 2025 00:00:00 +0000</pubDate><guid>/post/fundamentals/interview-dp-state-machine/</guid><description>&lt;p&gt;Most DP problems have a clean one-dimensional state: position in an array, remaining capacity, current index. State machine DP adds another dimension that isn&amp;rsquo;t just a number — it&amp;rsquo;s a &lt;em&gt;mode&lt;/em&gt; or &lt;em&gt;status&lt;/em&gt; that your system is in. &amp;ldquo;Am I currently holding a stock? Am I in a cooldown period? Which color did I just paint the last house?&amp;rdquo;&lt;/p&gt;
&lt;p&gt;The stock trading problems are the canonical example. LeetCode has an entire series of them (I, II, III, IV, with Cooldown, with Transaction Fee) that all share the same structure but add constraints one by one. Solving them all with the same mental framework is satisfying once the state machine clicks.&lt;/p&gt;</description></item><item><title>Lesson 28: Interval DP — Optimal strategy between boundaries</title><link>/post/fundamentals/interview-dp-intervals/</link><pubDate>Fri, 19 Sep 2025 00:00:00 +0000</pubDate><guid>/post/fundamentals/interview-dp-intervals/</guid><description>&lt;p&gt;Interval DP is the pattern that solves problems where you need to find the optimal way to process a contiguous segment, and the answer depends on how you choose to &amp;ldquo;split&amp;rdquo; or &amp;ldquo;last-process&amp;rdquo; within that segment. The classic examples — Burst Balloons, Matrix Chain Multiplication, Optimal BST — all share the same skeleton: try every possible &amp;ldquo;last operation&amp;rdquo; position k within [i, j], and combine the subproblems for [i, k] and [k, j].&lt;/p&gt;</description></item><item><title>Lesson 27: Knapsack Patterns — Pick or skip, that's the whole pattern</title><link>/post/fundamentals/interview-dp-knapsack/</link><pubDate>Fri, 05 Sep 2025 00:00:00 +0000</pubDate><guid>/post/fundamentals/interview-dp-knapsack/</guid><description>&lt;p&gt;Knapsack is one of those patterns that shows up in disguise constantly. You&amp;rsquo;ll see problems framed as &amp;ldquo;partition this array,&amp;rdquo; &amp;ldquo;find a subset with sum X,&amp;rdquo; &amp;ldquo;assign +/- signs to get target T&amp;rdquo; — and underneath each one is the same fundamental structure: for each item, decide whether to include it or exclude it.&lt;/p&gt;
&lt;p&gt;The standard 0/1 Knapsack is the reference problem. Every variant is a modification of it: different objective functions (count instead of max value), different constraints (exact sum instead of capacity), different item usage rules (unbounded instead of once). Once you internalize the base pattern, the variants fall into place.&lt;/p&gt;</description></item><item><title>Lesson 26: DP on Trees — Post-order traversal meets memoization</title><link>/post/fundamentals/interview-dp-trees/</link><pubDate>Wed, 20 Aug 2025 00:00:00 +0000</pubDate><guid>/post/fundamentals/interview-dp-trees/</guid><description>&lt;p&gt;Tree DP is the pattern that catches people by surprise. You&amp;rsquo;ve been thinking of DP as filling a 1D or 2D table from left to right — a sequential, iterative process. Trees are recursive by nature. The &amp;ldquo;table&amp;rdquo; is implicit in the call stack.&lt;/p&gt;
&lt;p&gt;The key insight: tree DP is just post-order traversal where each node computes its answer from its children&amp;rsquo;s answers. There&amp;rsquo;s no explicit table. The memoization (if you need it) is keyed by node pointer. The bottom-up order is inherently satisfied because post-order visits children before parents.&lt;/p&gt;</description></item><item><title>Lesson 25: DP on Strings — Palindromes and partitions</title><link>/post/fundamentals/interview-dp-strings/</link><pubDate>Sat, 09 Aug 2025 00:00:00 +0000</pubDate><guid>/post/fundamentals/interview-dp-strings/</guid><description>&lt;p&gt;String DP has its own flavor that&amp;rsquo;s distinct from the two-string comparison problems of the last lesson. Here, the state is typically a single string, but the subproblem is about a &lt;em&gt;range&lt;/em&gt; within that string: &amp;ldquo;what&amp;rsquo;s the answer for the substring s[i..j]?&amp;rdquo; That&amp;rsquo;s interval DP applied to strings, and palindromes are its most natural setting.&lt;/p&gt;
&lt;p&gt;The tricky part with palindrome problems is that you can approach them from the outside in (is s[i..j] a palindrome?) or the inside out (expand from a center). DP works from the outside in: small intervals first, then build up to larger ones. The expansion approach is often faster in practice but DP is more general and extends to partition problems naturally.&lt;/p&gt;</description></item><item><title>Lesson 24: 2D DP Advanced — String comparison is always 2D DP</title><link>/post/fundamentals/interview-dp-2d-advanced/</link><pubDate>Thu, 24 Jul 2025 00:00:00 +0000</pubDate><guid>/post/fundamentals/interview-dp-2d-advanced/</guid><description>&lt;p&gt;Edit Distance from the previous lesson is the template for a whole family of problems. Any time you&amp;rsquo;re comparing two strings — finding common parts, matching patterns, counting transformations — you&amp;rsquo;re drawing a 2D table where one string indexes the rows and the other indexes the columns.&lt;/p&gt;
&lt;p&gt;The structure is always the same: &lt;code&gt;dp[i][j]&lt;/code&gt; answers a question about the first i characters of one string and the first j characters of the other. The recurrence depends on whether the current characters match and what operations are allowed.&lt;/p&gt;</description></item><item><title>Lesson 23: 2D DP Basics — Two dimensions, one table</title><link>/post/fundamentals/interview-dp-2d-basics/</link><pubDate>Sun, 13 Jul 2025 00:00:00 +0000</pubDate><guid>/post/fundamentals/interview-dp-2d-basics/</guid><description>&lt;p&gt;1D DP was a single row — you filled it left to right and you were done. 2D DP extends that to a full table. You fill it row by row, and each cell depends on cells above it, to its left, or diagonally above-left. The shape of that dependency is the shape of the problem.&lt;/p&gt;
&lt;p&gt;I find 2D DP more intuitive than 1D once I got comfortable with the table visualization. Unique Paths is a perfect first problem because you can literally draw the grid and fill it in by hand in 30 seconds. Edit Distance is the crown jewel — once you understand why the three transitions correspond to the three edit operations, you&amp;rsquo;ll never forget the recurrence.&lt;/p&gt;</description></item><item><title>Lesson 22: 1D DP Advanced — When the state space gets interesting</title><link>/post/fundamentals/interview-dp-1d-advanced/</link><pubDate>Fri, 27 Jun 2025 00:00:00 +0000</pubDate><guid>/post/fundamentals/interview-dp-1d-advanced/</guid><description>&lt;p&gt;Climbing Stairs and House Robber have a comfortable property: &lt;code&gt;dp[i]&lt;/code&gt; depends on just the previous one or two positions. Each step you take looks back a fixed distance. These are the training wheels problems.&lt;/p&gt;
&lt;p&gt;Advanced 1D DP breaks that comfort. Word Break needs to look back up to the entire string. LIS needs to look back at every prior element. Coin Change loops over all denominations at every position. The look-back window is variable, sometimes unbounded. The dp array is still 1D — but you need a loop inside a loop.&lt;/p&gt;</description></item><item><title>Lesson 21: 1D DP Basics — If you can solve it recursively, you can DP it</title><link>/post/fundamentals/interview-dp-1d-basics/</link><pubDate>Fri, 06 Jun 2025 00:00:00 +0000</pubDate><guid>/post/fundamentals/interview-dp-1d-basics/</guid><description>&lt;p&gt;Every time I bombed a DP problem in a mock interview, the pattern was the same: I stared at the problem, thought &amp;ldquo;this looks like DP,&amp;rdquo; then froze because I couldn&amp;rsquo;t immediately write the recurrence. The fix wasn&amp;rsquo;t to memorize more recurrences. The fix was to stop trying to think bottom-up first.&lt;/p&gt;
&lt;p&gt;Here&amp;rsquo;s the approach that finally clicked: write the recursive solution first. Get it working. Then ask: &amp;ldquo;which subproblems am I solving multiple times?&amp;rdquo; Memoize those. Then, if you want to be clean about it, flip it into a bottom-up table. By the time you hit the bottom-up version, the recurrence is already obvious because you derived it from your own recursive code.&lt;/p&gt;</description></item><item><title>Interview Patterns L20: Shortest Path — When edges have weights</title><link>/post/fundamentals/interview-shortest-path/</link><pubDate>Sat, 17 May 2025 00:00:00 +0000</pubDate><guid>/post/fundamentals/interview-shortest-path/</guid><description>&lt;p&gt;In L16, I explained why BFS solves shortest path problems in unweighted graphs — every edge costs the same, so distance equals hop count, and BFS naturally explores in order of increasing hops. But the moment edges have different weights, BFS breaks. A path with two heavy edges can be longer than a path with ten light ones.&lt;/p&gt;
&lt;p&gt;This is the lesson where we graduate to weighted shortest path. Two algorithms matter most for interviews: Dijkstra&amp;rsquo;s (greedy, non-negative weights) and Bellman-Ford (dynamic programming, handles negative weights). A third problem shows a modified Dijkstra on a 2D grid. Understanding when to use each — and why the other would be wrong — is what separates candidates who have memorized code from candidates who actually understand the algorithms.&lt;/p&gt;</description></item><item><title>Interview Patterns L19: Union Find — Who belongs to whom?</title><link>/post/fundamentals/interview-union-find/</link><pubDate>Wed, 30 Apr 2025 00:00:00 +0000</pubDate><guid>/post/fundamentals/interview-union-find/</guid><description>&lt;p&gt;Union Find is one of those data structures that most people never implement until they need it in an interview, and then they discover it is both elegant and surprisingly short. The core idea is deceptively simple: maintain a &amp;ldquo;parent&amp;rdquo; array where each element points to its group&amp;rsquo;s representative. Two operations — &lt;code&gt;Find&lt;/code&gt; (who is the root of this group?) and &lt;code&gt;Union&lt;/code&gt; (merge two groups) — are all you need.&lt;/p&gt;</description></item><item><title>Interview Patterns L18: Topological Sort — Order tasks with dependencies</title><link>/post/fundamentals/interview-topo-sort/</link><pubDate>Mon, 14 Apr 2025 00:00:00 +0000</pubDate><guid>/post/fundamentals/interview-topo-sort/</guid><description>&lt;p&gt;Topological sort comes up whenever there is an ordering constraint — &amp;ldquo;A must happen before B,&amp;rdquo; &amp;ldquo;module X depends on module Y,&amp;rdquo; &amp;ldquo;this course is a prerequisite for that one.&amp;rdquo; The algorithm answers: given a directed acyclic graph (DAG) of dependencies, produce a linear ordering of all nodes such that every node appears after all its predecessors.&lt;/p&gt;
&lt;p&gt;The word &amp;ldquo;topological&amp;rdquo; makes it sound academic. In practice, it is one of the most industrially relevant algorithms: build systems, package managers, spreadsheet recalculation engines, compiler dependency resolution, task schedulers — they all use topological sort or something equivalent. Interviewers ask it because it tests whether you can model real-world dependency problems as graphs and then solve them correctly.&lt;/p&gt;</description></item><item><title>Interview Patterns L17: Graph DFS — Explore everything, mark what you've seen</title><link>/post/fundamentals/interview-graph-dfs/</link><pubDate>Mon, 24 Mar 2025 00:00:00 +0000</pubDate><guid>/post/fundamentals/interview-graph-dfs/</guid><description>&lt;p&gt;Graph DFS is the tool you reach for when you need to explore every reachable node, not just the closest ones. Unlike BFS, which expands in rings of increasing distance, DFS commits to one direction until it cannot go further, then backtracks. This makes it naturally suited for problems where you need to visit entire connected regions, detect cycles, or trace paths between specific source and destination sets.&lt;/p&gt;
&lt;p&gt;The three problems in this lesson each probe a different dimension of graph DFS. Cloning a graph tests your ability to handle shared references while building a new structure. Pacific Atlantic Water Flow introduces the reverse-DFS technique — instead of asking &amp;ldquo;where can this water flow?&amp;rdquo;, you ask &amp;ldquo;which cells can reach each ocean?&amp;rdquo; Course Schedule uses DFS to detect cycles, the canonical prerequisite for topological ordering. Together they cover the non-trivial uses of graph DFS that come up in senior engineering interviews.&lt;/p&gt;</description></item><item><title>Lesson 5: Debugging in Kubernetes — kubectl tricks that save hours</title><link>/post/fundamentals/k8s-debugging/</link><pubDate>Tue, 18 Mar 2025 00:00:00 +0000</pubDate><guid>/post/fundamentals/k8s-debugging/</guid><description>&lt;p&gt;The most stressful hours of my on-call life have been in Kubernetes clusters where something was wrong but nothing was obviously broken. Pods in &lt;code&gt;Pending&lt;/code&gt; for reasons that weren&amp;rsquo;t clear. Services returning 503s but all pods showing as &lt;code&gt;Running&lt;/code&gt;. Memory usage climbing slowly across a fleet for two days before things started dying. Over years of debugging production Kubernetes issues, I&amp;rsquo;ve accumulated a set of commands and mental models that cut through the noise. This article is that collection.&lt;/p&gt;</description></item><item><title>Interview Patterns L16: Graph BFS — Shortest path in unweighted graphs</title><link>/post/fundamentals/interview-graph-bfs/</link><pubDate>Fri, 07 Mar 2025 00:00:00 +0000</pubDate><guid>/post/fundamentals/interview-graph-bfs/</guid><description>&lt;p&gt;The moment someone says &amp;ldquo;shortest path&amp;rdquo; in an interview, I mentally split the problem into two cases: weighted or unweighted? If unweighted — every edge costs the same, every step is distance 1 — BFS gives the shortest path guarantee for free. No Dijkstra needed. BFS explores nodes in order of increasing distance from the source, so the first time you reach a destination, that is definitionally the shortest path.&lt;/p&gt;</description></item><item><title>Interview Patterns L15: Tree BFS Patterns — Level by level reveals structure</title><link>/post/fundamentals/interview-tree-bfs/</link><pubDate>Fri, 21 Feb 2025 00:00:00 +0000</pubDate><guid>/post/fundamentals/interview-tree-bfs/</guid><description>&lt;p&gt;Level by level. That phrase shows up in maybe a third of binary tree interview problems, sometimes explicitly and sometimes disguised. &amp;ldquo;What would you see from the right side?&amp;rdquo; Level by level. &amp;ldquo;What is the minimum number of steps from root to a leaf?&amp;rdquo; Level by level. &amp;ldquo;Connect each node to its right neighbor on the same level?&amp;rdquo; Level by level.&lt;/p&gt;
&lt;p&gt;BFS on trees feels simpler than BFS on graphs because trees have no cycles and no visited-set bookkeeping. But the interesting problems are not about the traversal itself — they are about what you do with the level structure that BFS naturally exposes. In this lesson, I focus on three problems that are each a variation on one theme: level order traversal plus one clever twist per problem.&lt;/p&gt;</description></item><item><title>Interview Patterns L14: Tree DFS Patterns — Every path question is DFS</title><link>/post/fundamentals/interview-tree-dfs/</link><pubDate>Thu, 30 Jan 2025 00:00:00 +0000</pubDate><guid>/post/fundamentals/interview-tree-dfs/</guid><description>&lt;p&gt;If a binary tree problem asks anything about paths — longest, shortest, sum along a path, common ancestor between two nodes — the solution is almost certainly DFS. Not because DFS is the only way, but because path problems require you to propagate information up from leaves to ancestors, and that is exactly what post-order DFS does. You compute the answer for children before combining it for the parent.&lt;/p&gt;
&lt;p&gt;What I find interesting about this cluster of problems is how they reveal a single reusable DFS skeleton. Once you internalize the pattern — recurse left, recurse right, combine and return — you can adapt it to maximum depth, path sum, diameter, and LCA with only surface-level changes. The thinking cost drops dramatically once you stop treating each problem as novel and start asking &amp;ldquo;what do I return from each recursive call, and how do I combine it?&amp;rdquo;&lt;/p&gt;</description></item><item><title>Interview Patterns L13: Tree Construction — Build trees from traversal arrays</title><link>/post/fundamentals/interview-tree-construction/</link><pubDate>Thu, 16 Jan 2025 00:00:00 +0000</pubDate><guid>/post/fundamentals/interview-tree-construction/</guid><description>&lt;p&gt;Construction problems break the mold. Most tree problems ask you to read a tree and return something — a traversal, a value, a boolean. Construction problems ask you to create a tree from some compact representation. That reversal in direction exposes a different set of thinking skills: can you reason backwards from output to structure? Can you identify what information each traversal encoding uniquely determines?&lt;/p&gt;
&lt;p&gt;I think of tree construction as a test of how well you understand what information each traversal preserves and destroys. Inorder alone cannot reconstruct a tree. Preorder alone cannot either. But together they contain exactly enough information. Serialize/deserialize is a different angle: design your own encoding so reconstruction is unambiguous. Both problems show up at Google, Meta, and Amazon with surprising regularity.&lt;/p&gt;</description></item><item><title>Lesson 4: Operators and CRDs — Extending Kubernetes with your own resources</title><link>/post/fundamentals/k8s-operators/</link><pubDate>Fri, 03 Jan 2025 00:00:00 +0000</pubDate><guid>/post/fundamentals/k8s-operators/</guid><description>&lt;p&gt;I used to manage PostgreSQL on Kubernetes with a collection of shell scripts and Helm hooks. Provisioning a new database instance meant running a script that created a StatefulSet, a Service, a ConfigMap, a Secret, and set up replication. Failover meant manually triggering another script. Backups were a cron job that required careful coordination with the stateful set. Every operational task was a script that a human had to run, remember, and maintain. Then I discovered the Zalando Postgres Operator, and I understood what the Operator pattern is actually for: encoding human operational knowledge into the Kubernetes control plane itself.&lt;/p&gt;</description></item><item><title>Interview Patterns L12: BST Operations — Sorted order hides in every BST</title><link>/post/fundamentals/interview-bst/</link><pubDate>Sun, 29 Dec 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/interview-bst/</guid><description>&lt;p&gt;BST problems are deceptively simple on the surface. The property is easy to state: every node&amp;rsquo;s left subtree contains only values less than it, every right subtree contains only values greater. You have known this since your data structures course. But FAANG interviewers do not ask you to recite the definition — they probe the edge cases, the constraints that cascade from parent to child rather than just between a node and its immediate children, and the elegant iterator pattern that wraps inorder traversal in an on-demand API.&lt;/p&gt;</description></item><item><title>Lesson 4: Code Generation — From AST to bytecode or machine code</title><link>/post/fundamentals/compiler-codegen/</link><pubDate>Fri, 20 Dec 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/compiler-codegen/</guid><description>&lt;p&gt;The tree-walking interpreter in Lesson 3 works, but it has a ceiling. Every time you evaluate an expression, you traverse the AST from scratch. No caching, no precomputed form, no optimization. For a REPL or a small scripting language this is fine. For a production language runtime — something that runs for hours, executes millions of operations — you want a tighter inner loop. That tighter loop is a virtual machine executing bytecode. And the step that produces bytecode from the AST is code generation.&lt;/p&gt;</description></item><item><title>Interview Patterns L11: Binary Tree Traversal — Four ways to walk a tree, four different answers</title><link>/post/fundamentals/interview-tree-traversal/</link><pubDate>Tue, 17 Dec 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/interview-tree-traversal/</guid><description>&lt;p&gt;The moment an interviewer draws a binary tree on the whiteboard, a clock starts. They are not just testing whether you know what inorder means. They are watching how you think about state, iteration, and the relationship between recursive and iterative logic. Four traversal orders — preorder, inorder, postorder, level order — each reveals something different about the tree. Knowing which one to reach for, and being able to implement it iteratively without hesitation, is what separates candidates who get offers from candidates who get &amp;ldquo;we&amp;rsquo;ll be in touch.&amp;rdquo;&lt;/p&gt;</description></item><item><title>Lesson 4: Recommendation Systems — Collaborative filtering, embeddings, and the cold start problem</title><link>/post/fundamentals/ml-recommendations/</link><pubDate>Fri, 13 Dec 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/ml-recommendations/</guid><description>&lt;p&gt;Recommendation systems are the most economically consequential ML systems most engineers will ever build. Netflix estimates that its recommendation system saves over a billion dollars per year in avoided cancellations. Amazon&amp;rsquo;s &amp;ldquo;customers who bought this also bought&amp;rdquo; drives a significant fraction of its revenue. Spotify&amp;rsquo;s Discover Weekly has become a user retention flywheel. The stakes are real, and the engineering is genuinely interesting.&lt;/p&gt;
&lt;p&gt;I got deep into recommendation systems when I was working on a content platform that had about 500,000 items and needed to surface relevant content for each user without overwhelming the ranking team with engineering requests. What I found was that the architecture of a production recommender is almost always the same shape, regardless of the domain — and the hardest problem is not the ML, it&amp;rsquo;s the cold start.&lt;/p&gt;</description></item><item><title>Lesson 3: GraphQL in Go — gqlgen, resolvers, and DataLoader for N+1</title><link>/post/fundamentals/graphql-go/</link><pubDate>Sat, 07 Dec 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/graphql-go/</guid><description>&lt;p&gt;Go is not the most common language for GraphQL tutorials. Most of the ecosystem documentation assumes you&amp;rsquo;re working in JavaScript or TypeScript, and a lot of the tooling is designed around Node&amp;rsquo;s event loop model. But Go is an excellent choice for a GraphQL server — statically typed, fast, and with gqlgen you get one of the cleanest schema-first GraphQL implementations I&amp;rsquo;ve used in any language.&lt;/p&gt;
&lt;p&gt;This lesson is about making it work in production: generating a working server from a schema, implementing resolvers correctly, and fixing the N+1 problem with DataLoader before it shows up in your latency graphs.&lt;/p&gt;</description></item><item><title>Lesson 10: Math Patterns — When the Answer Is Math, Not Code</title><link>/post/fundamentals/interview-math/</link><pubDate>Mon, 25 Nov 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/interview-math/</guid><description>&lt;p&gt;Some interview problems look like they need a data structure or a clever algorithm, but the answer is actually just math. I&amp;rsquo;ve watched candidates build hash maps and simulate processes for problems that have elegant O(1) or O(log n) solutions grounded in number theory or geometric reasoning. When I encountered Happy Number for the first time, I tried to detect cycles with a hash set — which works, but the Floyd&amp;rsquo;s cycle detection approach (same one as linked list cycle detection) is what separates a good answer from a great one.&lt;/p&gt;</description></item><item><title>Lesson 3: CQRS + Event Sourcing Together — When the combination makes sense</title><link>/post/fundamentals/es-cqrs-together/</link><pubDate>Sun, 24 Nov 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/es-cqrs-together/</guid><description>&lt;p&gt;CQRS and event sourcing are often discussed together, presented as a single package, and that conflation is responsible for a lot of unnecessary complexity in codebases that should have stayed simple. I&amp;rsquo;ve seen teams adopt both because they read a blog post, without being clear on why they needed either. I&amp;rsquo;ve also seen teams who should have used both and instead built elaborate workarounds that were essentially CQRS and event sourcing with worse ergonomics. The question isn&amp;rsquo;t &amp;ldquo;should I use CQRS and event sourcing?&amp;rdquo; — it&amp;rsquo;s &amp;ldquo;do I have the specific problems these patterns solve?&amp;rdquo;&lt;/p&gt;</description></item><item><title>Lesson 15: CAP Theorem in Practice — What It Actually Means for Your System</title><link>/post/fundamentals/sd-cap-theorem/</link><pubDate>Tue, 19 Nov 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/sd-cap-theorem/</guid><description>&lt;p&gt;CAP theorem is probably the most cited and most misunderstood concept in distributed systems interviews. Candidates memorize &amp;ldquo;you can only pick two of consistency, availability, and partition tolerance&amp;rdquo; and then either over-apply it (treating every design decision as a CAP trade-off) or under-apply it (never relating it to actual design choices). The theorem is real and important, but the way it&amp;rsquo;s usually taught in 30-second summaries strips out the nuance that makes it actually useful. This final lesson clears that up.&lt;/p&gt;</description></item><item><title>Lesson 5: Design Twitter/X — Tweet fanout, timeline ranking, trending topics at 500M users</title><link>/post/fundamentals/sd-deep-twitter/</link><pubDate>Thu, 14 Nov 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/sd-deep-twitter/</guid><description>&lt;p&gt;Twitter is the classic system design problem for good reason. It looks like a glorified blog until you start pulling on the threads: how does a tweet from a user with 100 million followers appear in every follower&amp;rsquo;s timeline within seconds? How do you rank timelines without reading millions of tweets per request? How do you identify trending topics across 500 million users in near real-time? Each of these is a genuinely hard problem, and they interact in non-obvious ways.&lt;/p&gt;</description></item><item><title>Lesson 3: Your Story — Connecting your career narrative to the role</title><link>/post/fundamentals/behavioral-story/</link><pubDate>Sat, 09 Nov 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/behavioral-story/</guid><description>&lt;p&gt;Almost every interview starts the same way: &amp;ldquo;So, tell me about yourself.&amp;rdquo; And almost every engineer I&amp;rsquo;ve talked to hates this question. Not because they don&amp;rsquo;t know themselves, but because it feels formless. You could say anything. You could say everything. What are they actually asking?&lt;/p&gt;
&lt;p&gt;What they&amp;rsquo;re asking is: give me a thesis statement for why you&amp;rsquo;re sitting in this chair. Not a resume recitation. Not a biography. A thread that connects where you&amp;rsquo;ve been to why this job is the logical next step. That thread is your career narrative.&lt;/p&gt;</description></item><item><title>Lesson 9: Bit Manipulation — The Trick Questions That Test Fundamentals</title><link>/post/fundamentals/interview-bit-manipulation/</link><pubDate>Thu, 07 Nov 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/interview-bit-manipulation/</guid><description>&lt;p&gt;Bit manipulation problems have a reputation for being tricks — the kind of problem where you either know the one-liner or you don&amp;rsquo;t, and if you don&amp;rsquo;t, no amount of reasoning will get you there. That&amp;rsquo;s mostly false. The problems in this lesson have elegant solutions, but those solutions come from understanding a small set of bitwise properties and applying them deliberately. I&amp;rsquo;ve seen candidates ace these problems in interviews not because they memorized the answer, but because they reasoned through the bit properties in real time.&lt;/p&gt;</description></item><item><title>Lesson 14: Designing for Failure — Circuit Breakers, Bulkheads, Chaos</title><link>/post/fundamentals/sd-designing-failure/</link><pubDate>Mon, 04 Nov 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/sd-designing-failure/</guid><description>&lt;p&gt;In distributed systems, failure is not an exceptional case — it&amp;rsquo;s the default condition. Networks partition. Hard drives fail. Memory fills up. Dependencies have bugs. Every distributed system that&amp;rsquo;s been running for more than a few years has experienced every kind of failure you can imagine, and many you can&amp;rsquo;t. The engineers who build resilient systems aren&amp;rsquo;t smarter than the ones who don&amp;rsquo;t. They&amp;rsquo;ve just internalized a single principle: design for when things go wrong, not for when things go right.&lt;/p&gt;</description></item><item><title>Lesson 14: Randomized Algorithms — Reservoir sampling, HyperLogLog, probabilistic counting</title><link>/post/fundamentals/algo-randomized/</link><pubDate>Fri, 01 Nov 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/algo-randomized/</guid><description>&lt;p&gt;There is a class of production problems where exact answers are either impossible or not worth the cost. You want to know approximately how many unique visitors hit your site today. You want to sample 1% of requests for tracing without reading every request into memory first. You want to check if a username is already taken without querying the database on every keystroke.&lt;/p&gt;
&lt;p&gt;Randomized algorithms provide exact answers with known probability bounds, or approximate answers with bounded error, using a fraction of the memory or time that exact computation would require. I was skeptical of &amp;ldquo;probabilistic&amp;rdquo; algorithms for a long time — it seemed like trading correctness for efficiency. Then I learned what the error bounds actually are. A HyperLogLog cardinality estimate with 1.5% error using 12KB of memory is a better engineering choice than a perfect count using 100MB, for the vast majority of use cases.&lt;/p&gt;</description></item><item><title>Lesson 13: Design a Payment System — Idempotency, Reconciliation, Double-Entry</title><link>/post/fundamentals/sd-payment-system/</link><pubDate>Mon, 21 Oct 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/sd-payment-system/</guid><description>&lt;p&gt;Payment systems have a property that almost no other software has: the cost of a bug isn&amp;rsquo;t a bad user experience — it&amp;rsquo;s a legal liability and a business catastrophe. Charging a customer twice, losing a transfer in a network failure, or crediting the wrong account can result in millions of dollars of loss and destroyed trust. Every other system we&amp;rsquo;ve covered tolerates a degree of eventual inconsistency. Payment systems, in most cases, do not. This lesson is about building for that level of correctness.&lt;/p&gt;</description></item><item><title>Lesson 13: Cryptographic Primitives — Hashing, HMAC, and never rolling your own</title><link>/post/fundamentals/algo-crypto/</link><pubDate>Thu, 17 Oct 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/algo-crypto/</guid><description>&lt;p&gt;There is a joke in security circles: every developer thinks they can write their own crypto. The punchline is that everyone who has tried has been wrong. Cryptography is the one area of computer science where being 99% correct is the same as being completely wrong. A subtle timing vulnerability, a nonce reuse, or a hash function with the wrong properties can completely destroy a security guarantee that looks solid on paper.&lt;/p&gt;</description></item><item><title>Lesson 3: ConfigMaps and Secrets — Configuration without rebuilding</title><link>/post/fundamentals/k8s-config-secrets/</link><pubDate>Wed, 16 Oct 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/k8s-config-secrets/</guid><description>&lt;p&gt;We had a bug where the same Docker image was running in staging and production but behaving differently. I spent an hour diffing the code before I checked the environment variables. The staging pod had &lt;code&gt;CACHE_TTL=60&lt;/code&gt; and the production pod had &lt;code&gt;CACHE_TTL=3600&lt;/code&gt;. The image was identical. The behavior was totally different. That kind of environment-specific configuration — the values that shouldn&amp;rsquo;t be baked into the image — is exactly what ConfigMaps and Secrets are for. But they have subtleties that bite you if you treat them as simple key-value stores.&lt;/p&gt;</description></item><item><title>Lesson 3: Model Serving — Latency, batching, A/B testing in production</title><link>/post/fundamentals/ml-model-serving/</link><pubDate>Mon, 14 Oct 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/ml-model-serving/</guid><description>&lt;p&gt;Training a model is satisfying. Deploying it to serve real traffic is humbling. I remember the first time I pushed a model to production and watched the p99 latency hover at 800ms on what was supposed to be a &amp;ldquo;fast&amp;rdquo; model. The benchmark had shown 12ms inference time. What happened? The benchmark ran the model on pre-loaded batches; production served one request at a time, loaded the model fresh on cold starts, and had no GPU batching. The gap between &amp;ldquo;the model is accurate&amp;rdquo; and &amp;ldquo;the model is fast enough to be useful&amp;rdquo; is where model serving engineering lives.&lt;/p&gt;</description></item><item><title>Lesson 2: Server-Side WASM — WASI, edge computing, and the universal binary</title><link>/post/fundamentals/wasm-server-side/</link><pubDate>Fri, 11 Oct 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/wasm-server-side/</guid><description>&lt;p&gt;When most people talk about WebAssembly, they mean the browser. That&amp;rsquo;s where it started, that&amp;rsquo;s where the tutorials are, and that&amp;rsquo;s where most of the public discourse still lives. But the more interesting story for backend engineers is what happens when you take the same sandboxed, portable binary model and apply it on the server.&lt;/p&gt;
&lt;p&gt;Server-side WebAssembly — specifically WebAssembly with WASI (the WebAssembly System Interface) — is not a toy. Cloudflare Workers runs WASM. Fastly Compute runs WASM. Fermyon Spin is built on it. The pattern is spreading from edge providers into general-purpose infrastructure. Understanding it now puts you ahead of where most backend engineers are.&lt;/p&gt;</description></item><item><title>Lesson 8: Sorting in Interviews — Can You Do Better Than O(n log n)?</title><link>/post/fundamentals/interview-sorting/</link><pubDate>Wed, 09 Oct 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/interview-sorting/</guid><description>&lt;p&gt;Most candidates treat sorting as a black box — call &lt;code&gt;sort.Ints&lt;/code&gt; and move on. That works until an interviewer says &amp;ldquo;can you do this without sorting?&amp;rdquo; or &amp;ldquo;implement this yourself&amp;rdquo; or &amp;ldquo;what&amp;rsquo;s the worst-case time?&amp;rdquo; At that point, not understanding sorting costs you the offer.&lt;/p&gt;
&lt;p&gt;I don&amp;rsquo;t mean you need to memorize every sorting algorithm. You need three things: understand merge sort well enough to implement it (divide-and-conquer is a pattern you&amp;rsquo;ll use in other contexts); know QuickSelect for kth-element problems (it&amp;rsquo;s O(n) average and comes up constantly); and know when sorting is all you need and when the comparison lower bound of O(n log n) is a limit you can break. This lesson teaches all three.&lt;/p&gt;</description></item><item><title>Lesson 12: Design a Search Engine — Inverted Index and Ranking</title><link>/post/fundamentals/sd-search-engine/</link><pubDate>Mon, 07 Oct 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/sd-search-engine/</guid><description>&lt;p&gt;Search is the feature that separates usable products from unusable ones at scale. When your application has thousands of documents, a sequential scan works. At millions, it doesn&amp;rsquo;t. The difference between a search box that works and one that times out is a data structure invented in the 1960s that every modern search engine still fundamentally relies on: the inverted index. Understanding how it works — and how to build a system around it — is what this lesson is about.&lt;/p&gt;</description></item><item><title>Lesson 3: Paxos and Beyond — When Raft isn't enough</title><link>/post/fundamentals/consensus-paxos/</link><pubDate>Fri, 04 Oct 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/consensus-paxos/</guid><description>&lt;p&gt;After spending time with Raft, I found myself curious about the algorithm it was designed to replace. Paxos has a reputation: brilliant, correct, nearly impossible to implement correctly, and even harder to extend to practical systems. Leslie Lamport published the original Paxos paper in 1989, got it rejected, submitted a revised version in 1998, and it became the theoretical foundation for a generation of distributed systems. Chubby (Google&amp;rsquo;s distributed lock service), Zookeeper (the coordination service), and the precursor to Spanner all descended from Paxos thinking. Understanding why Raft was necessary requires understanding what Paxos gets right and where it falls short in practice.&lt;/p&gt;</description></item><item><title>Lesson 12: Compression Basics — Why gzip works and entropy matters</title><link>/post/fundamentals/algo-compression/</link><pubDate>Thu, 03 Oct 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/algo-compression/</guid><description>&lt;p&gt;Every senior engineer I know makes compression decisions regularly: should this API response be gzip-compressed? Should logs be stored compressed? What compression level? Which algorithm? I made these decisions for years based on &amp;ldquo;gzip is standard, use it&amp;rdquo; without understanding why it worked or when something else might be better.&lt;/p&gt;
&lt;p&gt;Understanding the fundamentals of how compression works — entropy, Huffman coding, and LZ77 — changed how I think about data formats, wire protocols, and storage costs. It also helped me understand why some data compresses well and some data does not, which is critical for capacity planning.&lt;/p&gt;</description></item><item><title>Lesson 3: AST and Evaluation — Walking the tree to compute results</title><link>/post/fundamentals/compiler-ast-eval/</link><pubDate>Thu, 03 Oct 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/compiler-ast-eval/</guid><description>&lt;p&gt;After the lexer and parser, we have a tree. A beautiful, hierarchical, unambiguous representation of the program. Now we need to make it &lt;em&gt;do something&lt;/em&gt;. The simplest way to execute a program from its AST is to walk the tree recursively and compute results as you go. No intermediate representation, no bytecode, no machine code — just a recursive function that pattern-matches on node types and returns values.&lt;/p&gt;
&lt;p&gt;This is a tree-walking interpreter. It is not the fastest approach (we cover bytecode and code generation in Lesson 4), but it is the most direct. Many production language implementations started here — and some, like early Ruby and Python, stayed here for a long time.&lt;/p&gt;</description></item><item><title>Lesson 11: Design a Notification System — Push vs Pull, Priority, Dedup</title><link>/post/fundamentals/sd-notifications/</link><pubDate>Sat, 21 Sep 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/sd-notifications/</guid><description>&lt;p&gt;Notifications are the feature that can make or break user retention — and also destroy it. Done right, they bring users back at exactly the right moment. Done wrong, they&amp;rsquo;re spam that drives uninstalls. The system design challenge isn&amp;rsquo;t just the technical plumbing (though that&amp;rsquo;s interesting). It&amp;rsquo;s building infrastructure that&amp;rsquo;s fast for critical alerts, reliable for important messages, and smart enough to not overwhelm users with low-priority noise.&lt;/p&gt;
&lt;h2 id="the-core-concept"&gt;The Core Concept&lt;/h2&gt;
&lt;p&gt;A notification system has to handle multiple channels with wildly different characteristics:&lt;/p&gt;</description></item><item><title>Lesson 11: String Algorithms — KMP, Rabin-Karp, and why regex can be slow</title><link>/post/fundamentals/algo-string/</link><pubDate>Tue, 17 Sep 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/algo-string/</guid><description>&lt;p&gt;String processing sits at the foundation of almost every production system. Log parsing, protocol parsing, search, validation, templating — it all comes down to finding and transforming patterns in text. Most of the time, strings.Contains or a simple loop is fast enough. But when you are processing millions of log lines per second, or running user-supplied patterns against untrusted input, or implementing a search feature that needs to handle long documents, naive string matching becomes a bottleneck or a security hole.&lt;/p&gt;</description></item><item><title>Lesson 12: Ring Buffers — Fixed-size queues for real-time systems</title><link>/post/fundamentals/ds-ring-buffers/</link><pubDate>Mon, 16 Sep 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/ds-ring-buffers/</guid><description>&lt;p&gt;The ring buffer is the data structure that makes real-time systems possible. Audio processing, network packet capture, kernel I/O buffers, metrics collection — anywhere you have a producer and a consumer that need to exchange data with zero allocation and bounded latency, you&amp;rsquo;ll find a ring buffer.&lt;/p&gt;
&lt;p&gt;It&amp;rsquo;s also one of the most elegant structures in systems programming: a fixed-size array, two indices, and one invariant. Let me show you how it works and why it appears everywhere from Linux kernel drivers to Disruptor (the LMAX exchange&amp;rsquo;s million-transactions-per-second queue).&lt;/p&gt;</description></item><item><title>Lesson 4: Design WhatsApp — End-to-end encryption, message delivery guarantees, presence</title><link>/post/fundamentals/sd-deep-whatsapp/</link><pubDate>Sat, 14 Sep 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/sd-deep-whatsapp/</guid><description>&lt;p&gt;WhatsApp is deceptively simple from a user perspective: you send a message, it arrives. But building a messaging system that handles 100 billion messages per day with end-to-end encryption, reliable delivery semantics, and real-time presence for 2 billion users is a genuinely hard engineering problem. I find this problem particularly instructive because it forces you to confront three things simultaneously: cryptographic key management, message delivery guarantees, and the cost of maintaining online/offline state at massive scale.&lt;/p&gt;</description></item><item><title>Lesson 2: Projections and Read Models — Build any view from your event stream</title><link>/post/fundamentals/es-projections/</link><pubDate>Fri, 13 Sep 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/es-projections/</guid><description>&lt;p&gt;After I shipped the event-sourced account system, the first question from the product team was: &amp;ldquo;Can we have a page that shows all accounts that have been dormant for more than 90 days?&amp;rdquo; In a traditional system, this is a query: &lt;code&gt;SELECT * FROM accounts WHERE last_activity &amp;lt; NOW() - INTERVAL '90 days'&lt;/code&gt;. In event sourcing, there&amp;rsquo;s no &lt;code&gt;last_activity&lt;/code&gt; column — there&amp;rsquo;s an event stream. My first instinct was to query the event store directly, find the latest event per stream, and filter. It worked. Then they asked for the top 100 accounts by balance. Then active accounts by geography. Then a real-time dashboard with all of the above simultaneously. Querying the event store for each of these is either very slow or very complex. The answer is projections — pre-computed read models built from your event streams.&lt;/p&gt;</description></item><item><title>Lesson 10: Vacuum and Bloat — Why Postgres Tables Grow</title><link>/post/fundamentals/db-vacuum-bloat/</link><pubDate>Mon, 09 Sep 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/db-vacuum-bloat/</guid><description>&lt;p&gt;I watched a Postgres instance run out of disk space on a 500 GB SSD. The database had 50 GB of actual data. The other 450 GB was table bloat — dead row versions from MVCC that vacuum had failed to clean up. A batch job had been running long-running transactions for weeks, holding a transaction horizon that prevented vacuum from reclaiming anything. By the time we noticed, the disk was nearly full and autovacuum was struggling to catch up. Understanding why this happens — and how to prevent it — is one of the most important operational skills for running Postgres in production.&lt;/p&gt;</description></item><item><title>Lesson 2: Schema Design — Types, queries, mutations, and subscriptions</title><link>/post/fundamentals/graphql-schema/</link><pubDate>Sun, 08 Sep 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/graphql-schema/</guid><description>&lt;p&gt;The schema is the most important artifact in any GraphQL API. It is simultaneously the contract between your client and server, the documentation for every engineer who works with the API, and the boundary that forces you to think clearly about your domain before writing any implementation code. A well-designed schema makes everything easier. A poorly designed schema compounds every mistake downstream.&lt;/p&gt;
&lt;p&gt;I&amp;rsquo;ve seen both. The experience of inheriting a badly designed GraphQL schema — full of inconsistent naming, misused types, nullable fields everywhere because someone wasn&amp;rsquo;t sure — is one of the more persistent forms of technical debt I&amp;rsquo;ve encountered. Unlike a poorly written function, a bad schema is public-facing. You can&amp;rsquo;t just refactor it; you have to version and deprecate carefully.&lt;/p&gt;</description></item><item><title>Lesson 7: Recursion — Trust the Recursion, Define the Base Case</title><link>/post/fundamentals/interview-recursion/</link><pubDate>Sat, 07 Sep 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/interview-recursion/</guid><description>&lt;p&gt;Recursion trips people up not because the concept is hard, but because they try to trace through it. &amp;ldquo;If I call &lt;code&gt;f(3)&lt;/code&gt;, then &lt;code&gt;f(3)&lt;/code&gt; calls &lt;code&gt;f(2)&lt;/code&gt;, which calls &lt;code&gt;f(1)&lt;/code&gt;&amp;hellip;&amp;rdquo; and then they lose the thread. I used to do this. My interviewer at a Series B startup once watched me spend three minutes trying to mentally simulate a recursive call stack for a problem that had a two-line solution once I stopped simulating and started trusting.&lt;/p&gt;</description></item><item><title>Lesson 10: Design a News Feed — Fan-out on Write vs Read</title><link>/post/fundamentals/sd-news-feed/</link><pubDate>Fri, 06 Sep 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/sd-news-feed/</guid><description>&lt;p&gt;The news feed is the product feature that defines social media. Every time you open Instagram or Twitter, you see a personalized, ranked, real-time stream of content from people you follow. Behind that deceptively simple UI is one of the hardest distributed systems problems in consumer tech: how do you compute a personalized feed for hundreds of millions of users, where any piece of content needs to appear in potentially millions of feeds, within seconds of being posted?&lt;/p&gt;</description></item><item><title>Lesson 11: Skip Lists — How Redis sorted sets work</title><link>/post/fundamentals/ds-skip-lists/</link><pubDate>Tue, 03 Sep 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/ds-skip-lists/</guid><description>&lt;p&gt;Redis chose skip lists for its sorted set implementation, and that choice is more interesting than it first appears. When you have &lt;code&gt;ZADD&lt;/code&gt;, &lt;code&gt;ZRANGE&lt;/code&gt;, and &lt;code&gt;ZRANK&lt;/code&gt; all needing to run at O(log n), you might reach for a balanced BST. But Redis chose a probabilistic alternative that&amp;rsquo;s simpler to implement, easier to reason about in concurrent contexts, and performs comparably in practice.&lt;/p&gt;
&lt;p&gt;Understanding skip lists taught me something important about engineering tradeoffs: sometimes &amp;ldquo;good enough with simpler code&amp;rdquo; beats &amp;ldquo;optimal but complex.&amp;rdquo;&lt;/p&gt;</description></item><item><title>Lesson 10: Backtracking — Constraint satisfaction and config generation</title><link>/post/fundamentals/algo-backtracking/</link><pubDate>Mon, 02 Sep 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/algo-backtracking/</guid><description>&lt;p&gt;Backtracking is the algorithmic equivalent of &amp;ldquo;try everything, but be smart about giving up early.&amp;rdquo; It is a systematic way to explore a search space where you build a solution incrementally and abandon partial solutions as soon as you detect they cannot possibly lead to a valid result.&lt;/p&gt;
&lt;p&gt;I first encountered backtracking outside of textbooks when building a scheduling system that needed to assign employees to shifts while respecting availability constraints, skill requirements, and labor regulations. The brute force approach — enumerate all possible assignments — was computationally equivalent to exploring all permutations, which was 20! for 20 employees. Backtracking with constraint pruning reduced the search space by several orders of magnitude.&lt;/p&gt;</description></item><item><title>Lesson 9: Partitioning — Range, Hash, List and When Each Helps</title><link>/post/fundamentals/db-partitioning/</link><pubDate>Sat, 24 Aug 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/db-partitioning/</guid><description>&lt;p&gt;I have worked with a few systems that had grown their main tables to 500 million rows or more. At that scale, even well-indexed queries start slowing down — not because the indexes are wrong, but because the index itself becomes large and the buffer cache can only hold so many pages. Vacuum struggles to keep up. &lt;code&gt;EXPLAIN&lt;/code&gt; output looks fine but queries still feel sluggish. Partitioning is the architectural solution to this class of problem: instead of one big table, you have many smaller tables that look like one from the application&amp;rsquo;s perspective.&lt;/p&gt;</description></item><item><title>Lesson 8: Technical Debt — When to pay, when to live with it</title><link>/post/fundamentals/arch-tech-debt/</link><pubDate>Fri, 23 Aug 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/arch-tech-debt/</guid><description>&lt;p&gt;The term &amp;ldquo;technical debt&amp;rdquo; gets used to mean everything from &amp;ldquo;this code is a mess&amp;rdquo; to &amp;ldquo;we made a pragmatic shortcut we need to revisit&amp;rdquo; to &amp;ldquo;this system has grown organically and nobody understands it anymore.&amp;rdquo; These are different problems requiring different responses. Treating all technical debt as something that must be paid now, or all of it as something acceptable to defer indefinitely — both lead to bad outcomes. The skill is knowing which debt to address, when, and how.&lt;/p&gt;</description></item><item><title>Lesson 9: Design a Chat System — WebSocket, Presence, Message Ordering</title><link>/post/fundamentals/sd-chat-system/</link><pubDate>Wed, 21 Aug 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/sd-chat-system/</guid><description>&lt;p&gt;Building a chat system is the interview problem that catches people on protocol fundamentals. HTTP is a request-response protocol — the client asks, the server answers, and then the connection is idle. For chat, the server needs to push messages to clients the moment they arrive. This inversion of the HTTP model is what makes chat hard, and it&amp;rsquo;s what drives the entire architecture.&lt;/p&gt;
&lt;h2 id="the-core-concept"&gt;The Core Concept&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Why HTTP polling doesn&amp;rsquo;t work at scale&lt;/strong&gt;&lt;/p&gt;</description></item><item><title>Lesson 10: Bloom Filters — Probably yes, definitely no</title><link>/post/fundamentals/ds-bloom-filters/</link><pubDate>Mon, 19 Aug 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/ds-bloom-filters/</guid><description>&lt;p&gt;There&amp;rsquo;s a beautiful data structure that will tell you one of two things: &amp;ldquo;definitely not in the set&amp;rdquo; or &amp;ldquo;probably in the set.&amp;rdquo; That asymmetry — where false negatives are impossible but false positives are allowed — turns out to be useful in an enormous number of production scenarios.&lt;/p&gt;
&lt;p&gt;Bloom filters use a fraction of the memory of a hash set, and the math behind their false positive rate is surprisingly elegant. Once you understand them, you&amp;rsquo;ll see why databases, CDNs, and distributed caches reach for them constantly.&lt;/p&gt;</description></item><item><title>Lesson 9: Greedy Algorithms — When being selfish is optimal</title><link>/post/fundamentals/algo-greedy/</link><pubDate>Sat, 17 Aug 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/algo-greedy/</guid><description>&lt;p&gt;Greedy algorithms have a simple idea at their core: at each step, make the locally optimal choice. No looking ahead, no considering alternatives, no backtracking. Just take the best available option right now and trust that it leads to a globally optimal result.&lt;/p&gt;
&lt;p&gt;The catch is that this only works for specific problem structures. Use greedy when the problem has the &lt;strong&gt;greedy choice property&lt;/strong&gt; — the locally optimal choice is always part of a globally optimal solution. When this holds, greedy is elegant and fast. When it does not, greedy gives you a wrong answer with no warning.&lt;/p&gt;</description></item><item><title>Lesson 2: Leadership and Conflict — The questions that decide senior vs mid-level</title><link>/post/fundamentals/behavioral-leadership/</link><pubDate>Wed, 14 Aug 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/behavioral-leadership/</guid><description>&lt;p&gt;There&amp;rsquo;s a specific moment in every senior-level interview loop where the conversation shifts. The system design round wraps up, the coding is done, and then an interviewer leans back and asks something like: &amp;ldquo;Tell me about a time you had to push back on a product decision you thought was wrong.&amp;rdquo; Or: &amp;ldquo;Describe a situation where you disagreed with your tech lead and how you handled it.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;These are not softballs. For companies hiring at senior or staff level, behavioral questions about leadership and conflict are often the deciding signal. Everyone who makes it to that round can code. Not everyone can navigate the organizational and interpersonal dynamics of a senior role. These questions are trying to separate the two.&lt;/p&gt;</description></item><item><title>Lesson 8: Memory-Mapped IO — How Databases Read Files</title><link>/post/fundamentals/linux-mmap/</link><pubDate>Tue, 13 Aug 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/linux-mmap/</guid><description>&lt;p&gt;I was reading about how RocksDB works one evening and kept seeing references to &lt;code&gt;mmap&lt;/code&gt; for reads. I had heard of memory-mapped files but thought of them as a niche optimization. Then I realized: SQLite uses mmap. WiredTiger (MongoDB&amp;rsquo;s storage engine) uses mmap. LMDB is built almost entirely around mmap. Even some Postgres configurations use mmap for WAL. Understanding mmap explains a lot about how high-performance storage works, why some databases are so fast for random reads, and why memory and I/O are so deeply intertwined at the OS level.&lt;/p&gt;</description></item><item><title>Lesson 7: On-Call Engineering — Reducing toil, improving reliability</title><link>/post/fundamentals/eng-oncall/</link><pubDate>Sun, 11 Aug 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/eng-oncall/</guid><description>&lt;p&gt;I did 12 months of on-call on a team that hadn&amp;rsquo;t invested in reliability. The rotation was weekly. In a bad week, I&amp;rsquo;d get 15-20 pages. A good week was 5. I was exhausted by the end of my shift, and the paging frequency had barely changed over those 12 months. We were fixing incidents, not fixing the causes. The next team I joined approached on-call differently: on-call was treated as a reliability sensor, not a firefighting rotation. Pages were tracked, patterns identified, and root causes fixed. By month 6 I was averaging 2 pages per week on-call. The work we did during on-call made future on-call better.&lt;/p&gt;</description></item><item><title>Lesson 8: Replication — Streaming, Logical, and Failover</title><link>/post/fundamentals/db-replication/</link><pubDate>Fri, 09 Aug 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/db-replication/</guid><description>&lt;p&gt;The first time I had to fail over a Postgres primary, it was 2 AM, the primary was not responding, and I genuinely did not know whether the replica was up to date or 10 minutes behind. We recovered, but that experience drove me to actually understand replication — not just &amp;ldquo;the replica gets the writes somehow&amp;rdquo; but exactly what data is shipped, when it arrives, and what happens when the primary dies. It turns out the mechanism is elegant and directly connected to the WAL we covered in Lesson 3.&lt;/p&gt;</description></item><item><title>Lesson 2: Deployments and Scaling — Rolling updates, HPA, VPA</title><link>/post/fundamentals/k8s-deployments/</link><pubDate>Wed, 07 Aug 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/k8s-deployments/</guid><description>&lt;p&gt;The week before a major product launch, our Kubernetes cluster started evicting pods. Traffic was spiking as we ran load tests, but instead of scaling up, the HPA was oscillating — scaling up, then the new pods were getting OOMKilled, then scaled down, then up again. Pods were in CrashLoopBackOff, the HPA metrics were lagging behind the actual load, and I was watching health check failures cascade. It turned out our resource requests were wildly inaccurate — set once during initial deployment and never updated as the service&amp;rsquo;s actual usage changed. That day I learned that Deployments and autoscaling aren&amp;rsquo;t fire-and-forget configurations.&lt;/p&gt;</description></item><item><title>Lesson 7: Migration Strategies — Strangler fig and feature flags</title><link>/post/fundamentals/arch-migrations/</link><pubDate>Wed, 07 Aug 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/arch-migrations/</guid><description>&lt;p&gt;Rewrites are seductive. The existing system is messy, it&amp;rsquo;s slow to change, and the new system in your head is clean and fast and well-designed. Then you start the rewrite. Six months in, you&amp;rsquo;ve rebuilt 40% of the functionality and the remaining 60% is more complex than you thought. Meanwhile, the old system keeps shipping features. The new system falls behind, gets cancelled, and you&amp;rsquo;re back where you started, except now you&amp;rsquo;ve lost six months and the team is demoralized. I&amp;rsquo;ve seen this happen twice. The third time, we used the strangler fig pattern instead, and it worked.&lt;/p&gt;</description></item><item><title>Lesson 9: Tries — Prefix matching, autocomplete, routing tables</title><link>/post/fundamentals/ds-tries/</link><pubDate>Tue, 06 Aug 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/ds-tries/</guid><description>&lt;p&gt;When you type &amp;ldquo;ath&amp;rdquo; into a search box and it suggests &amp;ldquo;atharva,&amp;rdquo; &amp;ldquo;athens,&amp;rdquo; and &amp;ldquo;athletics,&amp;rdquo; that&amp;rsquo;s a trie. When an IP packet arrives at a router and the router decides which interface to forward it to, that&amp;rsquo;s a trie (specifically a Patricia trie). When your web framework matches &lt;code&gt;/api/users/:id&lt;/code&gt; against an incoming URL, the fast implementations use a trie.&lt;/p&gt;
&lt;p&gt;Tries (pronounced &amp;ldquo;try,&amp;rdquo; from &amp;ldquo;re&lt;em&gt;trie&lt;/em&gt;val&amp;rdquo;) are specialized trees for string keys. They trade memory for speed in prefix-matching scenarios, and they make certain string operations fundamentally faster than any other structure.&lt;/p&gt;</description></item><item><title>Lesson 8: Design a URL Shortener — The Classic, Done Properly</title><link>/post/fundamentals/sd-url-shortener/</link><pubDate>Sun, 04 Aug 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/sd-url-shortener/</guid><description>&lt;p&gt;The URL shortener is the &amp;ldquo;Hello World&amp;rdquo; of system design interviews. It appears deceptively simple: take a long URL, return a short one. But if you treat it superficially, you miss what the interviewer is actually testing: your ability to think through ID generation at scale, read-heavy caching, redirect semantics, analytics storage, and data modeling. Done well, the URL shortener problem touches nearly every fundamental we&amp;rsquo;ve covered so far.&lt;/p&gt;
&lt;h2 id="the-core-concept"&gt;The Core Concept&lt;/h2&gt;
&lt;p&gt;A URL shortener has two primary operations:&lt;/p&gt;</description></item><item><title>Lesson 7: Service Mesh — Sidecar proxies and mTLS without code</title><link>/post/fundamentals/net-service-mesh/</link><pubDate>Sat, 03 Aug 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/net-service-mesh/</guid><description>&lt;p&gt;A team I worked with had seventeen microservices. Each service had its own implementation of retry logic, circuit breaking, timeout handling, and mutual TLS. Some used libraries, some rolled their own. When we needed to update the TLS certificate rotation policy, it touched eleven different code repositories, four different languages, and took two months. Then we introduced Linkerd and moved all of that to the infrastructure layer. The services still did their jobs. The networking became someone else&amp;rsquo;s problem — specifically, the platform team&amp;rsquo;s.&lt;/p&gt;</description></item><item><title>Lesson 2: Feature Stores — Why feature engineering is 80% of ML</title><link>/post/fundamentals/ml-feature-stores/</link><pubDate>Fri, 02 Aug 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/ml-feature-stores/</guid><description>&lt;p&gt;There is a saying in machine learning that is so universally acknowledged it has become a cliché: 80% of the work in any ML project is feature engineering. I spent a long time thinking this referred to the cognitive labor — the domain expertise required to craft meaningful features. It does, in part. But the deeper meaning is operational. The 80% is not just about &lt;em&gt;what&lt;/em&gt; features to build; it&amp;rsquo;s about &lt;em&gt;how&lt;/em&gt; to compute them consistently, store them efficiently, retrieve them with sub-millisecond latency at serving time, and keep them synchronized between the training pipeline and the production system.&lt;/p&gt;</description></item><item><title>Lesson 8: Dynamic Programming Intuition — Memoization, not memorization</title><link>/post/fundamentals/algo-dynamic-programming/</link><pubDate>Fri, 02 Aug 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/algo-dynamic-programming/</guid><description>&lt;p&gt;Dynamic programming has an intimidating reputation. The name sounds academic, the problems in textbooks often involve sequences and matrices, and the &amp;ldquo;aha moment&amp;rdquo; is notoriously hard to force. I spent a long time treating DP as an interview preparation topic rather than a tool I would actually use. That changed when I built a pricing engine and realized I had been reimplementing DP badly — without knowing it — by computing the same values over and over in nested function calls.&lt;/p&gt;</description></item><item><title>Lesson 7: Containers from Scratch — Namespaces, Cgroups, What Docker Does</title><link>/post/fundamentals/linux-containers/</link><pubDate>Mon, 29 Jul 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/linux-containers/</guid><description>&lt;p&gt;I spent the first two years of my career using Docker without understanding what it was actually doing. I thought containers were a &amp;ldquo;lightweight VM&amp;rdquo; — some kind of virtualization. Then I read Liz Rice&amp;rsquo;s &amp;ldquo;Containers from Scratch&amp;rdquo; talk and wrote a 100-line container runtime in Go. Containers are not virtual machines. They are just Linux processes with restricted views of the system, implemented using two kernel features: namespaces and cgroups. Once you see how simple the underlying mechanism is, you understand exactly what Docker adds and why containers behave the way they do.&lt;/p&gt;</description></item><item><title>Lesson 6: Feature Flags — Progressive rollout and kill switches</title><link>/post/fundamentals/eng-feature-flags/</link><pubDate>Sun, 28 Jul 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/eng-feature-flags/</guid><description>&lt;p&gt;We deployed a new checkout flow on a Friday afternoon. It had passed code review, passed testing, passed staging load tests. By Friday evening, the error rate was climbing. The new flow had a race condition that only manifested under specific mobile browser timing that we hadn&amp;rsquo;t tested. Without a feature flag, the fix would have required an emergency deployment — 25 minutes of build time, deployment, and validation. With a feature flag, the rollback was turning off a switch. Thirty seconds.&lt;/p&gt;</description></item><item><title>Lesson 7: Connection Pooling — What PgBouncer Actually Does</title><link>/post/fundamentals/db-connection-pooling/</link><pubDate>Sat, 27 Jul 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/db-connection-pooling/</guid><description>&lt;p&gt;A Postgres connection is not cheap. When I first scaled a service from a single server to 20 replicas — each running a Go application with &lt;code&gt;database/sql&lt;/code&gt;&amp;rsquo;s default pool settings — the database became the bottleneck almost immediately. Not because the queries were slow. Because we had 20 × 100 = 2,000 open connections, each consuming RAM on the Postgres server, and Postgres was spending more time managing connections than executing queries. This is the problem PgBouncer solves, and understanding how it works makes you a much better architect of backend systems.&lt;/p&gt;</description></item><item><title>Lesson 2: Raft Consensus — The consensus algorithm you can actually understand</title><link>/post/fundamentals/consensus-raft/</link><pubDate>Fri, 26 Jul 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/consensus-raft/</guid><description>&lt;p&gt;When I first tried to understand Paxos — the original distributed consensus algorithm — I read the paper three times and still felt like I was missing something. I could follow each step individually, but I couldn&amp;rsquo;t build a mental model of why it worked or what the invariants were. Raft was designed specifically to fix that. Its paper is literally titled &amp;ldquo;In Search of an Understandability: The Raft Consensus Algorithm.&amp;rdquo; After reading it, I could explain it to someone else. That&amp;rsquo;s the bar Raft was designed to clear, and it does.&lt;/p&gt;</description></item><item><title>Lesson 2: Parsing — Tokens to AST, recursive descent</title><link>/post/fundamentals/compiler-parsing/</link><pubDate>Thu, 25 Jul 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/compiler-parsing/</guid><description>&lt;p&gt;The lexer gave us tokens. The tokens are still flat — a sequence with no hierarchy. The expression &lt;code&gt;1 + 2 * 3&lt;/code&gt; produces six tokens, but it does not yet say that &lt;code&gt;2 * 3&lt;/code&gt; should be computed before adding &lt;code&gt;1&lt;/code&gt;. That grouping, that hierarchy, is what the parser builds. By the time the parser is done, we have a tree. A specific kind of tree: an Abstract Syntax Tree, or AST. Everything after the parser works on this tree: evaluation, type checking, code generation, optimization. The tree is the program.&lt;/p&gt;</description></item><item><title>Lesson 6: API Versioning — URL, header, or content negotiation</title><link>/post/fundamentals/arch-api-versioning/</link><pubDate>Wed, 24 Jul 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/arch-api-versioning/</guid><description>&lt;p&gt;The first time I shipped a breaking change to a production API without a version, I got an incident at 2am because a partner integration stopped working. They&amp;rsquo;d been calling &lt;code&gt;/api/orders&lt;/code&gt; for six months. I changed the response shape. Their code broke. That was when &amp;ldquo;API versioning&amp;rdquo; moved from &amp;ldquo;thing to do eventually&amp;rdquo; to &amp;ldquo;thing to do before you ship anything external.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;There&amp;rsquo;s no universally right answer to how you version APIs. There are tradeoffs, and the choice you make will shape your codebase for years. Here&amp;rsquo;s how I think about it.&lt;/p&gt;</description></item><item><title>Lesson 8: Graphs — You're solving graph problems without knowing it</title><link>/post/fundamentals/ds-graphs/</link><pubDate>Tue, 23 Jul 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/ds-graphs/</guid><description>&lt;p&gt;Most engineers don&amp;rsquo;t think of themselves as solving graph problems. They think they&amp;rsquo;re deploying microservices, resolving package dependencies, routing network traffic, or modeling social connections. But the underlying structure in all of these is a graph, and the algorithms that make those systems work — Dijkstra&amp;rsquo;s, topological sort, BFS, DFS — are graph algorithms.&lt;/p&gt;
&lt;p&gt;Once I started seeing graphs everywhere, I started solving systems problems better. Let me show you the representation, the key algorithms, and the production contexts where this thinking pays off.&lt;/p&gt;</description></item><item><title>Lesson 7: Rate Limiting — Token Bucket, Sliding Window, Distributed</title><link>/post/fundamentals/sd-rate-limiting/</link><pubDate>Sun, 21 Jul 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/sd-rate-limiting/</guid><description>&lt;p&gt;Without rate limiting, a single misbehaving client can consume all your resources and deny service to everyone else. A bug in a client that retries in a tight loop. A competitor scraping your API. A DDoS attack. A feature that accidentally calls your endpoint ten times per button click. Rate limiting is the mechanism that prevents any of these from taking down your service — it&amp;rsquo;s the first line of defense between the internet and your origin.&lt;/p&gt;</description></item><item><title>Lesson 7: Shortest Path — Dijkstra in routing and network optimization</title><link>/post/fundamentals/algo-shortest-path/</link><pubDate>Fri, 19 Jul 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/algo-shortest-path/</guid><description>&lt;p&gt;When I joined a team building a multi-region traffic routing system, the first thing I had to understand was why the routing decisions were not always the geographically shortest path. The system was using Dijkstra&amp;rsquo;s algorithm, but the edge weights were not just latency — they incorporated bandwidth cost, current utilization, failure rates, and SLA constraints. Dijkstra did not care what the weights meant. It just found the minimum cost path. That is its power.&lt;/p&gt;</description></item><item><title>Lesson 6: gRPC and Protobuf — Binary protocols and streaming</title><link>/post/fundamentals/net-grpc/</link><pubDate>Wed, 17 Jul 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/net-grpc/</guid><description>&lt;p&gt;The first JSON API I replaced with gRPC was passing &lt;code&gt;[]Order&lt;/code&gt; objects around — each order had about 40 fields, most of which the caller never used. The JSON payload for a list of 100 orders was around 180KB. After the migration it was 22KB, and the serialization time in benchmarks dropped by 8x. But more than the performance numbers, what I noticed was the schema. Proto files are contracts. When a field changes, you know it. With JSON, you find out when things break in production.&lt;/p&gt;</description></item><item><title>Lesson 3: Design Google Docs — Real-time collaboration, OT vs CRDT, conflict resolution</title><link>/post/fundamentals/sd-deep-google-docs/</link><pubDate>Mon, 15 Jul 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/sd-deep-google-docs/</guid><description>&lt;p&gt;Google Docs is the problem I recommend to every engineer who thinks they understand distributed systems. The surface looks trivial: multiple users editing a document simultaneously. The depth is staggering. When two users type at the same position in a document at the same millisecond, what does each user see? How do you converge on a consistent state without a central lock? How do you preserve the intention behind each edit, not just the characters?&lt;/p&gt;</description></item><item><title>Lesson 6: Signals — SIGTERM vs SIGKILL and Graceful Shutdown</title><link>/post/fundamentals/linux-signals/</link><pubDate>Sun, 14 Jul 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/linux-signals/</guid><description>&lt;p&gt;The first time I deployed a service to Kubernetes and watched it restart, I noticed that requests in flight were sometimes failing with connection resets. The pod was receiving 502 responses for a few seconds before it disappeared. I had heard of &amp;ldquo;graceful shutdown&amp;rdquo; but hadn&amp;rsquo;t implemented it. Kubernetes sends &lt;code&gt;SIGTERM&lt;/code&gt; before killing a process, giving it time to finish in-flight requests — but my service was either ignoring the signal or exiting immediately, cutting off connections mid-request. Learning how signals work at the OS level, and then implementing correct signal handling in Go, fixed the issue permanently.&lt;/p&gt;</description></item><item><title>Lesson 5: Load Testing — k6, vegeta, realistic patterns</title><link>/post/fundamentals/eng-load-testing/</link><pubDate>Sat, 13 Jul 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/eng-load-testing/</guid><description>&lt;p&gt;We load-tested our new checkout service before launch. 1,000 virtual users, 10 minutes, all hitting &lt;code&gt;/v1/checkout&lt;/code&gt; sequentially. It passed with excellent numbers. Launch day: real traffic hit the service, and it fell over at 200 concurrent users. The problem was our test. Real users don&amp;rsquo;t all call the same endpoint in sequence. They browse, add to cart, apply discount codes, fill in addresses, and then checkout — a session that touches 8 different endpoints over 4 minutes. Our test didn&amp;rsquo;t model this. Our test was measuring the wrong thing.&lt;/p&gt;</description></item><item><title>Lesson 1: WASM from Go — Compile Go to WebAssembly and run it anywhere</title><link>/post/fundamentals/wasm-from-go/</link><pubDate>Fri, 12 Jul 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/wasm-from-go/</guid><description>&lt;p&gt;I came to WebAssembly skeptically. The pitch — &amp;ldquo;run code anywhere, at near-native speed, in a sandboxed environment&amp;rdquo; — sounded like the kind of claim that looks great in a conference talk and falls apart in production. It took a specific use case to make me take it seriously: I needed to run the same validation logic in three environments — a Go backend, a JavaScript frontend, and a CLI tool — without maintaining three separate implementations of the same business rules.&lt;/p&gt;</description></item><item><title>Lesson 6: Query Planning and EXPLAIN — Reading Execution Plans</title><link>/post/fundamentals/db-explain/</link><pubDate>Thu, 11 Jul 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/db-explain/</guid><description>&lt;p&gt;The first time I ran &lt;code&gt;EXPLAIN ANALYZE&lt;/code&gt; on a slow query, I stared at the output for five minutes and understood none of it. It looked like a compiler error message crossed with a financial report. Then a senior engineer walked me through it, and what had looked like noise resolved into a clear picture: this node was doing more work than the planner expected, this join strategy was wrong, this sort was spilling to disk. &lt;code&gt;EXPLAIN ANALYZE&lt;/code&gt; is the most powerful tool in database performance work, and it takes an hour to learn but pays back every week for the rest of your career.&lt;/p&gt;</description></item><item><title>Lesson 7: Heaps and Priority Queues — Scheduling, top-K, and rate limiters</title><link>/post/fundamentals/ds-heaps/</link><pubDate>Tue, 09 Jul 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/ds-heaps/</guid><description>&lt;p&gt;Every time your operating system picks the next process to run, it&amp;rsquo;s using a heap. Every time Kubernetes re-schedules a pod, it&amp;rsquo;s using a priority queue. Every time you&amp;rsquo;ve written a &amp;ldquo;get top 10 most frequent items from a billion-row stream,&amp;rdquo; the efficient solution involves a heap. These are workhorses of systems programming that look simple on paper and have genuinely tricky implementation details.&lt;/p&gt;
&lt;p&gt;The heap is also one of my favorite data structures to explain because it demonstrates a beautiful property: you can implement a tree efficiently inside an array, using only arithmetic to find parent and child nodes.&lt;/p&gt;</description></item><item><title>Lesson 5: DDD Essentials — Bounded contexts and aggregates</title><link>/post/fundamentals/arch-ddd/</link><pubDate>Sun, 07 Jul 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/arch-ddd/</guid><description>&lt;p&gt;The concept of &amp;ldquo;customer&amp;rdquo; meant something different in every team I talked to at one company. To the billing team, a customer was a billing account. To the support team, a customer was a person who filed tickets. To the identity team, a customer was an authenticated principal. Every team had their own &lt;code&gt;Customer&lt;/code&gt; struct, and syncing them was a full-time job. That&amp;rsquo;s the problem DDD&amp;rsquo;s bounded context concept solves — and it&amp;rsquo;s one of those ideas that sounds academic until you&amp;rsquo;ve suffered without it.&lt;/p&gt;</description></item><item><title>Lesson 6: CDNs — Put Your Bytes Close to Your Users</title><link>/post/fundamentals/sd-cdns/</link><pubDate>Thu, 04 Jul 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/sd-cdns/</guid><description>&lt;p&gt;Physics is the enemy of performance. Light travels through fiber optic cables at about 200,000 km/s — roughly two-thirds the speed of light in a vacuum. A round-trip from New York to London is about 11,000 km each way. That means the minimum possible latency for that trip is 55ms. You cannot engineer your way past it. What you can do is stop making the trip at all. Content Delivery Networks are the infrastructure that puts popular content at the edge — closer to users, so the bytes never have to travel far.&lt;/p&gt;</description></item><item><title>Lesson 6: BFS and DFS — Dependency resolution, crawlers, cycle detection</title><link>/post/fundamentals/algo-bfs-dfs/</link><pubDate>Wed, 03 Jul 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/algo-bfs-dfs/</guid><description>&lt;p&gt;Graph traversal sounds abstract until you realize that most interesting data in production is a graph. Service dependencies are a graph. Database foreign key relationships form a graph. Build tool dependencies, org charts, permission hierarchies, network topologies — all graphs. BFS and DFS are the two fundamental ways to walk them, and they show up in real engineering work more than almost any other algorithm.&lt;/p&gt;
&lt;p&gt;I have used DFS to detect circular imports in a build system, BFS to find the shortest migration path between two schema versions, and both to debug why a dependency injection container was resolving services in the wrong order.&lt;/p&gt;</description></item><item><title>Lesson 5: WebSockets — Upgrade, framing, vs SSE vs polling</title><link>/post/fundamentals/net-websockets/</link><pubDate>Mon, 01 Jul 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/net-websockets/</guid><description>&lt;p&gt;When I first built a real-time notification system, I reached for WebSockets because that&amp;rsquo;s what everyone said to use for &amp;ldquo;real-time.&amp;rdquo; Three months later I was debugging connection drops under load, wrestling with proxy timeouts, and fighting with nginx configuration. When I stepped back and actually thought about the access pattern — server pushing notifications to the browser, never the browser sending data to the server — I replaced WebSockets with Server-Sent Events in a weekend. The code got simpler, the proxies stopped complaining, and we had less to maintain. Picking the right tool requires understanding what each one actually does.&lt;/p&gt;</description></item><item><title>Lesson 1: Event Store Design — Append-only logs that are your source of truth</title><link>/post/fundamentals/es-event-store/</link><pubDate>Sat, 29 Jun 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/es-event-store/</guid><description>&lt;p&gt;My first encounter with event sourcing was a financial system that needed a complete audit trail. The business requirement was clear: every change to an account balance had to be traceable — who changed it, when, why. The first approach was an audit log table alongside the main accounts table. We&amp;rsquo;d write to both in a transaction. Within six months, the audit table was out of sync with the accounts table. Bugs in the dual-write logic, a migration that updated account records without touching the audit log, a direct database fix that bypassed the application layer. The audit log was nearly useless. The core insight I eventually arrived at: the audit log shouldn&amp;rsquo;t be a secondary record — it should be the primary record. That&amp;rsquo;s event sourcing.&lt;/p&gt;</description></item><item><title>Lesson 5: Epoll and IO Multiplexing — How Go''s Netpoller Works</title><link>/post/fundamentals/linux-epoll/</link><pubDate>Thu, 27 Jun 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/linux-epoll/</guid><description>&lt;p&gt;I used to wonder how a Go HTTP server could handle 100,000 concurrent connections with only 8 OS threads. If each connection required a dedicated thread, you would need 100,000 threads — which would require roughly 800 GB of stack space and would spend all their time in the kernel scheduler. The answer is epoll: a Linux kernel interface that lets a single thread wait on thousands of file descriptors simultaneously and be notified only when one is ready for I/O. Go&amp;rsquo;s runtime uses epoll internally as the foundation of its network I/O model. This lesson explains how.&lt;/p&gt;</description></item><item><title>Lesson 6: B-Trees and B+ Trees — How every database index actually works</title><link>/post/fundamentals/ds-btrees/</link><pubDate>Wed, 26 Jun 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/ds-btrees/</guid><description>&lt;p&gt;Every time PostgreSQL, MySQL, SQLite, or MongoDB uses an index, it&amp;rsquo;s almost certainly a B-tree or B+ tree underneath. Not a binary search tree — a B-tree. The difference matters more than most engineers realize, and it comes down to one thing: disk read costs.&lt;/p&gt;
&lt;p&gt;I&amp;rsquo;ve seen engineers add indexes blindly to fix slow queries without understanding what an index actually is. When you understand B-trees, you understand why some indexes help more than others, why certain query patterns can&amp;rsquo;t use indexes, and why index-heavy tables slow down on writes.&lt;/p&gt;</description></item><item><title>Lesson 4: Monitoring and Alerting — SLOs and alert fatigue</title><link>/post/fundamentals/eng-monitoring/</link><pubDate>Mon, 24 Jun 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/eng-monitoring/</guid><description>&lt;p&gt;The on-call rotation I inherited had 47 active alerts. On a bad week, the on-call engineer got paged 30 times. Most pages were &amp;ldquo;something might be wrong&amp;rdquo; noise — high CPU on one instance, a spike in error rate that self-resolved in 30 seconds, disk space at 70% on a server with months of capacity remaining. Engineers stopped taking the pages seriously. Then the one real incident got buried in the noise, and we had a 4-hour outage because no one treated the first alert seriously. Alert fatigue is not a monitoring problem. It&amp;rsquo;s an architecture-of-trust problem.&lt;/p&gt;</description></item><item><title>Lesson 5: Transaction Isolation — Read Committed vs Serializable</title><link>/post/fundamentals/db-isolation-levels/</link><pubDate>Sun, 23 Jun 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/db-isolation-levels/</guid><description>&lt;p&gt;I shipped a bug once that allowed a user to spend the same gift card balance twice. Two requests arrived nearly simultaneously, both read the same balance, both decided the balance was sufficient, both deducted it, and both succeeded. The database did exactly what I asked. The problem was what I asked for: I assumed reads were consistent across statements within a transaction, but I was running at the default isolation level. Understanding the four isolation levels — and what each one actually protects you from — is not academic. It is the difference between shipping correct financial code and shipping race conditions.&lt;/p&gt;</description></item><item><title>Lesson 4: CQRS — When reads and writes need different models</title><link>/post/fundamentals/arch-cqrs/</link><pubDate>Fri, 21 Jun 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/arch-cqrs/</guid><description>&lt;p&gt;I maintained an e-commerce admin dashboard that had a query so complex it took 12 seconds to run. It joined seven tables across the orders, inventory, and customers domains to build a summary view of &amp;ldquo;all orders pending fulfillment, with customer tier, item details, and warehouse stock levels.&amp;rdquo; Every time the admin loaded the page, twelve seconds. We tried indexes, caching, materialized views — all helped at the margins. The fundamental problem was that we were asking our write-optimized relational model to answer a read-optimized reporting question. Separating the read model was the only real fix.&lt;/p&gt;</description></item><item><title>Lesson 5: Message Queues — Decoupling Services Without Losing Messages</title><link>/post/fundamentals/sd-message-queues/</link><pubDate>Wed, 19 Jun 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/sd-message-queues/</guid><description>&lt;p&gt;Synchronous calls are elegant until they&amp;rsquo;re not. Service A calls Service B. B is slow. A waits. A&amp;rsquo;s request pool fills up. A becomes slow. The caller of A waits. The whole request chain stalls. Add a few more services in the chain and you have a distributed deadlock in slow motion. Message queues exist to break this dependency — to let a producer say &amp;ldquo;here&amp;rsquo;s some work&amp;rdquo; and move on, without caring whether the consumer is fast, slow, or temporarily down.&lt;/p&gt;</description></item><item><title>Lesson 5: Hashing and Consistent Hashing — How load balancers distribute traffic</title><link>/post/fundamentals/algo-hashing/</link><pubDate>Mon, 17 Jun 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/algo-hashing/</guid><description>&lt;p&gt;I did not think much about hashing until I had to explain why a cache cluster was unusable after we added two new nodes. We had doubled the cache hit rate over six months, added two boxes to handle the load, and immediately destroyed most of our cached data. Every key rehashed to a different server. Cache hit rate dropped from 85% to under 10%. We were hammering the database.&lt;/p&gt;</description></item><item><title>Lesson 4: DNS — Resolution, caching, and why changes take time</title><link>/post/fundamentals/net-dns/</link><pubDate>Sat, 15 Jun 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/net-dns/</guid><description>&lt;p&gt;We pushed a production incident fix at 3am — rotated to a new IP address for a critical service, updated the DNS record, set the TTL to 60 seconds. Thirty minutes later, half our users were still hitting the broken server. The other half were fine. We had updated the DNS record correctly. The TTL had expired. Yet somehow, stale answers were persisting. That night taught me more about DNS than any documentation had.&lt;/p&gt;</description></item><item><title>Lesson 5: Trees and BSTs — Why your database is a tree</title><link>/post/fundamentals/ds-trees-bst/</link><pubDate>Thu, 13 Jun 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/ds-trees-bst/</guid><description>&lt;p&gt;When you run &lt;code&gt;SELECT * FROM orders WHERE user_id = 42&lt;/code&gt;, your database doesn&amp;rsquo;t scan every row. It walks a tree. Understanding why databases chose trees over hash maps — and which kind of tree, and why — is one of those &amp;ldquo;oh, everything makes sense now&amp;rdquo; moments that changes how you design systems.&lt;/p&gt;
&lt;p&gt;Let me start with binary search trees, get honest about their weaknesses, and set up the foundation for B-trees in the next lesson.&lt;/p&gt;</description></item><item><title>Lesson 4: TCP/IP Stack — SYN Floods, TIME_WAIT, Connection Tuning</title><link>/post/fundamentals/linux-tcp-ip/</link><pubDate>Tue, 11 Jun 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/linux-tcp-ip/</guid><description>&lt;p&gt;A load test I ran against a Go HTTP server one afternoon produced a strange result: throughput leveled off at about 30,000 requests per second and couldn&amp;rsquo;t go higher, even though CPU was at 30%. Running &lt;code&gt;netstat -an | grep TIME_WAIT | wc -l&lt;/code&gt; showed over 28,000 connections in &lt;code&gt;TIME_WAIT&lt;/code&gt; state. The kernel was running out of local port numbers. Understanding TCP connection states — and how to tune Linux&amp;rsquo;s TCP stack — unblocked the test and taught me more about networking than any course had. This lesson covers the TCP mechanics that actually matter for backend engineers running high-throughput services.&lt;/p&gt;</description></item><item><title>Lesson 4: MVCC — How Postgres Handles Concurrent Reads and Writes</title><link>/post/fundamentals/db-mvcc/</link><pubDate>Sun, 09 Jun 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/db-mvcc/</guid><description>&lt;p&gt;One thing that puzzled me early on was how Postgres could let me read from a table while someone else was writing to it — without locking me out and without me seeing half-written data. Most systems I had worked with used explicit read locks, which meant readers and writers had to take turns. Postgres doesn&amp;rsquo;t do that. Reads never block writes, and writes never block reads. The mechanism that makes this possible is called Multiversion Concurrency Control, or MVCC, and it works by keeping multiple versions of every row simultaneously.&lt;/p&gt;</description></item><item><title>Lesson 3: Incident Response — Postmortems and blameless culture</title><link>/post/fundamentals/eng-incidents/</link><pubDate>Sat, 08 Jun 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/eng-incidents/</guid><description>&lt;p&gt;The first major incident I was on-call for, I spent 90 minutes trying to fix the problem and 30 minutes frantically communicating to stakeholders in a panic. The second one, I followed a runbook and spent the 90 minutes coordinating, communicating clearly, and delegating diagnosis — while the problem was resolved in 40 minutes. The difference wasn&amp;rsquo;t technical skill. It was process. Incident response is a skill you can learn and practice, and it makes a measurable difference in how quickly you restore service and how well the team learns from failures.&lt;/p&gt;</description></item><item><title>Lesson 1: GraphQL vs REST — When GraphQL wins and when it doesn't</title><link>/post/fundamentals/graphql-vs-rest/</link><pubDate>Thu, 06 Jun 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/graphql-vs-rest/</guid><description>&lt;p&gt;I spent about six months being wrong about GraphQL. My first exposure to it was a rewrite pitch from a frontend engineer who was tired of making five REST calls to assemble one screen. &amp;ldquo;GraphQL solves this,&amp;rdquo; they said, and the demo looked compelling. A single query, exactly the fields you need, one round trip. I was sold before thinking carefully about what we were buying.&lt;/p&gt;
&lt;p&gt;The rewrite happened. It took eight months instead of four. Some of the problems we were trying to solve got better. Others got worse. Looking back, the mistake wasn&amp;rsquo;t choosing GraphQL — it was not being rigorous about where GraphQL actually wins and where it doesn&amp;rsquo;t before committing.&lt;/p&gt;</description></item><item><title>Lesson 4: Database Scaling — Read Replicas, Sharding, and When Each Helps</title><link>/post/fundamentals/sd-database-scaling/</link><pubDate>Wed, 05 Jun 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/sd-database-scaling/</guid><description>&lt;p&gt;Your startup ships, gains traction, and one day your database becomes the bottleneck. Queries slow down. CPU spikes. The single machine hosting your Postgres instance can&amp;rsquo;t keep up. This is one of the most predictable problems in engineering, yet I&amp;rsquo;ve seen teams reach it completely unprepared because they&amp;rsquo;d never thought carefully about how databases scale. The right solution depends entirely on whether you&amp;rsquo;re read-heavy or write-heavy, whether your data is relational or not, and whether your workload is even or spiky.&lt;/p&gt;</description></item><item><title>Lesson 3: Event-Driven Architecture — Events vs commands</title><link>/post/fundamentals/arch-event-driven/</link><pubDate>Tue, 04 Jun 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/arch-event-driven/</guid><description>&lt;p&gt;I used to use &amp;ldquo;event&amp;rdquo; and &amp;ldquo;command&amp;rdquo; interchangeably. A message is a message, right? Then I started debugging an event-driven system where the payment service was publishing &lt;code&gt;ProcessPayment&lt;/code&gt; events, the order service was publishing &lt;code&gt;CreateShipment&lt;/code&gt; events, and two teams were arguing about who owned the workflow. The problem was naming: they were publishing commands disguised as events. The distinction isn&amp;rsquo;t pedantic — it determines who owns the workflow and how the system evolves.&lt;/p&gt;</description></item><item><title>Lesson 1: ML Pipelines — From raw data to deployed model</title><link>/post/fundamentals/ml-pipelines/</link><pubDate>Sun, 02 Jun 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/ml-pipelines/</guid><description>&lt;p&gt;The first time I built a &amp;ldquo;production ML pipeline,&amp;rdquo; it was a cron job that ran a Python script, trained a model, and saved the artifact to disk. It worked, until it didn&amp;rsquo;t. The model silently degraded over three weeks because the training data schema changed and nobody noticed. There were no tests, no validation, no monitoring. That experience, embarrassing as it was, taught me more about ML system design than any paper on model architecture.&lt;/p&gt;</description></item><item><title>Lesson 4: Two Pointers and Sliding Window — Stream processing in disguise</title><link>/post/fundamentals/algo-two-pointers/</link><pubDate>Sat, 01 Jun 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/algo-two-pointers/</guid><description>&lt;p&gt;The first time I recognized the sliding window pattern outside a textbook, I was reading the source code for a rate limiter. There was no comment saying &amp;ldquo;sliding window algorithm here.&amp;rdquo; There was just a loop, two indices into a circular buffer, and some simple arithmetic that maintained a count of events in the last N seconds. Once I saw it, I started seeing it everywhere — in network flow control, in moving average calculations, in deduplication logic.&lt;/p&gt;</description></item><item><title>Lesson 6: Linked Lists — Pointer Manipulation Is the Real Test</title><link>/post/fundamentals/interview-linked-lists/</link><pubDate>Fri, 31 May 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/interview-linked-lists/</guid><description>&lt;p&gt;Linked lists are where candidates reveal whether they actually understand pointers. You can read a hundred tutorials on reversing a linked list, but the first time you try to code it under pressure, you will likely lose a node, corrupt the list, or write an infinite loop. I know because it happened to me in a mock interview. The problem itself is not hard. The pointer mechanics are.&lt;/p&gt;
&lt;p&gt;This lesson is about building the mental model that makes pointer operations feel mechanical rather than scary. Meta asks linked list questions constantly — their systems rely heavily on custom allocators and cache implementations, and linked lists are the natural test vehicle. Amazon uses the LRU cache variant to test both data structure design and implementation discipline.&lt;/p&gt;</description></item><item><title>Lesson 3: TLS Handshake — What happens in those 2 round trips</title><link>/post/fundamentals/net-tls/</link><pubDate>Thu, 30 May 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/net-tls/</guid><description>&lt;p&gt;A colleague once asked me why adding TLS to a service increased our P99 latency by 50ms. She had measured it carefully, switching between HTTP and HTTPS in a load test. My first instinct was to say &amp;ldquo;encryption overhead&amp;rdquo; but that&amp;rsquo;s wrong — modern CPUs with AES-NI can encrypt gigabytes per second. The actual cost is the handshake. Once I explained what was actually happening in those first few round trips, the answer to &amp;ldquo;how do we fix it&amp;rdquo; became obvious: stop creating new connections.&lt;/p&gt;</description></item><item><title>Lesson 4: Stacks and Queues — The structures hiding in every system</title><link>/post/fundamentals/ds-stacks-queues/</link><pubDate>Tue, 28 May 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/ds-stacks-queues/</guid><description>&lt;p&gt;Every time you make a function call, your program uses a stack. Every HTTP request in a web server sits in a queue. Every undo operation in your IDE is a stack. These structures are so fundamental they&amp;rsquo;re baked into the hardware itself — the CPU has dedicated stack instructions.&lt;/p&gt;
&lt;p&gt;Yet I consistently see engineers reach for generic slices or channels when a properly implemented stack or queue would be both clearer and faster. Let me show you what these structures actually are, why they&amp;rsquo;re the shape they are, and where they show up in the systems you work on every day.&lt;/p&gt;</description></item><item><title>Lesson 3: File Descriptors — Why Too Many Open Files Kills Your Server</title><link>/post/fundamentals/linux-file-descriptors/</link><pubDate>Mon, 27 May 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/linux-file-descriptors/</guid><description>&lt;p&gt;I got a 3 AM page for a service that was returning connection errors to every client. The application logs said &lt;code&gt;dial tcp: lookup ...: too many open files&lt;/code&gt;. We hadn&amp;rsquo;t changed anything. Load was normal. But over 48 hours, something had been slowly accumulating open file descriptors and not closing them, and we hit the per-process limit. Restarting the service fixed the immediate problem; understanding why it happened required me to actually learn what file descriptors are and how the kernel manages them.&lt;/p&gt;</description></item><item><title>Lesson 2: Design Uber — Real-time matching, geospatial indexing, surge pricing</title><link>/post/fundamentals/sd-deep-uber/</link><pubDate>Sat, 25 May 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/sd-deep-uber/</guid><description>&lt;p&gt;Uber is one of the most instructive system design problems because it forces you to think about real-time data pipelines, spatial indexing, and latency-sensitive matching algorithms all at once. I spent a significant amount of time studying this problem specifically because the geospatial angle is something most candidates gloss over. They say &amp;ldquo;use a database with location queries&amp;rdquo; and move on. But at Uber&amp;rsquo;s scale — 5 million trips per day, hundreds of thousands of concurrent drivers and riders — the geospatial indexing strategy is the entire problem.&lt;/p&gt;</description></item><item><title>Lesson 1: Pod Design Patterns — Sidecar, ambassador, adapter</title><link>/post/fundamentals/k8s-pod-patterns/</link><pubDate>Fri, 24 May 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/k8s-pod-patterns/</guid><description>&lt;p&gt;The first time I deployed a service to Kubernetes, I put everything in a single container. Logging, metrics, TLS termination, the actual application — all in one Docker image. It worked, but the image was massive, updating any single concern meant rebuilding and redeploying the whole thing, and the application team had to understand infrastructure concerns they shouldn&amp;rsquo;t need to care about. The sidecar pattern was the thing that changed how I thought about container composition.&lt;/p&gt;</description></item><item><title>Lesson 3: Write-Ahead Log — How Databases Survive Crashes</title><link>/post/fundamentals/db-wal/</link><pubDate>Thu, 23 May 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/db-wal/</guid><description>&lt;p&gt;Databases promise durability. When &lt;code&gt;INSERT&lt;/code&gt; returns successfully, your data is safe — even if the server loses power a millisecond later. For a long time I accepted this as magic. Then I started reading about what actually happens when Postgres writes data, and the mechanism behind that promise is both elegant and counterintuitive: to make writes safe, you write them twice. The first write goes to a sequential log. The second write goes to the actual data file. And the log is what saves you when things go wrong.&lt;/p&gt;</description></item><item><title>Lesson 2: Code Review That Works — What to look for</title><link>/post/fundamentals/eng-code-review/</link><pubDate>Wed, 22 May 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/eng-code-review/</guid><description>&lt;p&gt;I&amp;rsquo;ve been on the wrong end of bad code reviews in both directions. Reviews that were nitpick sessions about variable naming while missing a race condition. Reviews that rubber-stamped everything because the reviewer was busy. And I&amp;rsquo;ve given both kinds myself. It took a few years, a few incidents, and a few honest retrospectives to develop a framework for reviews that actually improve code quality without burning out reviewers or demoralizing authors.&lt;/p&gt;</description></item><item><title>Lesson 3: Caching — The Hardest Easy Problem in CS</title><link>/post/fundamentals/sd-caching/</link><pubDate>Mon, 20 May 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/sd-caching/</guid><description>&lt;p&gt;Phil Karlton supposedly said: &amp;ldquo;There are only two hard things in Computer Science: cache invalidation and naming things.&amp;rdquo; He was joking, but also completely serious. Caching is the answer to nearly every &amp;ldquo;make it faster&amp;rdquo; problem in system design. It&amp;rsquo;s also the source of some of the most insidious bugs in production: stale data served with confidence, cache stampedes that take down your database, memory explosions from an unbounded cache. The concept is simple. Getting it right is not.&lt;/p&gt;</description></item><item><title>Lesson 2: Clean Architecture — Not the textbook version</title><link>/post/fundamentals/arch-clean/</link><pubDate>Sun, 19 May 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/arch-clean/</guid><description>&lt;p&gt;I&amp;rsquo;ve read the Clean Architecture book. I&amp;rsquo;ve also seen teams implement it so literally that they had six layers of indirection for a CRUD endpoint: a controller called a use case, which called a domain service, which called a port, which went through an adapter, which called a repository, which hit the database. Each hop had its own error mapping. Changing a database column required touching eight files. That&amp;rsquo;s not clean. That&amp;rsquo;s engineering theater.&lt;/p&gt;</description></item><item><title>Lesson 3: Binary Search — The most useful algorithm you'll use weekly</title><link>/post/fundamentals/algo-binary-search/</link><pubDate>Sat, 18 May 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/algo-binary-search/</guid><description>&lt;p&gt;Binary search has a reputation as a simple algorithm — and it is, conceptually. Divide the search space in half, check the middle, repeat. Every programmer knows this. Yet I have seen engineers reach for a linear scan when binary search would have solved the problem in a fraction of the time, and I have also seen subtly buggy binary search implementations that work 99.9% of the time and silently fail on edge cases.&lt;/p&gt;</description></item><item><title>Lesson 1: Leader Election — Someone has to be in charge</title><link>/post/fundamentals/consensus-leader-election/</link><pubDate>Thu, 16 May 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/consensus-leader-election/</guid><description>&lt;p&gt;I spent three days debugging a production incident where two nodes in our cluster both believed they were the primary. Each was accepting writes. Each was replicating to followers. Each was convinced the other was dead. By the time we noticed, we had diverged state that took two more days to reconcile. That incident made me obsessive about leader election — not as an academic concept, but as a concrete engineering problem with real failure modes.&lt;/p&gt;</description></item><item><title>Lesson 3: Hash Maps — O(1) with asterisks</title><link>/post/fundamentals/ds-hash-maps/</link><pubDate>Wed, 15 May 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/ds-hash-maps/</guid><description>&lt;p&gt;Hash maps are the workhorse of production software. Nearly every caching layer, session store, deduplication system, and configuration lookup you&amp;rsquo;ve ever written relies on one. They&amp;rsquo;re fast, they&amp;rsquo;re flexible, and they&amp;rsquo;re genuinely O(1) — with some asterisks that matter enormously in production.&lt;/p&gt;
&lt;p&gt;I want to walk through how they actually work, because the &amp;ldquo;O(1) lookup&amp;rdquo; claim hides a lot of complexity that shows up in the worst possible moments: under load, with adversarial input, or when your hash function is subtly wrong.&lt;/p&gt;</description></item><item><title>Lesson 2: HTTP/2 and HTTP/3 — Multiplexing and QUIC</title><link>/post/fundamentals/net-http2-http3/</link><pubDate>Tue, 14 May 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/net-http2-http3/</guid><description>&lt;p&gt;The first time I looked at a waterfall chart in Chrome DevTools and saw a hundred requests stacking up like rush-hour traffic, I knew HTTP/1.1 was the problem. We had a dashboard loading twelve API calls and a handful of assets, and they were all waiting in line. Not because the server was slow — it was idle — but because the browser had a six-connection-per-host limit and every request had to wait its turn. That was my introduction to why HTTP/2 exists, and it changed how I think about protocol design.&lt;/p&gt;</description></item><item><title>Lesson 2: Virtual Memory — 1GB RSS but Only 50MB is Real</title><link>/post/fundamentals/linux-virtual-memory/</link><pubDate>Sun, 12 May 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/linux-virtual-memory/</guid><description>&lt;p&gt;Early in my career I deployed a Go service and checked &lt;code&gt;top&lt;/code&gt;. The &lt;code&gt;VIRT&lt;/code&gt; column showed 1.2 GB. I nearly had a heart attack — our server had 4 GB of RAM and I thought the service was consuming nearly a third of it. A senior engineer laughed and told me to look at &lt;code&gt;RES&lt;/code&gt; instead: 52 MB. I had no idea what the difference was. Understanding virtual memory is fundamental to reading memory metrics correctly, debugging out-of-memory kills, and understanding how processes interact with the kernel.&lt;/p&gt;</description></item><item><title>Lesson 1: The STAR Method — "Tell me about a time you..." and how to actually answer</title><link>/post/fundamentals/behavioral-star/</link><pubDate>Fri, 10 May 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/behavioral-star/</guid><description>&lt;p&gt;I used to dread behavioral interviews. I was confident in system design rounds, reasonably calm under algorithm pressure, but the moment someone said &amp;ldquo;tell me about a time you disagreed with a technical decision,&amp;rdquo; I felt my brain empty out completely. I would ramble for two minutes, lose the thread halfway through, and trail off into something like &amp;ldquo;&amp;hellip;so yeah, it worked out eventually.&amp;rdquo; Not exactly compelling.&lt;/p&gt;
&lt;p&gt;The STAR method fixed that — not because it&amp;rsquo;s magic, but because it gives a structure that stops you from wandering. Situation, Task, Action, Result. Four buckets. Every behavioral answer you&amp;rsquo;ll ever give fits into them.&lt;/p&gt;</description></item><item><title>Lesson 5: Binary Search — When the Search Space Is Sorted or Monotonic</title><link>/post/fundamentals/interview-binary-search/</link><pubDate>Thu, 09 May 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/interview-binary-search/</guid><description>&lt;p&gt;Binary search has a reputation as the thing you learned in your first CS course and never thought about deeply again. Then you encounter a rotated array or a capacity-minimization problem in an interview, and suddenly it is not obvious at all. I&amp;rsquo;ve seen engineers who could recite the textbook implementation fail completely on problems that are, at their core, binary search — because the search space isn&amp;rsquo;t an array of values, it&amp;rsquo;s something more abstract.&lt;/p&gt;</description></item><item><title>Lesson 2: B-Tree Indexes — O(n) to O(log n) in one CREATE INDEX</title><link>/post/fundamentals/db-btree-indexes/</link><pubDate>Wed, 08 May 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/db-btree-indexes/</guid><description>&lt;p&gt;I once joined a team that had a &lt;code&gt;users&lt;/code&gt; table with 8 million rows. The &lt;code&gt;GET /users/{email}&lt;/code&gt; endpoint was consistently timing out under load. When I ran &lt;code&gt;EXPLAIN ANALYZE&lt;/code&gt; on the query, I saw it: &lt;code&gt;Seq Scan on users (cost=0.00..180000.00 rows=1 width=120) (actual rows=1 loops=1)&lt;/code&gt;. It was reading every single row — all 8 million of them — to find one user by email. Adding a single index fixed it in under five minutes. That experience made me want to actually understand what an index is, not just that it makes things faster.&lt;/p&gt;</description></item><item><title>Lesson 1: Git Beyond Basics — Rebase, bisect, worktrees</title><link>/post/fundamentals/eng-git/</link><pubDate>Tue, 07 May 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/eng-git/</guid><description>&lt;p&gt;Most engineers use maybe 15% of Git. &lt;code&gt;add&lt;/code&gt;, &lt;code&gt;commit&lt;/code&gt;, &lt;code&gt;push&lt;/code&gt;, &lt;code&gt;pull&lt;/code&gt;, &lt;code&gt;branch&lt;/code&gt;, &lt;code&gt;merge&lt;/code&gt;, and &lt;code&gt;status&lt;/code&gt; covers daily work. That&amp;rsquo;s fine until you need to find which commit introduced a regression across 300 commits, or you need to untangle a messy history before merging, or you want to work on three features simultaneously without context-switching overhead. The commands I&amp;rsquo;m covering here don&amp;rsquo;t come up every day. When they do, they save hours.&lt;/p&gt;</description></item><item><title>Lesson 1: Lexing — Turning text into tokens</title><link>/post/fundamentals/compiler-lexing/</link><pubDate>Mon, 06 May 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/compiler-lexing/</guid><description>&lt;p&gt;I wrote my first lexer as a side project after reading the first chapter of &amp;ldquo;Writing an Interpreter in Go.&amp;rdquo; I expected it to be hard. It was not. A lexer is conceptually one of the simpler pieces of a compiler: read characters, group them into meaningful chunks, throw away whitespace. What surprised me was how much clarity it brought to everything downstream. Once I had tokens instead of characters, every subsequent step became easier to reason about. The text became structured. And the structure was entirely my design.&lt;/p&gt;</description></item><item><title>Lesson 2: Load Balancing — L4 vs L7 and Why It Matters</title><link>/post/fundamentals/sd-load-balancing/</link><pubDate>Sun, 05 May 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/sd-load-balancing/</guid><description>&lt;p&gt;The first time I drew a system design diagram in an interview, I drew a box labeled &amp;ldquo;load balancer&amp;rdquo; and drew arrows from clients to it, and from it to servers. My interviewer asked, &amp;ldquo;What kind of load balancer?&amp;rdquo; I didn&amp;rsquo;t have an answer. I knew load balancers existed. I didn&amp;rsquo;t know they made fundamentally different decisions at different network layers — and that the choice between them shapes what your system can and cannot do.&lt;/p&gt;</description></item><item><title>Lesson 1: Monolith First — Starting with microservices is usually wrong</title><link>/post/fundamentals/arch-monolith-first/</link><pubDate>Fri, 03 May 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/arch-monolith-first/</guid><description>&lt;p&gt;I&amp;rsquo;ve seen two teams start new products with microservices. One spent four months before they had anything deployed — Kubernetes setup, service discovery, distributed tracing, a CI/CD pipeline for twelve repos, and debates about how to split domains they hadn&amp;rsquo;t fully understood yet. They ran out of runway. The other team I worked on started with a monolith, shipped their first real user feature in three weeks, and only extracted services when the actual pain of growth made the benefit obvious. We&amp;rsquo;re still running that system, and it handles tens of millions of requests a day.&lt;/p&gt;</description></item><item><title>Lesson 2: Sorting in Practice — When to sort and why TimSort won</title><link>/post/fundamentals/algo-sorting/</link><pubDate>Wed, 01 May 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/algo-sorting/</guid><description>&lt;p&gt;Sorting is one of those topics that feels like a solved problem until you actually need to care about it. Every language ships a standard sort. You call it, it works, you move on. But I have run into subtle production bugs caused by not understanding what the sort is actually doing — unstable sorts breaking tie-breaking logic, sorts on large datasets consuming unexpected memory, and sort comparators with subtle bugs that triggered Go&amp;rsquo;s sort to panic.&lt;/p&gt;</description></item><item><title>Lesson 1: TCP Deep Dive — Three-way handshake, congestion, Nagle</title><link>/post/fundamentals/net-tcp/</link><pubDate>Mon, 29 Apr 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/net-tcp/</guid><description>&lt;p&gt;I used to treat TCP as a black box. Data goes in one side, data comes out the other — reliably, in order, no duplicates. That was all I needed to know, right? Then I started debugging latency spikes in a payment service and spent three days chasing a 200ms tail latency that turned out to be Nagle&amp;rsquo;s algorithm fighting with delayed ACKs. After that, I stopped treating TCP as a black box.&lt;/p&gt;</description></item><item><title>Lesson 2: Linked Lists — Almost never the right choice</title><link>/post/fundamentals/ds-linked-lists/</link><pubDate>Sun, 28 Apr 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/ds-linked-lists/</guid><description>&lt;p&gt;Linked lists are the most over-taught data structure in computer science and the most under-used structure in production systems. I&amp;rsquo;ve reviewed hundreds of pull requests across distributed systems codebases, and I can count on one hand the times a linked list was genuinely the right call.&lt;/p&gt;
&lt;p&gt;That said, understanding why linked lists are usually wrong teaches you something profound about how computers actually work. And there are a handful of situations where they&amp;rsquo;re exactly right — and when those situations come up, you need to recognize them fast.&lt;/p&gt;</description></item><item><title>Lesson 1: Processes and Threads — What Goroutines Map To</title><link>/post/fundamentals/linux-processes/</link><pubDate>Thu, 25 Apr 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/linux-processes/</guid><description>&lt;p&gt;When I first started using Go seriously, I accepted goroutines as &amp;ldquo;lightweight threads&amp;rdquo; without really understanding what that meant. The Go runtime creates them, schedules them, and I launch them with &lt;code&gt;go&lt;/code&gt;. Then I got curious: what does the OS actually see? When I run 10,000 goroutines, does the kernel manage 10,000 things? The answer is no — and understanding why requires understanding the difference between processes, kernel threads, and userspace threads. It also explains why goroutines scale so much better than Java threads or Python threads.&lt;/p&gt;</description></item><item><title>Lesson 1: How a Query Executes — Parser to Planner to Disk</title><link>/post/fundamentals/db-query-execution/</link><pubDate>Mon, 22 Apr 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/db-query-execution/</guid><description>&lt;p&gt;Every time I call &lt;code&gt;db.QueryContext(ctx, &amp;quot;SELECT * FROM orders WHERE user_id = $1&amp;quot;, userID)&lt;/code&gt; in Go, I used to think the database just &amp;ldquo;found&amp;rdquo; the rows. It wasn&amp;rsquo;t until I started debugging a production slowdown — a query that was fast for months and then suddenly wasn&amp;rsquo;t — that I actually traced what happens between the moment my application hands off that SQL string and the moment rows come back. Understanding that pipeline changed how I write queries, design schemas, and diagnose performance problems.&lt;/p&gt;</description></item><item><title>Lesson 1: How the Internet Works — DNS, TCP, HTTP, TLS in 15 Minutes</title><link>/post/fundamentals/sd-internet-works/</link><pubDate>Sat, 20 Apr 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/sd-internet-works/</guid><description>&lt;p&gt;Every system design interview starts with the same silent assumption: you already know how the internet works. Interviewers won&amp;rsquo;t ask you to explain DNS. But when you confidently say &amp;ldquo;the client calls the API&amp;rdquo; without being able to say what actually happens between those words, the cracks show up in your design. Understanding the layers — DNS, TCP, HTTP, TLS — isn&amp;rsquo;t trivia. It&amp;rsquo;s the mental model that tells you where latency hides, why connections are expensive, and what breaks under load.&lt;/p&gt;</description></item><item><title>Lesson 1: Big-O Thinking — Will this scale to 1M records?</title><link>/post/fundamentals/algo-big-o/</link><pubDate>Thu, 18 Apr 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/algo-big-o/</guid><description>&lt;p&gt;I spent two years writing Go services before I really internalized Big-O. Not because I didn&amp;rsquo;t know the notation — every CS course teaches you to recite O(n log n) — but because I never tied it to a real decision I had to make in production. It clicked for me the day a coworker asked, &amp;ldquo;will this work when we have a million users?&amp;rdquo; and I had no honest answer.&lt;/p&gt;</description></item><item><title>Lesson 1: Arrays and Memory Layout — Cache lines decide your performance</title><link>/post/fundamentals/ds-arrays-memory/</link><pubDate>Mon, 15 Apr 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/ds-arrays-memory/</guid><description>&lt;p&gt;I spent two years writing Go services before I genuinely understood why iterating over a two-dimensional slice in the wrong order could tank my throughput by 5x. It wasn&amp;rsquo;t a bug. It wasn&amp;rsquo;t a bad algorithm. It was cache lines.&lt;/p&gt;
&lt;p&gt;Arrays are the first data structure everyone learns and the last one most engineers actually understand. This is my attempt to fix that — not with theory, but with the reasoning that makes you a better systems engineer.&lt;/p&gt;</description></item><item><title>Lesson 4: Stack — Last In, First Out Solves More Than You Think</title><link>/post/fundamentals/interview-stack/</link><pubDate>Sat, 13 Apr 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/interview-stack/</guid><description>&lt;p&gt;The stack is one of those data structures that seems too simple to be interesting — push, pop, peek, done. Then you sit in an interview, stare at a problem about brackets or expression evaluation, and realize you&amp;rsquo;re reaching for exactly this tool. I&amp;rsquo;ve seen candidates overcomplicate parenthesis validation with counters and flags when a stack makes it four lines of logic. I&amp;rsquo;ve also seen the reverse: candidates who knew the stack solution for valid parentheses but couldn&amp;rsquo;t extend the thinking to a harder variant.&lt;/p&gt;</description></item><item><title>Lesson 1: Design YouTube — Video upload, transcoding, streaming at scale</title><link>/post/fundamentals/sd-deep-youtube/</link><pubDate>Tue, 09 Apr 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/sd-deep-youtube/</guid><description>&lt;p&gt;YouTube serves over 500 hours of video uploaded every minute and delivers billions of views per day. When I first studied this problem seriously, I made the mistake of treating it as a simple &amp;ldquo;upload file, store it, serve it&amp;rdquo; exercise. It isn&amp;rsquo;t. The interesting engineering is in what happens between the moment a creator hits upload and the moment a viewer&amp;rsquo;s video starts playing seamlessly on a 3G connection in rural India. That gap is where the system design lives.&lt;/p&gt;</description></item><item><title>Lesson 3: Sliding Window — Fixed or Variable, the Window Always Moves Right</title><link>/post/fundamentals/interview-sliding-window/</link><pubDate>Thu, 28 Mar 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/interview-sliding-window/</guid><description>&lt;p&gt;Sliding window problems have a tell: the problem asks about a contiguous subarray or substring, and there&amp;rsquo;s a constraint that makes the brute force obvious but slow. When I was interviewing at a mid-size fintech that fancied itself FAANG-adjacent, I got a stock price problem in my first round. I almost panicked — &amp;ldquo;is this dynamic programming?&amp;rdquo; It wasn&amp;rsquo;t. It was a sliding window in disguise, and once I saw it, the solution wrote itself in about four minutes.&lt;/p&gt;</description></item><item><title>Lesson 2: Two Pointers — When One Pointer Isn't Enough</title><link>/post/fundamentals/interview-two-pointers/</link><pubDate>Thu, 14 Mar 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/interview-two-pointers/</guid><description>&lt;p&gt;One of my early interview mistakes was throwing a hash map at every array problem. It works a surprising amount of the time, but there&amp;rsquo;s a whole class of problems where you don&amp;rsquo;t need extra space at all — where the structure of the input lets two indices do the work together. Two pointers is that technique, and once you see it, you cannot unsee it.&lt;/p&gt;
&lt;p&gt;The pattern shows up at every company. Amazon loves it for stream-processing problems. Google uses it for geometry and partition questions. Meta favors it in string validation and deduplication. Three Sum alone has appeared in more on-site rounds than I can count.&lt;/p&gt;</description></item><item><title>Lesson 1: Arrays and Hashing — The Pattern Behind 30% of All Interview Questions</title><link>/post/fundamentals/interview-arrays-hashing/</link><pubDate>Sun, 03 Mar 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/interview-arrays-hashing/</guid><description>&lt;p&gt;When I was preparing for my first round of FAANG interviews, I was overwhelmed. Hundreds of LeetCode problems, dozens of patterns, no clear starting point. Then I noticed something: roughly a third of the medium-difficulty problems I encountered could be cracked with one insight — trading time for space using a hash map. Once I internalized that, arrays and hashing stopped feeling like a category and started feeling like a reflex.&lt;/p&gt;</description></item></channel></rss>