<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Algorithms on</title><link>/tags/algorithms/</link><description>Recent content in Algorithms on</description><generator>Hugo</generator><language>en</language><lastBuildDate>Fri, 01 Nov 2024 00:00:00 +0000</lastBuildDate><atom:link href="/tags/algorithms/index.xml" rel="self" type="application/rss+xml"/><item><title>Lesson 14: Randomized Algorithms — Reservoir sampling, HyperLogLog, probabilistic counting</title><link>/post/fundamentals/algo-randomized/</link><pubDate>Fri, 01 Nov 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/algo-randomized/</guid><description>&lt;p&gt;There is a class of production problems where exact answers are either impossible or not worth the cost. You want to know approximately how many unique visitors hit your site today. You want to sample 1% of requests for tracing without reading every request into memory first. You want to check if a username is already taken without querying the database on every keystroke.&lt;/p&gt;
&lt;p&gt;Randomized algorithms provide exact answers with known probability bounds, or approximate answers with bounded error, using a fraction of the memory or time that exact computation would require. I was skeptical of &amp;ldquo;probabilistic&amp;rdquo; algorithms for a long time — it seemed like trading correctness for efficiency. Then I learned what the error bounds actually are. A HyperLogLog cardinality estimate with 1.5% error using 12KB of memory is a better engineering choice than a perfect count using 100MB, for the vast majority of use cases.&lt;/p&gt;</description></item><item><title>Lesson 13: Cryptographic Primitives — Hashing, HMAC, and never rolling your own</title><link>/post/fundamentals/algo-crypto/</link><pubDate>Thu, 17 Oct 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/algo-crypto/</guid><description>&lt;p&gt;There is a joke in security circles: every developer thinks they can write their own crypto. The punchline is that everyone who has tried has been wrong. Cryptography is the one area of computer science where being 99% correct is the same as being completely wrong. A subtle timing vulnerability, a nonce reuse, or a hash function with the wrong properties can completely destroy a security guarantee that looks solid on paper.&lt;/p&gt;</description></item><item><title>Lesson 12: Compression Basics — Why gzip works and entropy matters</title><link>/post/fundamentals/algo-compression/</link><pubDate>Thu, 03 Oct 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/algo-compression/</guid><description>&lt;p&gt;Every senior engineer I know makes compression decisions regularly: should this API response be gzip-compressed? Should logs be stored compressed? What compression level? Which algorithm? I made these decisions for years based on &amp;ldquo;gzip is standard, use it&amp;rdquo; without understanding why it worked or when something else might be better.&lt;/p&gt;
&lt;p&gt;Understanding the fundamentals of how compression works — entropy, Huffman coding, and LZ77 — changed how I think about data formats, wire protocols, and storage costs. It also helped me understand why some data compresses well and some data does not, which is critical for capacity planning.&lt;/p&gt;</description></item><item><title>Lesson 11: String Algorithms — KMP, Rabin-Karp, and why regex can be slow</title><link>/post/fundamentals/algo-string/</link><pubDate>Tue, 17 Sep 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/algo-string/</guid><description>&lt;p&gt;String processing sits at the foundation of almost every production system. Log parsing, protocol parsing, search, validation, templating — it all comes down to finding and transforming patterns in text. Most of the time, strings.Contains or a simple loop is fast enough. But when you are processing millions of log lines per second, or running user-supplied patterns against untrusted input, or implementing a search feature that needs to handle long documents, naive string matching becomes a bottleneck or a security hole.&lt;/p&gt;</description></item><item><title>Lesson 10: Backtracking — Constraint satisfaction and config generation</title><link>/post/fundamentals/algo-backtracking/</link><pubDate>Mon, 02 Sep 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/algo-backtracking/</guid><description>&lt;p&gt;Backtracking is the algorithmic equivalent of &amp;ldquo;try everything, but be smart about giving up early.&amp;rdquo; It is a systematic way to explore a search space where you build a solution incrementally and abandon partial solutions as soon as you detect they cannot possibly lead to a valid result.&lt;/p&gt;
&lt;p&gt;I first encountered backtracking outside of textbooks when building a scheduling system that needed to assign employees to shifts while respecting availability constraints, skill requirements, and labor regulations. The brute force approach — enumerate all possible assignments — was computationally equivalent to exploring all permutations, which was 20! for 20 employees. Backtracking with constraint pruning reduced the search space by several orders of magnitude.&lt;/p&gt;</description></item><item><title>Lesson 9: Greedy Algorithms — When being selfish is optimal</title><link>/post/fundamentals/algo-greedy/</link><pubDate>Sat, 17 Aug 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/algo-greedy/</guid><description>&lt;p&gt;Greedy algorithms have a simple idea at their core: at each step, make the locally optimal choice. No looking ahead, no considering alternatives, no backtracking. Just take the best available option right now and trust that it leads to a globally optimal result.&lt;/p&gt;
&lt;p&gt;The catch is that this only works for specific problem structures. Use greedy when the problem has the &lt;strong&gt;greedy choice property&lt;/strong&gt; — the locally optimal choice is always part of a globally optimal solution. When this holds, greedy is elegant and fast. When it does not, greedy gives you a wrong answer with no warning.&lt;/p&gt;</description></item><item><title>Lesson 8: Dynamic Programming Intuition — Memoization, not memorization</title><link>/post/fundamentals/algo-dynamic-programming/</link><pubDate>Fri, 02 Aug 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/algo-dynamic-programming/</guid><description>&lt;p&gt;Dynamic programming has an intimidating reputation. The name sounds academic, the problems in textbooks often involve sequences and matrices, and the &amp;ldquo;aha moment&amp;rdquo; is notoriously hard to force. I spent a long time treating DP as an interview preparation topic rather than a tool I would actually use. That changed when I built a pricing engine and realized I had been reimplementing DP badly — without knowing it — by computing the same values over and over in nested function calls.&lt;/p&gt;</description></item><item><title>Lesson 7: Shortest Path — Dijkstra in routing and network optimization</title><link>/post/fundamentals/algo-shortest-path/</link><pubDate>Fri, 19 Jul 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/algo-shortest-path/</guid><description>&lt;p&gt;When I joined a team building a multi-region traffic routing system, the first thing I had to understand was why the routing decisions were not always the geographically shortest path. The system was using Dijkstra&amp;rsquo;s algorithm, but the edge weights were not just latency — they incorporated bandwidth cost, current utilization, failure rates, and SLA constraints. Dijkstra did not care what the weights meant. It just found the minimum cost path. That is its power.&lt;/p&gt;</description></item><item><title>Lesson 6: BFS and DFS — Dependency resolution, crawlers, cycle detection</title><link>/post/fundamentals/algo-bfs-dfs/</link><pubDate>Wed, 03 Jul 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/algo-bfs-dfs/</guid><description>&lt;p&gt;Graph traversal sounds abstract until you realize that most interesting data in production is a graph. Service dependencies are a graph. Database foreign key relationships form a graph. Build tool dependencies, org charts, permission hierarchies, network topologies — all graphs. BFS and DFS are the two fundamental ways to walk them, and they show up in real engineering work more than almost any other algorithm.&lt;/p&gt;
&lt;p&gt;I have used DFS to detect circular imports in a build system, BFS to find the shortest migration path between two schema versions, and both to debug why a dependency injection container was resolving services in the wrong order.&lt;/p&gt;</description></item><item><title>Lesson 5: Hashing and Consistent Hashing — How load balancers distribute traffic</title><link>/post/fundamentals/algo-hashing/</link><pubDate>Mon, 17 Jun 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/algo-hashing/</guid><description>&lt;p&gt;I did not think much about hashing until I had to explain why a cache cluster was unusable after we added two new nodes. We had doubled the cache hit rate over six months, added two boxes to handle the load, and immediately destroyed most of our cached data. Every key rehashed to a different server. Cache hit rate dropped from 85% to under 10%. We were hammering the database.&lt;/p&gt;</description></item><item><title>Lesson 4: Two Pointers and Sliding Window — Stream processing in disguise</title><link>/post/fundamentals/algo-two-pointers/</link><pubDate>Sat, 01 Jun 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/algo-two-pointers/</guid><description>&lt;p&gt;The first time I recognized the sliding window pattern outside a textbook, I was reading the source code for a rate limiter. There was no comment saying &amp;ldquo;sliding window algorithm here.&amp;rdquo; There was just a loop, two indices into a circular buffer, and some simple arithmetic that maintained a count of events in the last N seconds. Once I saw it, I started seeing it everywhere — in network flow control, in moving average calculations, in deduplication logic.&lt;/p&gt;</description></item><item><title>Lesson 3: Binary Search — The most useful algorithm you'll use weekly</title><link>/post/fundamentals/algo-binary-search/</link><pubDate>Sat, 18 May 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/algo-binary-search/</guid><description>&lt;p&gt;Binary search has a reputation as a simple algorithm — and it is, conceptually. Divide the search space in half, check the middle, repeat. Every programmer knows this. Yet I have seen engineers reach for a linear scan when binary search would have solved the problem in a fraction of the time, and I have also seen subtly buggy binary search implementations that work 99.9% of the time and silently fail on edge cases.&lt;/p&gt;</description></item><item><title>Lesson 2: Sorting in Practice — When to sort and why TimSort won</title><link>/post/fundamentals/algo-sorting/</link><pubDate>Wed, 01 May 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/algo-sorting/</guid><description>&lt;p&gt;Sorting is one of those topics that feels like a solved problem until you actually need to care about it. Every language ships a standard sort. You call it, it works, you move on. But I have run into subtle production bugs caused by not understanding what the sort is actually doing — unstable sorts breaking tie-breaking logic, sorts on large datasets consuming unexpected memory, and sort comparators with subtle bugs that triggered Go&amp;rsquo;s sort to panic.&lt;/p&gt;</description></item><item><title>Lesson 1: Big-O Thinking — Will this scale to 1M records?</title><link>/post/fundamentals/algo-big-o/</link><pubDate>Thu, 18 Apr 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/algo-big-o/</guid><description>&lt;p&gt;I spent two years writing Go services before I really internalized Big-O. Not because I didn&amp;rsquo;t know the notation — every CS course teaches you to recite O(n log n) — but because I never tied it to a real decision I had to make in production. It clicked for me the day a coworker asked, &amp;ldquo;will this work when we have a million users?&amp;rdquo; and I had no honest answer.&lt;/p&gt;</description></item></channel></rss>