<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>System Design on</title><link>/tags/system-design/</link><description>Recent content in System Design on</description><generator>Hugo</generator><language>en</language><lastBuildDate>Fri, 13 Dec 2024 00:00:00 +0000</lastBuildDate><atom:link href="/tags/system-design/index.xml" rel="self" type="application/rss+xml"/><item><title>Lesson 4: Recommendation Systems — Collaborative filtering, embeddings, and the cold start problem</title><link>/post/fundamentals/ml-recommendations/</link><pubDate>Fri, 13 Dec 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/ml-recommendations/</guid><description>&lt;p&gt;Recommendation systems are the most economically consequential ML systems most engineers will ever build. Netflix estimates that its recommendation system saves over a billion dollars per year in avoided cancellations. Amazon&amp;rsquo;s &amp;ldquo;customers who bought this also bought&amp;rdquo; drives a significant fraction of its revenue. Spotify&amp;rsquo;s Discover Weekly has become a user retention flywheel. The stakes are real, and the engineering is genuinely interesting.&lt;/p&gt;
&lt;p&gt;I got deep into recommendation systems when I was working on a content platform that had about 500,000 items and needed to surface relevant content for each user without overwhelming the ranking team with engineering requests. What I found was that the architecture of a production recommender is almost always the same shape, regardless of the domain — and the hardest problem is not the ML, it&amp;rsquo;s the cold start.&lt;/p&gt;</description></item><item><title>Lesson 15: CAP Theorem in Practice — What It Actually Means for Your System</title><link>/post/fundamentals/sd-cap-theorem/</link><pubDate>Tue, 19 Nov 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/sd-cap-theorem/</guid><description>&lt;p&gt;CAP theorem is probably the most cited and most misunderstood concept in distributed systems interviews. Candidates memorize &amp;ldquo;you can only pick two of consistency, availability, and partition tolerance&amp;rdquo; and then either over-apply it (treating every design decision as a CAP trade-off) or under-apply it (never relating it to actual design choices). The theorem is real and important, but the way it&amp;rsquo;s usually taught in 30-second summaries strips out the nuance that makes it actually useful. This final lesson clears that up.&lt;/p&gt;</description></item><item><title>Lesson 5: Design Twitter/X — Tweet fanout, timeline ranking, trending topics at 500M users</title><link>/post/fundamentals/sd-deep-twitter/</link><pubDate>Thu, 14 Nov 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/sd-deep-twitter/</guid><description>&lt;p&gt;Twitter is the classic system design problem for good reason. It looks like a glorified blog until you start pulling on the threads: how does a tweet from a user with 100 million followers appear in every follower&amp;rsquo;s timeline within seconds? How do you rank timelines without reading millions of tweets per request? How do you identify trending topics across 500 million users in near real-time? Each of these is a genuinely hard problem, and they interact in non-obvious ways.&lt;/p&gt;</description></item><item><title>Lesson 14: Designing for Failure — Circuit Breakers, Bulkheads, Chaos</title><link>/post/fundamentals/sd-designing-failure/</link><pubDate>Mon, 04 Nov 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/sd-designing-failure/</guid><description>&lt;p&gt;In distributed systems, failure is not an exceptional case — it&amp;rsquo;s the default condition. Networks partition. Hard drives fail. Memory fills up. Dependencies have bugs. Every distributed system that&amp;rsquo;s been running for more than a few years has experienced every kind of failure you can imagine, and many you can&amp;rsquo;t. The engineers who build resilient systems aren&amp;rsquo;t smarter than the ones who don&amp;rsquo;t. They&amp;rsquo;ve just internalized a single principle: design for when things go wrong, not for when things go right.&lt;/p&gt;</description></item><item><title>Lesson 13: Design a Payment System — Idempotency, Reconciliation, Double-Entry</title><link>/post/fundamentals/sd-payment-system/</link><pubDate>Mon, 21 Oct 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/sd-payment-system/</guid><description>&lt;p&gt;Payment systems have a property that almost no other software has: the cost of a bug isn&amp;rsquo;t a bad user experience — it&amp;rsquo;s a legal liability and a business catastrophe. Charging a customer twice, losing a transfer in a network failure, or crediting the wrong account can result in millions of dollars of loss and destroyed trust. Every other system we&amp;rsquo;ve covered tolerates a degree of eventual inconsistency. Payment systems, in most cases, do not. This lesson is about building for that level of correctness.&lt;/p&gt;</description></item><item><title>Lesson 3: Model Serving — Latency, batching, A/B testing in production</title><link>/post/fundamentals/ml-model-serving/</link><pubDate>Mon, 14 Oct 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/ml-model-serving/</guid><description>&lt;p&gt;Training a model is satisfying. Deploying it to serve real traffic is humbling. I remember the first time I pushed a model to production and watched the p99 latency hover at 800ms on what was supposed to be a &amp;ldquo;fast&amp;rdquo; model. The benchmark had shown 12ms inference time. What happened? The benchmark ran the model on pre-loaded batches; production served one request at a time, loaded the model fresh on cold starts, and had no GPU batching. The gap between &amp;ldquo;the model is accurate&amp;rdquo; and &amp;ldquo;the model is fast enough to be useful&amp;rdquo; is where model serving engineering lives.&lt;/p&gt;</description></item><item><title>Lesson 12: Design a Search Engine — Inverted Index and Ranking</title><link>/post/fundamentals/sd-search-engine/</link><pubDate>Mon, 07 Oct 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/sd-search-engine/</guid><description>&lt;p&gt;Search is the feature that separates usable products from unusable ones at scale. When your application has thousands of documents, a sequential scan works. At millions, it doesn&amp;rsquo;t. The difference between a search box that works and one that times out is a data structure invented in the 1960s that every modern search engine still fundamentally relies on: the inverted index. Understanding how it works — and how to build a system around it — is what this lesson is about.&lt;/p&gt;</description></item><item><title>Lesson 11: Design a Notification System — Push vs Pull, Priority, Dedup</title><link>/post/fundamentals/sd-notifications/</link><pubDate>Sat, 21 Sep 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/sd-notifications/</guid><description>&lt;p&gt;Notifications are the feature that can make or break user retention — and also destroy it. Done right, they bring users back at exactly the right moment. Done wrong, they&amp;rsquo;re spam that drives uninstalls. The system design challenge isn&amp;rsquo;t just the technical plumbing (though that&amp;rsquo;s interesting). It&amp;rsquo;s building infrastructure that&amp;rsquo;s fast for critical alerts, reliable for important messages, and smart enough to not overwhelm users with low-priority noise.&lt;/p&gt;
&lt;h2 id="the-core-concept"&gt;The Core Concept&lt;/h2&gt;
&lt;p&gt;A notification system has to handle multiple channels with wildly different characteristics:&lt;/p&gt;</description></item><item><title>Lesson 4: Design WhatsApp — End-to-end encryption, message delivery guarantees, presence</title><link>/post/fundamentals/sd-deep-whatsapp/</link><pubDate>Sat, 14 Sep 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/sd-deep-whatsapp/</guid><description>&lt;p&gt;WhatsApp is deceptively simple from a user perspective: you send a message, it arrives. But building a messaging system that handles 100 billion messages per day with end-to-end encryption, reliable delivery semantics, and real-time presence for 2 billion users is a genuinely hard engineering problem. I find this problem particularly instructive because it forces you to confront three things simultaneously: cryptographic key management, message delivery guarantees, and the cost of maintaining online/offline state at massive scale.&lt;/p&gt;</description></item><item><title>Lesson 10: Design a News Feed — Fan-out on Write vs Read</title><link>/post/fundamentals/sd-news-feed/</link><pubDate>Fri, 06 Sep 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/sd-news-feed/</guid><description>&lt;p&gt;The news feed is the product feature that defines social media. Every time you open Instagram or Twitter, you see a personalized, ranked, real-time stream of content from people you follow. Behind that deceptively simple UI is one of the hardest distributed systems problems in consumer tech: how do you compute a personalized feed for hundreds of millions of users, where any piece of content needs to appear in potentially millions of feeds, within seconds of being posted?&lt;/p&gt;</description></item><item><title>Lesson 9: Design a Chat System — WebSocket, Presence, Message Ordering</title><link>/post/fundamentals/sd-chat-system/</link><pubDate>Wed, 21 Aug 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/sd-chat-system/</guid><description>&lt;p&gt;Building a chat system is the interview problem that catches people on protocol fundamentals. HTTP is a request-response protocol — the client asks, the server answers, and then the connection is idle. For chat, the server needs to push messages to clients the moment they arrive. This inversion of the HTTP model is what makes chat hard, and it&amp;rsquo;s what drives the entire architecture.&lt;/p&gt;
&lt;h2 id="the-core-concept"&gt;The Core Concept&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Why HTTP polling doesn&amp;rsquo;t work at scale&lt;/strong&gt;&lt;/p&gt;</description></item><item><title>Lesson 8: Design a URL Shortener — The Classic, Done Properly</title><link>/post/fundamentals/sd-url-shortener/</link><pubDate>Sun, 04 Aug 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/sd-url-shortener/</guid><description>&lt;p&gt;The URL shortener is the &amp;ldquo;Hello World&amp;rdquo; of system design interviews. It appears deceptively simple: take a long URL, return a short one. But if you treat it superficially, you miss what the interviewer is actually testing: your ability to think through ID generation at scale, read-heavy caching, redirect semantics, analytics storage, and data modeling. Done well, the URL shortener problem touches nearly every fundamental we&amp;rsquo;ve covered so far.&lt;/p&gt;
&lt;h2 id="the-core-concept"&gt;The Core Concept&lt;/h2&gt;
&lt;p&gt;A URL shortener has two primary operations:&lt;/p&gt;</description></item><item><title>Lesson 2: Feature Stores — Why feature engineering is 80% of ML</title><link>/post/fundamentals/ml-feature-stores/</link><pubDate>Fri, 02 Aug 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/ml-feature-stores/</guid><description>&lt;p&gt;There is a saying in machine learning that is so universally acknowledged it has become a cliché: 80% of the work in any ML project is feature engineering. I spent a long time thinking this referred to the cognitive labor — the domain expertise required to craft meaningful features. It does, in part. But the deeper meaning is operational. The 80% is not just about &lt;em&gt;what&lt;/em&gt; features to build; it&amp;rsquo;s about &lt;em&gt;how&lt;/em&gt; to compute them consistently, store them efficiently, retrieve them with sub-millisecond latency at serving time, and keep them synchronized between the training pipeline and the production system.&lt;/p&gt;</description></item><item><title>Lesson 7: Rate Limiting — Token Bucket, Sliding Window, Distributed</title><link>/post/fundamentals/sd-rate-limiting/</link><pubDate>Sun, 21 Jul 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/sd-rate-limiting/</guid><description>&lt;p&gt;Without rate limiting, a single misbehaving client can consume all your resources and deny service to everyone else. A bug in a client that retries in a tight loop. A competitor scraping your API. A DDoS attack. A feature that accidentally calls your endpoint ten times per button click. Rate limiting is the mechanism that prevents any of these from taking down your service — it&amp;rsquo;s the first line of defense between the internet and your origin.&lt;/p&gt;</description></item><item><title>Lesson 3: Design Google Docs — Real-time collaboration, OT vs CRDT, conflict resolution</title><link>/post/fundamentals/sd-deep-google-docs/</link><pubDate>Mon, 15 Jul 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/sd-deep-google-docs/</guid><description>&lt;p&gt;Google Docs is the problem I recommend to every engineer who thinks they understand distributed systems. The surface looks trivial: multiple users editing a document simultaneously. The depth is staggering. When two users type at the same position in a document at the same millisecond, what does each user see? How do you converge on a consistent state without a central lock? How do you preserve the intention behind each edit, not just the characters?&lt;/p&gt;</description></item><item><title>Lesson 6: CDNs — Put Your Bytes Close to Your Users</title><link>/post/fundamentals/sd-cdns/</link><pubDate>Thu, 04 Jul 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/sd-cdns/</guid><description>&lt;p&gt;Physics is the enemy of performance. Light travels through fiber optic cables at about 200,000 km/s — roughly two-thirds the speed of light in a vacuum. A round-trip from New York to London is about 11,000 km each way. That means the minimum possible latency for that trip is 55ms. You cannot engineer your way past it. What you can do is stop making the trip at all. Content Delivery Networks are the infrastructure that puts popular content at the edge — closer to users, so the bytes never have to travel far.&lt;/p&gt;</description></item><item><title>Lesson 5: Message Queues — Decoupling Services Without Losing Messages</title><link>/post/fundamentals/sd-message-queues/</link><pubDate>Wed, 19 Jun 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/sd-message-queues/</guid><description>&lt;p&gt;Synchronous calls are elegant until they&amp;rsquo;re not. Service A calls Service B. B is slow. A waits. A&amp;rsquo;s request pool fills up. A becomes slow. The caller of A waits. The whole request chain stalls. Add a few more services in the chain and you have a distributed deadlock in slow motion. Message queues exist to break this dependency — to let a producer say &amp;ldquo;here&amp;rsquo;s some work&amp;rdquo; and move on, without caring whether the consumer is fast, slow, or temporarily down.&lt;/p&gt;</description></item><item><title>Lesson 4: Database Scaling — Read Replicas, Sharding, and When Each Helps</title><link>/post/fundamentals/sd-database-scaling/</link><pubDate>Wed, 05 Jun 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/sd-database-scaling/</guid><description>&lt;p&gt;Your startup ships, gains traction, and one day your database becomes the bottleneck. Queries slow down. CPU spikes. The single machine hosting your Postgres instance can&amp;rsquo;t keep up. This is one of the most predictable problems in engineering, yet I&amp;rsquo;ve seen teams reach it completely unprepared because they&amp;rsquo;d never thought carefully about how databases scale. The right solution depends entirely on whether you&amp;rsquo;re read-heavy or write-heavy, whether your data is relational or not, and whether your workload is even or spiky.&lt;/p&gt;</description></item><item><title>Lesson 1: ML Pipelines — From raw data to deployed model</title><link>/post/fundamentals/ml-pipelines/</link><pubDate>Sun, 02 Jun 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/ml-pipelines/</guid><description>&lt;p&gt;The first time I built a &amp;ldquo;production ML pipeline,&amp;rdquo; it was a cron job that ran a Python script, trained a model, and saved the artifact to disk. It worked, until it didn&amp;rsquo;t. The model silently degraded over three weeks because the training data schema changed and nobody noticed. There were no tests, no validation, no monitoring. That experience, embarrassing as it was, taught me more about ML system design than any paper on model architecture.&lt;/p&gt;</description></item><item><title>Lesson 2: Design Uber — Real-time matching, geospatial indexing, surge pricing</title><link>/post/fundamentals/sd-deep-uber/</link><pubDate>Sat, 25 May 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/sd-deep-uber/</guid><description>&lt;p&gt;Uber is one of the most instructive system design problems because it forces you to think about real-time data pipelines, spatial indexing, and latency-sensitive matching algorithms all at once. I spent a significant amount of time studying this problem specifically because the geospatial angle is something most candidates gloss over. They say &amp;ldquo;use a database with location queries&amp;rdquo; and move on. But at Uber&amp;rsquo;s scale — 5 million trips per day, hundreds of thousands of concurrent drivers and riders — the geospatial indexing strategy is the entire problem.&lt;/p&gt;</description></item><item><title>Lesson 3: Caching — The Hardest Easy Problem in CS</title><link>/post/fundamentals/sd-caching/</link><pubDate>Mon, 20 May 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/sd-caching/</guid><description>&lt;p&gt;Phil Karlton supposedly said: &amp;ldquo;There are only two hard things in Computer Science: cache invalidation and naming things.&amp;rdquo; He was joking, but also completely serious. Caching is the answer to nearly every &amp;ldquo;make it faster&amp;rdquo; problem in system design. It&amp;rsquo;s also the source of some of the most insidious bugs in production: stale data served with confidence, cache stampedes that take down your database, memory explosions from an unbounded cache. The concept is simple. Getting it right is not.&lt;/p&gt;</description></item><item><title>Lesson 2: Load Balancing — L4 vs L7 and Why It Matters</title><link>/post/fundamentals/sd-load-balancing/</link><pubDate>Sun, 05 May 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/sd-load-balancing/</guid><description>&lt;p&gt;The first time I drew a system design diagram in an interview, I drew a box labeled &amp;ldquo;load balancer&amp;rdquo; and drew arrows from clients to it, and from it to servers. My interviewer asked, &amp;ldquo;What kind of load balancer?&amp;rdquo; I didn&amp;rsquo;t have an answer. I knew load balancers existed. I didn&amp;rsquo;t know they made fundamentally different decisions at different network layers — and that the choice between them shapes what your system can and cannot do.&lt;/p&gt;</description></item><item><title>Lesson 1: How the Internet Works — DNS, TCP, HTTP, TLS in 15 Minutes</title><link>/post/fundamentals/sd-internet-works/</link><pubDate>Sat, 20 Apr 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/sd-internet-works/</guid><description>&lt;p&gt;Every system design interview starts with the same silent assumption: you already know how the internet works. Interviewers won&amp;rsquo;t ask you to explain DNS. But when you confidently say &amp;ldquo;the client calls the API&amp;rdquo; without being able to say what actually happens between those words, the cracks show up in your design. Understanding the layers — DNS, TCP, HTTP, TLS — isn&amp;rsquo;t trivia. It&amp;rsquo;s the mental model that tells you where latency hides, why connections are expensive, and what breaks under load.&lt;/p&gt;</description></item><item><title>Lesson 1: Design YouTube — Video upload, transcoding, streaming at scale</title><link>/post/fundamentals/sd-deep-youtube/</link><pubDate>Tue, 09 Apr 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/sd-deep-youtube/</guid><description>&lt;p&gt;YouTube serves over 500 hours of video uploaded every minute and delivers billions of views per day. When I first studied this problem seriously, I made the mistake of treating it as a simple &amp;ldquo;upload file, store it, serve it&amp;rdquo; exercise. It isn&amp;rsquo;t. The interesting engineering is in what happens between the moment a creator hits upload and the moment a viewer&amp;rsquo;s video starts playing seamlessly on a 3G connection in rural India. That gap is where the system design lives.&lt;/p&gt;</description></item></channel></rss>