<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>System Design From First Principles on Atharva Pandey</title><link>https://atharva.page/series/system-design-from-first-principles/</link><description>Recent content in System Design From First Principles on Atharva Pandey</description><generator>Hugo</generator><language>en-us</language><copyright>Copyright ©</copyright><lastBuildDate>Tue, 19 Nov 2024 00:00:00 +0000</lastBuildDate><atom:link href="https://atharva.page/series/system-design-from-first-principles/index.xml" rel="self" type="application/rss+xml"/><item><title>Lesson 15: CAP Theorem in Practice — What It Actually Means for Your System</title><link>https://atharva.page/post/fundamentals/sd-cap-theorem/</link><pubDate>Tue, 19 Nov 2024 00:00:00 +0000</pubDate><guid>https://atharva.page/post/fundamentals/sd-cap-theorem/</guid><description>&lt;p&gt;CAP theorem is probably the most cited and most misunderstood concept in distributed systems interviews. Candidates memorize &amp;ldquo;you can only pick two of consistency, availability, and partition tolerance&amp;rdquo; and then either over-apply it (treating every design decision as a CAP trade-off) or under-apply it (never relating it to actual design choices). The theorem is real and important, but the way it&amp;rsquo;s usually taught in 30-second summaries strips out the nuance that makes it actually useful. This final lesson clears that up.&lt;/p&gt;</description></item><item><title>Lesson 14: Designing for Failure — Circuit Breakers, Bulkheads, Chaos</title><link>https://atharva.page/post/fundamentals/sd-designing-failure/</link><pubDate>Mon, 04 Nov 2024 00:00:00 +0000</pubDate><guid>https://atharva.page/post/fundamentals/sd-designing-failure/</guid><description>&lt;p&gt;In distributed systems, failure is not an exceptional case — it&amp;rsquo;s the default condition. Networks partition. Hard drives fail. Memory fills up. Dependencies have bugs. Every distributed system that&amp;rsquo;s been running for more than a few years has experienced every kind of failure you can imagine, and many you can&amp;rsquo;t. The engineers who build resilient systems aren&amp;rsquo;t smarter than the ones who don&amp;rsquo;t. They&amp;rsquo;ve just internalized a single principle: design for when things go wrong, not for when things go right.&lt;/p&gt;</description></item><item><title>Lesson 13: Design a Payment System — Idempotency, Reconciliation, Double-Entry</title><link>https://atharva.page/post/fundamentals/sd-payment-system/</link><pubDate>Mon, 21 Oct 2024 00:00:00 +0000</pubDate><guid>https://atharva.page/post/fundamentals/sd-payment-system/</guid><description>&lt;p&gt;Payment systems have a property that almost no other software has: the cost of a bug isn&amp;rsquo;t a bad user experience — it&amp;rsquo;s a legal liability and a business catastrophe. Charging a customer twice, losing a transfer in a network failure, or crediting the wrong account can result in millions of dollars of loss and destroyed trust. Every other system we&amp;rsquo;ve covered tolerates a degree of eventual inconsistency. Payment systems, in most cases, do not. This lesson is about building for that level of correctness.&lt;/p&gt;</description></item><item><title>Lesson 12: Design a Search Engine — Inverted Index and Ranking</title><link>https://atharva.page/post/fundamentals/sd-search-engine/</link><pubDate>Mon, 07 Oct 2024 00:00:00 +0000</pubDate><guid>https://atharva.page/post/fundamentals/sd-search-engine/</guid><description>&lt;p&gt;Search is the feature that separates usable products from unusable ones at scale. When your application has thousands of documents, a sequential scan works. At millions, it doesn&amp;rsquo;t. The difference between a search box that works and one that times out is a data structure invented in the 1960s that every modern search engine still fundamentally relies on: the inverted index. Understanding how it works — and how to build a system around it — is what this lesson is about.&lt;/p&gt;</description></item><item><title>Lesson 11: Design a Notification System — Push vs Pull, Priority, Dedup</title><link>https://atharva.page/post/fundamentals/sd-notifications/</link><pubDate>Sat, 21 Sep 2024 00:00:00 +0000</pubDate><guid>https://atharva.page/post/fundamentals/sd-notifications/</guid><description>&lt;p&gt;Notifications are the feature that can make or break user retention — and also destroy it. Done right, they bring users back at exactly the right moment. Done wrong, they&amp;rsquo;re spam that drives uninstalls. The system design challenge isn&amp;rsquo;t just the technical plumbing (though that&amp;rsquo;s interesting). It&amp;rsquo;s building infrastructure that&amp;rsquo;s fast for critical alerts, reliable for important messages, and smart enough to not overwhelm users with low-priority noise.&lt;/p&gt;
&lt;h2 id="the-core-concept"&gt;The Core Concept&lt;/h2&gt;
&lt;p&gt;A notification system has to handle multiple channels with wildly different characteristics:&lt;/p&gt;</description></item><item><title>Lesson 10: Design a News Feed — Fan-out on Write vs Read</title><link>https://atharva.page/post/fundamentals/sd-news-feed/</link><pubDate>Fri, 06 Sep 2024 00:00:00 +0000</pubDate><guid>https://atharva.page/post/fundamentals/sd-news-feed/</guid><description>&lt;p&gt;The news feed is the product feature that defines social media. Every time you open Instagram or Twitter, you see a personalized, ranked, real-time stream of content from people you follow. Behind that deceptively simple UI is one of the hardest distributed systems problems in consumer tech: how do you compute a personalized feed for hundreds of millions of users, where any piece of content needs to appear in potentially millions of feeds, within seconds of being posted?&lt;/p&gt;</description></item><item><title>Lesson 9: Design a Chat System — WebSocket, Presence, Message Ordering</title><link>https://atharva.page/post/fundamentals/sd-chat-system/</link><pubDate>Wed, 21 Aug 2024 00:00:00 +0000</pubDate><guid>https://atharva.page/post/fundamentals/sd-chat-system/</guid><description>&lt;p&gt;Building a chat system is the interview problem that catches people on protocol fundamentals. HTTP is a request-response protocol — the client asks, the server answers, and then the connection is idle. For chat, the server needs to push messages to clients the moment they arrive. This inversion of the HTTP model is what makes chat hard, and it&amp;rsquo;s what drives the entire architecture.&lt;/p&gt;
&lt;h2 id="the-core-concept"&gt;The Core Concept&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Why HTTP polling doesn&amp;rsquo;t work at scale&lt;/strong&gt;&lt;/p&gt;</description></item><item><title>Lesson 8: Design a URL Shortener — The Classic, Done Properly</title><link>https://atharva.page/post/fundamentals/sd-url-shortener/</link><pubDate>Sun, 04 Aug 2024 00:00:00 +0000</pubDate><guid>https://atharva.page/post/fundamentals/sd-url-shortener/</guid><description>&lt;p&gt;The URL shortener is the &amp;ldquo;Hello World&amp;rdquo; of system design interviews. It appears deceptively simple: take a long URL, return a short one. But if you treat it superficially, you miss what the interviewer is actually testing: your ability to think through ID generation at scale, read-heavy caching, redirect semantics, analytics storage, and data modeling. Done well, the URL shortener problem touches nearly every fundamental we&amp;rsquo;ve covered so far.&lt;/p&gt;
&lt;h2 id="the-core-concept"&gt;The Core Concept&lt;/h2&gt;
&lt;p&gt;A URL shortener has two primary operations:&lt;/p&gt;</description></item><item><title>Lesson 7: Rate Limiting — Token Bucket, Sliding Window, Distributed</title><link>https://atharva.page/post/fundamentals/sd-rate-limiting/</link><pubDate>Sun, 21 Jul 2024 00:00:00 +0000</pubDate><guid>https://atharva.page/post/fundamentals/sd-rate-limiting/</guid><description>&lt;p&gt;Without rate limiting, a single misbehaving client can consume all your resources and deny service to everyone else. A bug in a client that retries in a tight loop. A competitor scraping your API. A DDoS attack. A feature that accidentally calls your endpoint ten times per button click. Rate limiting is the mechanism that prevents any of these from taking down your service — it&amp;rsquo;s the first line of defense between the internet and your origin.&lt;/p&gt;</description></item><item><title>Lesson 6: CDNs — Put Your Bytes Close to Your Users</title><link>https://atharva.page/post/fundamentals/sd-cdns/</link><pubDate>Thu, 04 Jul 2024 00:00:00 +0000</pubDate><guid>https://atharva.page/post/fundamentals/sd-cdns/</guid><description>&lt;p&gt;Physics is the enemy of performance. Light travels through fiber optic cables at about 200,000 km/s — roughly two-thirds the speed of light in a vacuum. A round-trip from New York to London is about 11,000 km each way. That means the minimum possible latency for that trip is 55ms. You cannot engineer your way past it. What you can do is stop making the trip at all. Content Delivery Networks are the infrastructure that puts popular content at the edge — closer to users, so the bytes never have to travel far.&lt;/p&gt;</description></item><item><title>Lesson 5: Message Queues — Decoupling Services Without Losing Messages</title><link>https://atharva.page/post/fundamentals/sd-message-queues/</link><pubDate>Wed, 19 Jun 2024 00:00:00 +0000</pubDate><guid>https://atharva.page/post/fundamentals/sd-message-queues/</guid><description>&lt;p&gt;Synchronous calls are elegant until they&amp;rsquo;re not. Service A calls Service B. B is slow. A waits. A&amp;rsquo;s request pool fills up. A becomes slow. The caller of A waits. The whole request chain stalls. Add a few more services in the chain and you have a distributed deadlock in slow motion. Message queues exist to break this dependency — to let a producer say &amp;ldquo;here&amp;rsquo;s some work&amp;rdquo; and move on, without caring whether the consumer is fast, slow, or temporarily down.&lt;/p&gt;</description></item><item><title>Lesson 4: Database Scaling — Read Replicas, Sharding, and When Each Helps</title><link>https://atharva.page/post/fundamentals/sd-database-scaling/</link><pubDate>Wed, 05 Jun 2024 00:00:00 +0000</pubDate><guid>https://atharva.page/post/fundamentals/sd-database-scaling/</guid><description>&lt;p&gt;Your startup ships, gains traction, and one day your database becomes the bottleneck. Queries slow down. CPU spikes. The single machine hosting your Postgres instance can&amp;rsquo;t keep up. This is one of the most predictable problems in engineering, yet I&amp;rsquo;ve seen teams reach it completely unprepared because they&amp;rsquo;d never thought carefully about how databases scale. The right solution depends entirely on whether you&amp;rsquo;re read-heavy or write-heavy, whether your data is relational or not, and whether your workload is even or spiky.&lt;/p&gt;</description></item><item><title>Lesson 3: Caching — The Hardest Easy Problem in CS</title><link>https://atharva.page/post/fundamentals/sd-caching/</link><pubDate>Mon, 20 May 2024 00:00:00 +0000</pubDate><guid>https://atharva.page/post/fundamentals/sd-caching/</guid><description>&lt;p&gt;Phil Karlton supposedly said: &amp;ldquo;There are only two hard things in Computer Science: cache invalidation and naming things.&amp;rdquo; He was joking, but also completely serious. Caching is the answer to nearly every &amp;ldquo;make it faster&amp;rdquo; problem in system design. It&amp;rsquo;s also the source of some of the most insidious bugs in production: stale data served with confidence, cache stampedes that take down your database, memory explosions from an unbounded cache. The concept is simple. Getting it right is not.&lt;/p&gt;</description></item><item><title>Lesson 2: Load Balancing — L4 vs L7 and Why It Matters</title><link>https://atharva.page/post/fundamentals/sd-load-balancing/</link><pubDate>Sun, 05 May 2024 00:00:00 +0000</pubDate><guid>https://atharva.page/post/fundamentals/sd-load-balancing/</guid><description>&lt;p&gt;The first time I drew a system design diagram in an interview, I drew a box labeled &amp;ldquo;load balancer&amp;rdquo; and drew arrows from clients to it, and from it to servers. My interviewer asked, &amp;ldquo;What kind of load balancer?&amp;rdquo; I didn&amp;rsquo;t have an answer. I knew load balancers existed. I didn&amp;rsquo;t know they made fundamentally different decisions at different network layers — and that the choice between them shapes what your system can and cannot do.&lt;/p&gt;</description></item><item><title>Lesson 1: How the Internet Works — DNS, TCP, HTTP, TLS in 15 Minutes</title><link>https://atharva.page/post/fundamentals/sd-internet-works/</link><pubDate>Sat, 20 Apr 2024 00:00:00 +0000</pubDate><guid>https://atharva.page/post/fundamentals/sd-internet-works/</guid><description>&lt;p&gt;Every system design interview starts with the same silent assumption: you already know how the internet works. Interviewers won&amp;rsquo;t ask you to explain DNS. But when you confidently say &amp;ldquo;the client calls the API&amp;rdquo; without being able to say what actually happens between those words, the cracks show up in your design. Understanding the layers — DNS, TCP, HTTP, TLS — isn&amp;rsquo;t trivia. It&amp;rsquo;s the mental model that tells you where latency hides, why connections are expensive, and what breaks under load.&lt;/p&gt;</description></item></channel></rss>