<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Distributed Systems on</title><link>/tags/distributed-systems/</link><description>Recent content in Distributed Systems on</description><generator>Hugo</generator><language>en</language><lastBuildDate>Wed, 04 Jun 2025 11:08:00 +0000</lastBuildDate><atom:link href="/tags/distributed-systems/index.xml" rel="self" type="application/rss+xml"/><item><title>Lesson 10: Distributed System Patterns — Consensus, CRDTs, and consistency</title><link>/post/rust/rust-net-distributed-patterns/</link><pubDate>Wed, 04 Jun 2025 11:08:00 +0000</pubDate><guid>/post/rust/rust-net-distributed-patterns/</guid><description>&lt;p&gt;I once watched a team spend six months building a &amp;ldquo;distributed database&amp;rdquo; that was really just PostgreSQL with a cron job that copied rows between data centers. It worked until it didn&amp;rsquo;t — conflicting writes, lost updates, and an incident where the same order was fulfilled twice from different warehouses. They learned the hard way that distributed systems aren&amp;rsquo;t just &amp;ldquo;run it on multiple machines.&amp;rdquo; They&amp;rsquo;re a fundamentally different programming model with different guarantees, different failure modes, and different mental models.&lt;/p&gt;</description></item><item><title>Lesson 9: Message Queues — NATS, Kafka, RabbitMQ from Rust</title><link>/post/rust/rust-net-message-queues/</link><pubDate>Sun, 01 Jun 2025 15:25:00 +0000</pubDate><guid>/post/rust/rust-net-message-queues/</guid><description>&lt;p&gt;The moment I stopped thinking of services as calling each other and started thinking of them as reacting to events, my architecture got dramatically simpler. Instead of service A calling service B calling service C in a synchronous chain that&amp;rsquo;s as fragile as it sounds, service A publishes an event. Services B and C subscribe and react independently. A doesn&amp;rsquo;t even know they exist. B can be down for maintenance without affecting A. C can be added next month without changing A&amp;rsquo;s code. Message queues are the backbone of this pattern.&lt;/p&gt;</description></item><item><title>Lesson 8: Circuit Breakers in Rust — Failing fast</title><link>/post/rust/rust-net-circuit-breaker/</link><pubDate>Thu, 29 May 2025 07:35:00 +0000</pubDate><guid>/post/rust/rust-net-circuit-breaker/</guid><description>&lt;p&gt;Picture this: your payment service depends on a fraud detection API that&amp;rsquo;s completely down. Every request to it takes 30 seconds to timeout. Your payment service has 200 requests queued up, each holding a thread and a database connection while waiting for fraud detection to respond. Within minutes, you&amp;rsquo;re out of connections, the payment service itself starts failing, and now the checkout service that depends on payments starts failing too. One dead service has cascaded into a full outage.&lt;/p&gt;</description></item><item><title>Lesson 7: Retry Strategies and Exponential Backoff — Resilient clients</title><link>/post/rust/rust-net-retries/</link><pubDate>Mon, 26 May 2025 10:40:00 +0000</pubDate><guid>/post/rust/rust-net-retries/</guid><description>&lt;p&gt;Here&amp;rsquo;s a scenario that&amp;rsquo;s burned me more than once: a downstream service has a brief hiccup — maybe a pod is restarting, maybe there&amp;rsquo;s a momentary network partition — and instead of gracefully retrying, my service immediately returns a 500 to every caller. The hiccup lasts 3 seconds. My P99 latency graph spikes. On-call gets paged. Everyone&amp;rsquo;s unhappy. The fix? A retry loop with exponential backoff. Three lines of logic that would&amp;rsquo;ve made the entire incident invisible.&lt;/p&gt;</description></item><item><title>Lesson 6: TLS — rustls and native TLS</title><link>/post/rust/rust-net-tls/</link><pubDate>Fri, 23 May 2025 19:10:00 +0000</pubDate><guid>/post/rust/rust-net-tls/</guid><description>&lt;p&gt;I&amp;rsquo;ll never forget the 3am page that turned out to be an expired TLS certificate. Our automated renewal had been silently failing for two weeks, nobody noticed because the cert was still valid, and then at 2:47am on a Sunday it expired and every client started getting connection errors. We had monitoring for CPU, memory, disk, latency, error rates — but not for certificate expiry. That was the day I decided to actually understand TLS instead of just copy-pasting cert paths into config files.&lt;/p&gt;</description></item><item><title>Lesson 5: DNS Resolution and Custom Resolvers — Understanding name resolution</title><link>/post/rust/rust-net-dns/</link><pubDate>Wed, 21 May 2025 13:55:00 +0000</pubDate><guid>/post/rust/rust-net-dns/</guid><description>&lt;p&gt;A few months back, our entire staging environment went down for an hour. Not because any service crashed — because someone changed a DNS record and forgot that our Kubernetes ingress had a 5-minute TTL cache while the CDN had a 24-hour cache. Half our traffic was going to the old IP, half to the new one. Debugging it took forever because &lt;code&gt;dig&lt;/code&gt; on my laptop showed the correct answer, but the services inside the cluster were seeing stale records.&lt;/p&gt;</description></item><item><title>Lesson 4: WebSocket Servers and Clients — Real-time communication</title><link>/post/rust/rust-net-websockets/</link><pubDate>Sun, 18 May 2025 08:20:00 +0000</pubDate><guid>/post/rust/rust-net-websockets/</guid><description>&lt;p&gt;I built my first WebSocket server to power a live dashboard that showed deployment status across our fleet. The alternative was polling every 2 seconds — 500 browser tabs hitting the API, each getting back the same &amp;ldquo;nothing changed&amp;rdquo; response 99% of the time. WebSockets turned that from 250 requests/second of wasted work into a handful of persistent connections that only sent data when something actually happened.&lt;/p&gt;
&lt;h2 id="http-vs-websockets--when-do-you-need-them"&gt;HTTP vs WebSockets — When Do You Need Them?&lt;/h2&gt;
&lt;p&gt;HTTP is request-response. Client asks, server answers. Great for most things. But some use cases fundamentally don&amp;rsquo;t fit that model:&lt;/p&gt;</description></item><item><title>Lesson 3: gRPC with tonic — High-performance RPC</title><link>/post/rust/rust-net-grpc/</link><pubDate>Fri, 16 May 2025 11:30:00 +0000</pubDate><guid>/post/rust/rust-net-grpc/</guid><description>&lt;p&gt;The first time I used gRPC in production, I was skeptical. We already had REST APIs that worked fine — why add protobuf compilation, code generation, and an entirely new protocol? Then our team grew to four services in three languages, and the answer became painfully obvious. Every REST endpoint had slightly different JSON field naming, different error formats, and documentation that was always a version behind. gRPC eliminated all of that overnight.&lt;/p&gt;</description></item><item><title>Lesson 2: HTTP Clients — reqwest and hyper</title><link>/post/rust/rust-net-http-client/</link><pubDate>Wed, 14 May 2025 16:45:00 +0000</pubDate><guid>/post/rust/rust-net-http-client/</guid><description>&lt;p&gt;I once spent three hours debugging a production issue that turned out to be an HTTP client with no timeout configured. Three hours. The client was happily waiting forever for a response from a service that had crashed, holding a database connection open the entire time. That experience permanently changed how I think about HTTP clients — they&amp;rsquo;re not just &amp;ldquo;make a request, get a response.&amp;rdquo; They&amp;rsquo;re complex state machines with connection pools, redirect policies, timeout hierarchies, and a dozen other knobs that matter when things go wrong.&lt;/p&gt;</description></item><item><title>Lesson 1: Building a TCP Server from Scratch — Raw sockets</title><link>/post/rust/rust-net-tcp-server/</link><pubDate>Mon, 12 May 2025 09:14:00 +0000</pubDate><guid>/post/rust/rust-net-tcp-server/</guid><description>&lt;p&gt;Last month I was debugging a flaky microservice at work and realized I couldn&amp;rsquo;t explain what was actually happening between &lt;code&gt;bind()&lt;/code&gt; and the first byte arriving. I&amp;rsquo;d been using high-level frameworks for years — Actix, Axum, you name it — but I&amp;rsquo;d never actually built a TCP server from raw sockets in Rust. That bothered me. So I spent a weekend doing exactly that, and honestly, it changed how I think about every networked service I write.&lt;/p&gt;</description></item><item><title>Lesson 3: Paxos and Beyond — When Raft isn't enough</title><link>/post/fundamentals/consensus-paxos/</link><pubDate>Fri, 04 Oct 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/consensus-paxos/</guid><description>&lt;p&gt;After spending time with Raft, I found myself curious about the algorithm it was designed to replace. Paxos has a reputation: brilliant, correct, nearly impossible to implement correctly, and even harder to extend to practical systems. Leslie Lamport published the original Paxos paper in 1989, got it rejected, submitted a revised version in 1998, and it became the theoretical foundation for a generation of distributed systems. Chubby (Google&amp;rsquo;s distributed lock service), Zookeeper (the coordination service), and the precursor to Spanner all descended from Paxos thinking. Understanding why Raft was necessary requires understanding what Paxos gets right and where it falls short in practice.&lt;/p&gt;</description></item><item><title>Lesson 2: Raft Consensus — The consensus algorithm you can actually understand</title><link>/post/fundamentals/consensus-raft/</link><pubDate>Fri, 26 Jul 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/consensus-raft/</guid><description>&lt;p&gt;When I first tried to understand Paxos — the original distributed consensus algorithm — I read the paper three times and still felt like I was missing something. I could follow each step individually, but I couldn&amp;rsquo;t build a mental model of why it worked or what the invariants were. Raft was designed specifically to fix that. Its paper is literally titled &amp;ldquo;In Search of an Understandability: The Raft Consensus Algorithm.&amp;rdquo; After reading it, I could explain it to someone else. That&amp;rsquo;s the bar Raft was designed to clear, and it does.&lt;/p&gt;</description></item><item><title>Lesson 1: Leader Election — Someone has to be in charge</title><link>/post/fundamentals/consensus-leader-election/</link><pubDate>Thu, 16 May 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/consensus-leader-election/</guid><description>&lt;p&gt;I spent three days debugging a production incident where two nodes in our cluster both believed they were the primary. Each was accepting writes. Each was replicating to followers. Each was convinced the other was dead. By the time we noticed, we had diverged state that took two more days to reconcile. That incident made me obsessive about leader election — not as an academic concept, but as a concrete engineering problem with real failure modes.&lt;/p&gt;</description></item></channel></rss>