<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Networking on</title><link>/tags/networking/</link><description>Recent content in Networking on</description><generator>Hugo</generator><language>en</language><lastBuildDate>Wed, 04 Jun 2025 11:08:00 +0000</lastBuildDate><atom:link href="/tags/networking/index.xml" rel="self" type="application/rss+xml"/><item><title>Lesson 10: Distributed System Patterns — Consensus, CRDTs, and consistency</title><link>/post/rust/rust-net-distributed-patterns/</link><pubDate>Wed, 04 Jun 2025 11:08:00 +0000</pubDate><guid>/post/rust/rust-net-distributed-patterns/</guid><description>&lt;p&gt;I once watched a team spend six months building a &amp;ldquo;distributed database&amp;rdquo; that was really just PostgreSQL with a cron job that copied rows between data centers. It worked until it didn&amp;rsquo;t — conflicting writes, lost updates, and an incident where the same order was fulfilled twice from different warehouses. They learned the hard way that distributed systems aren&amp;rsquo;t just &amp;ldquo;run it on multiple machines.&amp;rdquo; They&amp;rsquo;re a fundamentally different programming model with different guarantees, different failure modes, and different mental models.&lt;/p&gt;</description></item><item><title>Lesson 9: Message Queues — NATS, Kafka, RabbitMQ from Rust</title><link>/post/rust/rust-net-message-queues/</link><pubDate>Sun, 01 Jun 2025 15:25:00 +0000</pubDate><guid>/post/rust/rust-net-message-queues/</guid><description>&lt;p&gt;The moment I stopped thinking of services as calling each other and started thinking of them as reacting to events, my architecture got dramatically simpler. Instead of service A calling service B calling service C in a synchronous chain that&amp;rsquo;s as fragile as it sounds, service A publishes an event. Services B and C subscribe and react independently. A doesn&amp;rsquo;t even know they exist. B can be down for maintenance without affecting A. C can be added next month without changing A&amp;rsquo;s code. Message queues are the backbone of this pattern.&lt;/p&gt;</description></item><item><title>Lesson 8: Circuit Breakers in Rust — Failing fast</title><link>/post/rust/rust-net-circuit-breaker/</link><pubDate>Thu, 29 May 2025 07:35:00 +0000</pubDate><guid>/post/rust/rust-net-circuit-breaker/</guid><description>&lt;p&gt;Picture this: your payment service depends on a fraud detection API that&amp;rsquo;s completely down. Every request to it takes 30 seconds to timeout. Your payment service has 200 requests queued up, each holding a thread and a database connection while waiting for fraud detection to respond. Within minutes, you&amp;rsquo;re out of connections, the payment service itself starts failing, and now the checkout service that depends on payments starts failing too. One dead service has cascaded into a full outage.&lt;/p&gt;</description></item><item><title>Lesson 7: Retry Strategies and Exponential Backoff — Resilient clients</title><link>/post/rust/rust-net-retries/</link><pubDate>Mon, 26 May 2025 10:40:00 +0000</pubDate><guid>/post/rust/rust-net-retries/</guid><description>&lt;p&gt;Here&amp;rsquo;s a scenario that&amp;rsquo;s burned me more than once: a downstream service has a brief hiccup — maybe a pod is restarting, maybe there&amp;rsquo;s a momentary network partition — and instead of gracefully retrying, my service immediately returns a 500 to every caller. The hiccup lasts 3 seconds. My P99 latency graph spikes. On-call gets paged. Everyone&amp;rsquo;s unhappy. The fix? A retry loop with exponential backoff. Three lines of logic that would&amp;rsquo;ve made the entire incident invisible.&lt;/p&gt;</description></item><item><title>Lesson 6: TLS — rustls and native TLS</title><link>/post/rust/rust-net-tls/</link><pubDate>Fri, 23 May 2025 19:10:00 +0000</pubDate><guid>/post/rust/rust-net-tls/</guid><description>&lt;p&gt;I&amp;rsquo;ll never forget the 3am page that turned out to be an expired TLS certificate. Our automated renewal had been silently failing for two weeks, nobody noticed because the cert was still valid, and then at 2:47am on a Sunday it expired and every client started getting connection errors. We had monitoring for CPU, memory, disk, latency, error rates — but not for certificate expiry. That was the day I decided to actually understand TLS instead of just copy-pasting cert paths into config files.&lt;/p&gt;</description></item><item><title>Lesson 5: DNS Resolution and Custom Resolvers — Understanding name resolution</title><link>/post/rust/rust-net-dns/</link><pubDate>Wed, 21 May 2025 13:55:00 +0000</pubDate><guid>/post/rust/rust-net-dns/</guid><description>&lt;p&gt;A few months back, our entire staging environment went down for an hour. Not because any service crashed — because someone changed a DNS record and forgot that our Kubernetes ingress had a 5-minute TTL cache while the CDN had a 24-hour cache. Half our traffic was going to the old IP, half to the new one. Debugging it took forever because &lt;code&gt;dig&lt;/code&gt; on my laptop showed the correct answer, but the services inside the cluster were seeing stale records.&lt;/p&gt;</description></item><item><title>Lesson 4: WebSocket Servers and Clients — Real-time communication</title><link>/post/rust/rust-net-websockets/</link><pubDate>Sun, 18 May 2025 08:20:00 +0000</pubDate><guid>/post/rust/rust-net-websockets/</guid><description>&lt;p&gt;I built my first WebSocket server to power a live dashboard that showed deployment status across our fleet. The alternative was polling every 2 seconds — 500 browser tabs hitting the API, each getting back the same &amp;ldquo;nothing changed&amp;rdquo; response 99% of the time. WebSockets turned that from 250 requests/second of wasted work into a handful of persistent connections that only sent data when something actually happened.&lt;/p&gt;
&lt;h2 id="http-vs-websockets--when-do-you-need-them"&gt;HTTP vs WebSockets — When Do You Need Them?&lt;/h2&gt;
&lt;p&gt;HTTP is request-response. Client asks, server answers. Great for most things. But some use cases fundamentally don&amp;rsquo;t fit that model:&lt;/p&gt;</description></item><item><title>Lesson 3: gRPC with tonic — High-performance RPC</title><link>/post/rust/rust-net-grpc/</link><pubDate>Fri, 16 May 2025 11:30:00 +0000</pubDate><guid>/post/rust/rust-net-grpc/</guid><description>&lt;p&gt;The first time I used gRPC in production, I was skeptical. We already had REST APIs that worked fine — why add protobuf compilation, code generation, and an entirely new protocol? Then our team grew to four services in three languages, and the answer became painfully obvious. Every REST endpoint had slightly different JSON field naming, different error formats, and documentation that was always a version behind. gRPC eliminated all of that overnight.&lt;/p&gt;</description></item><item><title>Lesson 8: gRPC Basics and Streaming — Protobuf on the wire, types in your code</title><link>/post/go/go-net-grpc/</link><pubDate>Thu, 15 May 2025 00:00:00 +0000</pubDate><guid>/post/go/go-net-grpc/</guid><description>&lt;p&gt;My team migrated a set of internal service APIs from JSON over HTTP/1.1 to gRPC roughly two years ago. The motivating factors were type safety across service boundaries, binary serialization that was 5-10x smaller on the wire, and bidirectional streaming that HTTP/1.1 can&amp;rsquo;t do at all. The migration took about a week per service and the performance improvements were immediately visible in our latency percentiles.&lt;/p&gt;
&lt;p&gt;gRPC is a Remote Procedure Call framework that runs over HTTP/2. The interface is defined in Protocol Buffers — a language-neutral schema language — and the gRPC toolchain generates client and server stubs in Go. You call a method; the framework handles serialization, connection management, and streaming.&lt;/p&gt;</description></item><item><title>Lesson 2: HTTP Clients — reqwest and hyper</title><link>/post/rust/rust-net-http-client/</link><pubDate>Wed, 14 May 2025 16:45:00 +0000</pubDate><guid>/post/rust/rust-net-http-client/</guid><description>&lt;p&gt;I once spent three hours debugging a production issue that turned out to be an HTTP client with no timeout configured. Three hours. The client was happily waiting forever for a response from a service that had crashed, holding a database connection open the entire time. That experience permanently changed how I think about HTTP clients — they&amp;rsquo;re not just &amp;ldquo;make a request, get a response.&amp;rdquo; They&amp;rsquo;re complex state machines with connection pools, redirect policies, timeout hierarchies, and a dozen other knobs that matter when things go wrong.&lt;/p&gt;</description></item><item><title>Lesson 1: Building a TCP Server from Scratch — Raw sockets</title><link>/post/rust/rust-net-tcp-server/</link><pubDate>Mon, 12 May 2025 09:14:00 +0000</pubDate><guid>/post/rust/rust-net-tcp-server/</guid><description>&lt;p&gt;Last month I was debugging a flaky microservice at work and realized I couldn&amp;rsquo;t explain what was actually happening between &lt;code&gt;bind()&lt;/code&gt; and the first byte arriving. I&amp;rsquo;d been using high-level frameworks for years — Actix, Axum, you name it — but I&amp;rsquo;d never actually built a TCP server from raw sockets in Rust. That bothered me. So I spent a weekend doing exactly that, and honestly, it changed how I think about every networked service I write.&lt;/p&gt;</description></item><item><title>Lesson 7: Outbox Pattern — Reliable events without distributed transactions</title><link>/post/go/go-net-outbox/</link><pubDate>Thu, 20 Mar 2025 00:00:00 +0000</pubDate><guid>/post/go/go-net-outbox/</guid><description>&lt;p&gt;Here&amp;rsquo;s a bug that&amp;rsquo;s bitten almost every distributed system I&amp;rsquo;ve worked on: a service saves a record to the database and then publishes an event to Kafka. The record is saved. The process crashes before the publish. Now the database says the order exists, but the fulfillment service never heard about it. The order sits in limbo forever.&lt;/p&gt;
&lt;p&gt;The naive fix — wrapping both operations in a transaction — doesn&amp;rsquo;t work because Kafka isn&amp;rsquo;t a participant in your database transaction. You can&amp;rsquo;t two-phase commit across Postgres and Kafka without a distributed transaction coordinator, and nobody wants to operate one of those in production.&lt;/p&gt;</description></item><item><title>Lesson 6: Eventual Consistency — Your data will be wrong, temporarily</title><link>/post/go/go-net-eventual-consistency/</link><pubDate>Sat, 25 Jan 2025 00:00:00 +0000</pubDate><guid>/post/go/go-net-eventual-consistency/</guid><description>&lt;p&gt;I used to believe that eventual consistency was an exotic property of distributed databases that I&amp;rsquo;d only encounter at Google scale. Then I added a Redis cache to a simple CRUD service and immediately had a bug where users couldn&amp;rsquo;t see their own updates. Welcome to eventual consistency — it shows up the moment you have two places that store the same data.&lt;/p&gt;
&lt;p&gt;Eventual consistency means that if you stop writing to a system, all replicas will eventually converge to the same value. The word &amp;ldquo;eventually&amp;rdquo; can mean milliseconds or minutes, depending on the system. The hard part isn&amp;rsquo;t the definition — it&amp;rsquo;s designing application code that works correctly even when replicas haven&amp;rsquo;t converged yet.&lt;/p&gt;</description></item><item><title>Lesson 5: Message Queues in Go — NATS, RabbitMQ, Kafka — pick your tradeoff</title><link>/post/go/go-net-message-queues/</link><pubDate>Wed, 20 Nov 2024 00:00:00 +0000</pubDate><guid>/post/go/go-net-message-queues/</guid><description>&lt;p&gt;I&amp;rsquo;ve integrated all three of these systems in production Go services, and the question I get most often is: &amp;ldquo;Which one should I use?&amp;rdquo; The honest answer is that it depends on what guarantee you actually need — and most engineers pick based on familiarity or hype rather than requirements. NATS, RabbitMQ, and Kafka solve meaningfully different problems, and choosing the wrong one creates operational pain that no amount of clever application code can fix.&lt;/p&gt;</description></item><item><title>Lesson 4: Connection Pooling — One connection per request is a performance bug</title><link>/post/go/go-net-conn-pooling/</link><pubDate>Sat, 12 Oct 2024 00:00:00 +0000</pubDate><guid>/post/go/go-net-conn-pooling/</guid><description>&lt;p&gt;The first time I benchmarked a Go service against a database, the numbers were embarrassing. Five hundred requests per second, each one taking 30 milliseconds to execute a trivially simple query. The query itself took 2 milliseconds on the database server. The other 28 milliseconds were TCP handshake plus TLS plus PostgreSQL authentication — repeated for every single request because I had no connection pool.&lt;/p&gt;
&lt;p&gt;Connection pooling is one of those topics that feels like an advanced optimization until you discover that nearly every database driver in Go already pools connections by default — you just have to configure the pool instead of leaving it at its default settings, which are almost always wrong for your workload.&lt;/p&gt;</description></item><item><title>Lesson 3: Circuit Breaking — Stop calling the service that's already down</title><link>/post/go/go-net-circuit-breaker/</link><pubDate>Wed, 28 Aug 2024 00:00:00 +0000</pubDate><guid>/post/go/go-net-circuit-breaker/</guid><description>&lt;p&gt;There&amp;rsquo;s a particular kind of production incident that goes like this: Service A calls Service B. Service B starts responding slowly because its database is under pressure. Service A&amp;rsquo;s goroutines pile up waiting for responses. After a few minutes, Service A is out of memory, and now two services are down instead of one. The postmortem note says &amp;ldquo;cascading failure.&amp;rdquo; The fix, which nobody implemented, is a circuit breaker.&lt;/p&gt;
&lt;p&gt;The circuit breaker pattern comes from electrical engineering. When too much current flows through a circuit, the breaker trips — it opens the circuit and prevents more current from flowing until someone resets it. In software, the &amp;ldquo;current&amp;rdquo; is requests, and &amp;ldquo;tripping&amp;rdquo; means refusing to make calls to a failing downstream instead of piling up timeouts and consuming resources.&lt;/p&gt;</description></item><item><title>Lesson 7: Service Mesh — Sidecar proxies and mTLS without code</title><link>/post/fundamentals/net-service-mesh/</link><pubDate>Sat, 03 Aug 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/net-service-mesh/</guid><description>&lt;p&gt;A team I worked with had seventeen microservices. Each service had its own implementation of retry logic, circuit breaking, timeout handling, and mutual TLS. Some used libraries, some rolled their own. When we needed to update the TLS certificate rotation policy, it touched eleven different code repositories, four different languages, and took two months. Then we introduced Linkerd and moved all of that to the infrastructure layer. The services still did their jobs. The networking became someone else&amp;rsquo;s problem — specifically, the platform team&amp;rsquo;s.&lt;/p&gt;</description></item><item><title>Lesson 2: Retries with Exponential Backoff — Retry right or retry forever</title><link>/post/go/go-net-retries/</link><pubDate>Sat, 20 Jul 2024 00:00:00 +0000</pubDate><guid>/post/go/go-net-retries/</guid><description>&lt;p&gt;I once worked on a service that retried failed HTTP requests in a tight loop with no backoff and no jitter. When our upstream had a brief outage, every instance of our service fired retries simultaneously, on the same cadence, forever. The upstream came back online, got immediately hammered by a synchronized retry storm from fifty instances, went down again, and the cycle repeated for forty minutes. The &amp;ldquo;retry&amp;rdquo; logic had turned a five-minute outage into a cascading failure.&lt;/p&gt;</description></item><item><title>Lesson 6: gRPC and Protobuf — Binary protocols and streaming</title><link>/post/fundamentals/net-grpc/</link><pubDate>Wed, 17 Jul 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/net-grpc/</guid><description>&lt;p&gt;The first JSON API I replaced with gRPC was passing &lt;code&gt;[]Order&lt;/code&gt; objects around — each order had about 40 fields, most of which the caller never used. The JSON payload for a list of 100 orders was around 180KB. After the migration it was 22KB, and the serialization time in benchmarks dropped by 8x. But more than the performance numbers, what I noticed was the schema. Proto files are contracts. When a field changes, you know it. With JSON, you find out when things break in production.&lt;/p&gt;</description></item><item><title>Lesson 5: WebSockets — Upgrade, framing, vs SSE vs polling</title><link>/post/fundamentals/net-websockets/</link><pubDate>Mon, 01 Jul 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/net-websockets/</guid><description>&lt;p&gt;When I first built a real-time notification system, I reached for WebSockets because that&amp;rsquo;s what everyone said to use for &amp;ldquo;real-time.&amp;rdquo; Three months later I was debugging connection drops under load, wrestling with proxy timeouts, and fighting with nginx configuration. When I stepped back and actually thought about the access pattern — server pushing notifications to the browser, never the browser sending data to the server — I replaced WebSockets with Server-Sent Events in a weekend. The code got simpler, the proxies stopped complaining, and we had less to maintain. Picking the right tool requires understanding what each one actually does.&lt;/p&gt;</description></item><item><title>Lesson 4: DNS — Resolution, caching, and why changes take time</title><link>/post/fundamentals/net-dns/</link><pubDate>Sat, 15 Jun 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/net-dns/</guid><description>&lt;p&gt;We pushed a production incident fix at 3am — rotated to a new IP address for a critical service, updated the DNS record, set the TTL to 60 seconds. Thirty minutes later, half our users were still hitting the broken server. The other half were fine. We had updated the DNS record correctly. The TTL had expired. Yet somehow, stale answers were persisting. That night taught me more about DNS than any documentation had.&lt;/p&gt;</description></item><item><title>Lesson 1: HTTP Client Internals — Transport, connection reuse, and the pool you forgot to configure</title><link>/post/go/go-net-http-client/</link><pubDate>Fri, 14 Jun 2024 00:00:00 +0000</pubDate><guid>/post/go/go-net-http-client/</guid><description>&lt;p&gt;The first production Go service I wrote hammered our database proxy with thousands of new TCP connections every minute. The proxy started refusing connections. The on-call engineer thought it was the database. It wasn&amp;rsquo;t. It was me — specifically, my decision to create a new &lt;code&gt;http.Client&lt;/code&gt; on every request and then never read the response body to completion. I spent three hours debugging what turned out to be two lines of misunderstood standard library behavior.&lt;/p&gt;</description></item><item><title>Lesson 3: TLS Handshake — What happens in those 2 round trips</title><link>/post/fundamentals/net-tls/</link><pubDate>Thu, 30 May 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/net-tls/</guid><description>&lt;p&gt;A colleague once asked me why adding TLS to a service increased our P99 latency by 50ms. She had measured it carefully, switching between HTTP and HTTPS in a load test. My first instinct was to say &amp;ldquo;encryption overhead&amp;rdquo; but that&amp;rsquo;s wrong — modern CPUs with AES-NI can encrypt gigabytes per second. The actual cost is the handshake. Once I explained what was actually happening in those first few round trips, the answer to &amp;ldquo;how do we fix it&amp;rdquo; became obvious: stop creating new connections.&lt;/p&gt;</description></item><item><title>Lesson 2: HTTP/2 and HTTP/3 — Multiplexing and QUIC</title><link>/post/fundamentals/net-http2-http3/</link><pubDate>Tue, 14 May 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/net-http2-http3/</guid><description>&lt;p&gt;The first time I looked at a waterfall chart in Chrome DevTools and saw a hundred requests stacking up like rush-hour traffic, I knew HTTP/1.1 was the problem. We had a dashboard loading twelve API calls and a handful of assets, and they were all waiting in line. Not because the server was slow — it was idle — but because the browser had a six-connection-per-host limit and every request had to wait its turn. That was my introduction to why HTTP/2 exists, and it changed how I think about protocol design.&lt;/p&gt;</description></item><item><title>Lesson 1: TCP Deep Dive — Three-way handshake, congestion, Nagle</title><link>/post/fundamentals/net-tcp/</link><pubDate>Mon, 29 Apr 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/net-tcp/</guid><description>&lt;p&gt;I used to treat TCP as a black box. Data goes in one side, data comes out the other — reliably, in order, no duplicates. That was all I needed to know, right? Then I started debugging latency spikes in a payment service and spent three days chasing a 200ms tail latency that turned out to be Nagle&amp;rsquo;s algorithm fighting with delayed ACKs. After that, I stopped treating TCP as a black box.&lt;/p&gt;</description></item></channel></rss>