<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Rust Concurrency Masterclass on Atharva Pandey</title><link>https://atharva.page/series/rust-concurrency-masterclass/</link><description>Recent content in Rust Concurrency Masterclass on Atharva Pandey</description><generator>Hugo</generator><language>en-us</language><copyright>Copyright ©</copyright><lastBuildDate>Sun, 22 Dec 2024 10:00:00 +0000</lastBuildDate><atom:link href="https://atharva.page/series/rust-concurrency-masterclass/index.xml" rel="self" type="application/rss+xml"/><item><title>Lesson 25: Production Concurrency Architecture — Putting it all together</title><link>https://atharva.page/post/rust/rust-conc-production-patterns/</link><pubDate>Sun, 22 Dec 2024 10:00:00 +0000</pubDate><guid>https://atharva.page/post/rust/rust-conc-production-patterns/</guid><description>&lt;p&gt;After 24 lessons of building blocks, let&amp;rsquo;s talk about how they compose in real systems. I&amp;rsquo;ve shipped concurrent Rust services handling millions of requests per day, and the architecture patterns that survive production are surprisingly consistent. Not because they&amp;rsquo;re clever — because they&amp;rsquo;re boring. Boring is good when your pager is involved.&lt;/p&gt;
&lt;p&gt;This lesson is the blueprint I wish I&amp;rsquo;d had when I started building concurrent Rust systems for real.&lt;/p&gt;</description></item><item><title>Lesson 24: Testing Concurrent Code — Loom and beyond</title><link>https://atharva.page/post/rust/rust-conc-testing/</link><pubDate>Sat, 21 Dec 2024 09:15:00 +0000</pubDate><guid>https://atharva.page/post/rust/rust-conc-testing/</guid><description>&lt;p&gt;I spent a full week writing tests for a lock-free queue. Ran them a thousand times — all green. Shipped it. Two days later, a production crash. A race condition that occurred roughly once every 50,000 operations under specific timing. My tests never hit it because standard testing can&amp;rsquo;t explore all possible thread interleavings. That&amp;rsquo;s when I found Loom.&lt;/p&gt;
&lt;p&gt;Testing concurrent code is fundamentally different from testing sequential code. A test that passes doesn&amp;rsquo;t mean the code is correct — it means the code was correct &lt;em&gt;for that particular thread scheduling&lt;/em&gt;. Run the same test with different timing and you might get a different result.&lt;/p&gt;</description></item><item><title>Lesson 23: GPU Computing from Rust — wgpu and compute shaders</title><link>https://atharva.page/post/rust/rust-conc-gpu/</link><pubDate>Thu, 19 Dec 2024 11:35:00 +0000</pubDate><guid>https://atharva.page/post/rust/rust-conc-gpu/</guid><description>&lt;p&gt;The first time I ran a matrix multiplication on a GPU, my jaw dropped. A computation that took 8 seconds on an 8-core CPU finished in 40 milliseconds on a mid-range GPU. That&amp;rsquo;s a 200x speedup. GPUs have thousands of cores — small, simple cores designed for massively parallel uniform computation.&lt;/p&gt;
&lt;p&gt;Rust&amp;rsquo;s GPU story has gotten remarkably good. &lt;code&gt;wgpu&lt;/code&gt; gives you cross-platform GPU compute that works on Vulkan, Metal, DX12, and even in the browser via WebGPU. No CUDA lock-in.&lt;/p&gt;</description></item><item><title>Lesson 22: SIMD — Explicit vectorization</title><link>https://atharva.page/post/rust/rust-conc-simd/</link><pubDate>Tue, 17 Dec 2024 08:20:00 +0000</pubDate><guid>https://atharva.page/post/rust/rust-conc-simd/</guid><description>&lt;p&gt;I had an image processing pipeline that took 4.2 seconds per frame on a single core. After rewriting the hot loop with SIMD intrinsics, it dropped to 0.9 seconds. Same core, same algorithm, 4.7x faster. SIMD doesn&amp;rsquo;t add more cores — it makes each core do more work per clock cycle.&lt;/p&gt;
&lt;p&gt;SIMD (Single Instruction, Multiple Data) is the other kind of parallelism. While threads run code on different cores, SIMD runs the same operation on multiple data elements simultaneously within a single core.&lt;/p&gt;</description></item><item><title>Lesson 21: CSP-Style Concurrency — Go channels in Rust</title><link>https://atharva.page/post/rust/rust-conc-csp/</link><pubDate>Sun, 15 Dec 2024 10:40:00 +0000</pubDate><guid>https://atharva.page/post/rust/rust-conc-csp/</guid><description>&lt;p&gt;Before I wrote Rust full-time, I spent two years writing Go. The goroutine-plus-channel model gets into your brain. You start thinking about problems as independent processes connected by typed pipes. When I switched to Rust, the first thing I looked for was the equivalent of Go&amp;rsquo;s &lt;code&gt;select&lt;/code&gt; statement. Crossbeam has it — and in some ways it&amp;rsquo;s even better.&lt;/p&gt;
&lt;p&gt;CSP (Communicating Sequential Processes) is the formal model behind Go&amp;rsquo;s concurrency. The idea: concurrent processes interact only by passing messages through channels. No shared memory. Each process is sequential internally. Concurrency comes from composition.&lt;/p&gt;</description></item><item><title>Lesson 20: The Actor Model with Rust — Message-passing architectures</title><link>https://atharva.page/post/rust/rust-conc-actor-model/</link><pubDate>Fri, 13 Dec 2024 14:20:00 +0000</pubDate><guid>https://atharva.page/post/rust/rust-conc-actor-model/</guid><description>&lt;p&gt;The first time I saw Erlang&amp;rsquo;s actor model, I thought it was over-engineered. Every piece of state behind a process, every interaction a message. Then I worked on a system with 40 mutexes, 12 deadlock-prone code paths, and a debugging story that involved printf-ing thread IDs into a file and diffing them. I became an actor model convert that week.&lt;/p&gt;
&lt;p&gt;Rust doesn&amp;rsquo;t have a built-in actor system like Erlang or Akka. But the building blocks — channels, threads, ownership transfer — make building one surprisingly natural.&lt;/p&gt;</description></item><item><title>Lesson 19: Thread-Local Storage — Per-thread state</title><link>https://atharva.page/post/rust/rust-conc-thread-local/</link><pubDate>Wed, 11 Dec 2024 09:55:00 +0000</pubDate><guid>https://atharva.page/post/rust/rust-conc-thread-local/</guid><description>&lt;p&gt;I was optimizing a JSON serializer that allocated a buffer for every call. Under profiling, those allocations were 30% of the cost. The fix? A thread-local buffer that gets reused across calls on the same thread. No synchronization needed. No contention. Each thread has its own buffer. Throughput doubled.&lt;/p&gt;
&lt;p&gt;Thread-local storage is the ultimate escape hatch from synchronization overhead. If each thread has its own copy, there&amp;rsquo;s nothing to synchronize.&lt;/p&gt;</description></item><item><title>Lesson 18: Debugging Deadlocks and Data Races — Tools and techniques</title><link>https://atharva.page/post/rust/rust-conc-deadlocks/</link><pubDate>Mon, 09 Dec 2024 11:30:00 +0000</pubDate><guid>https://atharva.page/post/rust/rust-conc-deadlocks/</guid><description>&lt;p&gt;The worst deadlock I ever encountered wasn&amp;rsquo;t between two mutexes. It was between a mutex and a channel. Thread A held a lock and tried to send on a full bounded channel. Thread B was the consumer for that channel but was waiting to acquire the same lock before it could receive. Clean deadlock. Took me four hours to find because I was looking at mutex ordering and the real problem was a channel.&lt;/p&gt;</description></item><item><title>Lesson 17: Lock-Free Data Structures in Rust — Beyond mutexes</title><link>https://atharva.page/post/rust/rust-conc-lock-free/</link><pubDate>Sat, 07 Dec 2024 08:50:00 +0000</pubDate><guid>https://atharva.page/post/rust/rust-conc-lock-free/</guid><description>&lt;p&gt;A few years ago I was profiling a metrics collection service and found that 60% of the CPU time was spent on mutex contention. Thirty-two threads, one mutex protecting a counter map. The actual &amp;ldquo;work&amp;rdquo; — incrementing counters — took nanoseconds. The locking overhead was three orders of magnitude more expensive than the operation it protected.&lt;/p&gt;
&lt;p&gt;That&amp;rsquo;s when lock-free data structures become worth the complexity. When the lock is the bottleneck, remove the lock.&lt;/p&gt;</description></item><item><title>Lesson 16: Barriers and Once — Synchronization primitives</title><link>https://atharva.page/post/rust/rust-conc-barriers/</link><pubDate>Thu, 05 Dec 2024 10:15:00 +0000</pubDate><guid>https://atharva.page/post/rust/rust-conc-barriers/</guid><description>&lt;p&gt;I was building a benchmark suite once — eight threads, each measuring throughput of a different operation. The problem was that threads started at different times depending on OS scheduling. Thread 0 might start 50ms before thread 7, skewing the results. I needed all threads to start their measurement at exactly the same time.&lt;/p&gt;
&lt;p&gt;That&amp;rsquo;s what barriers do. Everyone waits at the barrier until the last thread arrives, then they all proceed together.&lt;/p&gt;</description></item><item><title>Lesson 15: Condvar — Waiting for conditions</title><link>https://atharva.page/post/rust/rust-conc-condvar/</link><pubDate>Tue, 03 Dec 2024 14:40:00 +0000</pubDate><guid>https://atharva.page/post/rust/rust-conc-condvar/</guid><description>&lt;p&gt;Early in my career, I wrote a producer-consumer queue using a mutex and a busy-wait loop. The consumer would lock the mutex, check if there&amp;rsquo;s data, unlock, sleep for 10 milliseconds, and repeat. It worked, but it burned CPU doing nothing and had up to 10ms of latency on every message. My tech lead pointed me to condition variables, and the latency dropped to microseconds while CPU usage went to near zero.&lt;/p&gt;</description></item><item><title>Lesson 14: parking_lot — Faster mutexes</title><link>https://atharva.page/post/rust/rust-conc-parking-lot/</link><pubDate>Sun, 01 Dec 2024 09:25:00 +0000</pubDate><guid>https://atharva.page/post/rust/rust-conc-parking-lot/</guid><description>&lt;p&gt;I switched a high-contention service from &lt;code&gt;std::sync::Mutex&lt;/code&gt; to &lt;code&gt;parking_lot::Mutex&lt;/code&gt; and saw lock acquisition time drop by 30% under load. The API is nearly identical — it was a find-and-replace job. That&amp;rsquo;s the kind of optimization I like: massive payoff, zero complexity cost.&lt;/p&gt;
&lt;p&gt;&lt;code&gt;parking_lot&lt;/code&gt; is one of those crates that probably should have been in the standard library. It provides the same synchronization primitives as std, but faster, smaller, and with more features.&lt;/p&gt;</description></item><item><title>Lesson 13: Fan-Out Fan-In in Rust — Parallel pipelines</title><link>https://atharva.page/post/rust/rust-conc-fan-out-fan-in/</link><pubDate>Fri, 29 Nov 2024 12:50:00 +0000</pubDate><guid>https://atharva.page/post/rust/rust-conc-fan-out-fan-in/</guid><description>&lt;p&gt;One of the most satisfying architectures I&amp;rsquo;ve built was a real-time analytics pipeline. Events came in through a single ingestion point, fanned out to eight processing workers, then fanned back in to a single aggregator that wrote results to the database. Throughput went from 2,000 events/second to 14,000 with zero data loss.&lt;/p&gt;
&lt;p&gt;Fan-out/fan-in is one of those patterns that looks simple on a whiteboard but has real subtlety in implementation. Getting the channel lifecycle right, handling errors, managing backpressure — that&amp;rsquo;s where things get interesting.&lt;/p&gt;</description></item><item><title>Lesson 12: Worker Pool Patterns — Bounded concurrency</title><link>https://atharva.page/post/rust/rust-conc-worker-pools/</link><pubDate>Wed, 27 Nov 2024 08:35:00 +0000</pubDate><guid>https://atharva.page/post/rust/rust-conc-worker-pools/</guid><description>&lt;p&gt;I blew up a production database once by spawning unlimited concurrent connections. A batch job that normally processed 100 items suddenly got 50,000. Each item opened a database connection. The connection pool maxed out, then the OS ran out of file descriptors. The database server stopped accepting connections from any service. All because I wrote &lt;code&gt;for item in items { thread::spawn(...) }&lt;/code&gt; without thinking about bounds.&lt;/p&gt;
&lt;p&gt;Worker pools solve this. Fixed number of threads, bounded queue, backpressure when the system is overloaded. After that incident, I never write unbounded concurrency again.&lt;/p&gt;</description></item><item><title>Lesson 11: Crossbeam — Scoped threads and lock-free structures</title><link>https://atharva.page/post/rust/rust-conc-crossbeam/</link><pubDate>Mon, 25 Nov 2024 11:05:00 +0000</pubDate><guid>https://atharva.page/post/rust/rust-conc-crossbeam/</guid><description>&lt;p&gt;Before Rust 1.63 added &lt;code&gt;thread::scope&lt;/code&gt; to the standard library, crossbeam was the only ergonomic way to spawn threads that could borrow local data. Even now that std has scoped threads, crossbeam remains essential. Its channels are faster than &lt;code&gt;std::sync::mpsc&lt;/code&gt;, it provides lock-free data structures, and its utilities fill gaps the standard library doesn&amp;rsquo;t cover.&lt;/p&gt;
&lt;p&gt;I reach for crossbeam in nearly every concurrent Rust project. Here&amp;rsquo;s why.&lt;/p&gt;
&lt;h2 id="crossbeam-channels"&gt;Crossbeam Channels&lt;/h2&gt;
&lt;p&gt;The biggest win: crossbeam&amp;rsquo;s channels are multi-producer, multi-consumer (MPMC) and significantly faster than &lt;code&gt;std::sync::mpsc&lt;/code&gt;.&lt;/p&gt;</description></item><item><title>Lesson 10: Rayon — Data parallelism made easy</title><link>https://atharva.page/post/rust/rust-conc-rayon/</link><pubDate>Sat, 23 Nov 2024 15:20:00 +0000</pubDate><guid>https://atharva.page/post/rust/rust-conc-rayon/</guid><description>&lt;p&gt;I had a batch processing job — reading 50,000 JSON files, parsing them, running validation, writing results. Single-threaded, it took 12 minutes. I added Rayon, changed &lt;code&gt;.iter()&lt;/code&gt; to &lt;code&gt;.par_iter()&lt;/code&gt;, and it dropped to 90 seconds. Five characters added to my code. That&amp;rsquo;s it.&lt;/p&gt;
&lt;p&gt;Rayon is probably the most impressive crate in the Rust ecosystem for the effort-to-impact ratio. If you have CPU-bound work that operates on collections, Rayon makes parallelism trivial.&lt;/p&gt;</description></item><item><title>Lesson 9: Send and Sync — The traits behind thread safety</title><link>https://atharva.page/post/rust/rust-conc-send-sync/</link><pubDate>Thu, 21 Nov 2024 09:40:00 +0000</pubDate><guid>https://atharva.page/post/rust/rust-conc-send-sync/</guid><description>&lt;p&gt;There&amp;rsquo;s a moment in every Rust developer&amp;rsquo;s life when they try to send an &lt;code&gt;Rc&amp;lt;RefCell&amp;lt;Vec&amp;lt;String&amp;gt;&amp;gt;&amp;gt;&lt;/code&gt; to another thread and get a wall of compiler errors. They Google the error, see something about &lt;code&gt;Send&lt;/code&gt; and &lt;code&gt;Sync&lt;/code&gt;, patch the type to &lt;code&gt;Arc&amp;lt;Mutex&amp;lt;Vec&amp;lt;String&amp;gt;&amp;gt;&amp;gt;&lt;/code&gt;, and move on without understanding why.&lt;/p&gt;
&lt;p&gt;I did exactly that for six months. Then I needed to write a custom type that crossed thread boundaries, and I had to actually learn what these traits mean. Turns out they&amp;rsquo;re beautifully simple once you see the design.&lt;/p&gt;</description></item><item><title>Lesson 8: Memory Ordering — Relaxed, Acquire, Release, SeqCst</title><link>https://atharva.page/post/rust/rust-conc-ordering/</link><pubDate>Tue, 19 Nov 2024 13:15:00 +0000</pubDate><guid>https://atharva.page/post/rust/rust-conc-ordering/</guid><description>&lt;p&gt;Memory ordering is the thing that separates people who &lt;em&gt;use&lt;/em&gt; concurrent code from people who &lt;em&gt;write&lt;/em&gt; concurrent primitives. I avoided understanding it for years, using &lt;code&gt;SeqCst&lt;/code&gt; everywhere like a safety blanket. It worked. But it left performance on the table and, more importantly, left me unable to read half the lock-free code I encountered.&lt;/p&gt;
&lt;p&gt;So here&amp;rsquo;s the actual explanation — no hand-waving.&lt;/p&gt;
&lt;h2 id="why-ordering-matters"&gt;Why Ordering Matters&lt;/h2&gt;
&lt;p&gt;Modern CPUs and compilers reorder instructions for performance. Your code says &amp;ldquo;write A, then write B,&amp;rdquo; but the CPU might execute B first if there&amp;rsquo;s no dependency between them. On a single thread, this is invisible — the final result is the same.&lt;/p&gt;</description></item><item><title>Lesson 7: Atomics — Lock-free primitives</title><link>https://atharva.page/post/rust/rust-conc-atomics/</link><pubDate>Sun, 17 Nov 2024 07:45:00 +0000</pubDate><guid>https://atharva.page/post/rust/rust-conc-atomics/</guid><description>&lt;p&gt;I once replaced a &lt;code&gt;Mutex&amp;lt;u64&amp;gt;&lt;/code&gt; counter in a hot path with an &lt;code&gt;AtomicU64&lt;/code&gt; and saw throughput jump 40%. Not because mutexes are slow — they&amp;rsquo;re fast. But for a single integer being incremented by 32 threads, the overhead of acquiring and releasing a lock millions of times per second adds up to real time.&lt;/p&gt;
&lt;p&gt;Atomics are the foundation of lock-free programming. They let you do thread-safe operations on primitive values without any lock at all.&lt;/p&gt;</description></item><item><title>Lesson 6: Arc&lt;Mutex&lt;T&gt;&gt; — The shared mutable state pattern</title><link>https://atharva.page/post/rust/rust-conc-arc-mutex/</link><pubDate>Fri, 15 Nov 2024 10:30:00 +0000</pubDate><guid>https://atharva.page/post/rust/rust-conc-arc-mutex/</guid><description>&lt;p&gt;I remember staring at &lt;code&gt;Arc&amp;lt;Mutex&amp;lt;HashMap&amp;lt;String, Vec&amp;lt;u8&amp;gt;&amp;gt;&amp;gt;&amp;gt;&lt;/code&gt; in a codebase and thinking &amp;ldquo;this is the ugliest type I&amp;rsquo;ve ever seen.&amp;rdquo; Three months later, after debugging a similar system in Go that had zero type safety around its concurrent map access, I came crawling back to Rust&amp;rsquo;s ugly-but-correct approach.&lt;/p&gt;
&lt;p&gt;The &lt;code&gt;Arc&amp;lt;Mutex&amp;lt;T&amp;gt;&amp;gt;&lt;/code&gt; pattern is everywhere in Rust concurrent code. Understanding &lt;em&gt;why&lt;/em&gt; it exists — not just how to type it — is the key.&lt;/p&gt;</description></item><item><title>Lesson 5: Shared State — Mutex, RwLock, and poisoning</title><link>https://atharva.page/post/rust/rust-conc-shared-state/</link><pubDate>Wed, 13 Nov 2024 16:55:00 +0000</pubDate><guid>https://atharva.page/post/rust/rust-conc-shared-state/</guid><description>&lt;p&gt;The nastiest production bug I ever tracked down involved a Java ConcurrentHashMap that was &amp;ldquo;thread-safe&amp;rdquo; in the API sense but not in the logic sense. Two threads would read a key, both see it&amp;rsquo;s absent, both insert with computed values, and one would silently overwrite the other. The map itself was fine — the &lt;em&gt;access pattern&lt;/em&gt; was broken.&lt;/p&gt;
&lt;p&gt;Rust&amp;rsquo;s Mutex won&amp;rsquo;t save you from logic bugs either. But it will absolutely prevent you from forgetting the lock in the first place.&lt;/p&gt;</description></item><item><title>Lesson 4: Channels — mpsc and beyond</title><link>https://atharva.page/post/rust/rust-conc-channels/</link><pubDate>Mon, 11 Nov 2024 09:10:00 +0000</pubDate><guid>https://atharva.page/post/rust/rust-conc-channels/</guid><description>&lt;p&gt;The first concurrent system I built that actually worked well was a log aggregation pipeline. Multiple producers writing log lines, one consumer batching and flushing to disk. No shared state, no locks, no races. Just messages flowing through a pipe.&lt;/p&gt;
&lt;p&gt;That experience sold me on message passing. And Rust&amp;rsquo;s channel implementation makes it surprisingly ergonomic.&lt;/p&gt;
&lt;h2 id="the-problem-shared-state-is-hard"&gt;The Problem: Shared State Is Hard&lt;/h2&gt;
&lt;p&gt;You &lt;em&gt;can&lt;/em&gt; share state between threads with mutexes. But every shared mutable variable is a coordination point, a potential bottleneck, and a source of bugs. The more threads touching the same data, the harder the code is to reason about.&lt;/p&gt;</description></item><item><title>Lesson 3: move Closures — Sending data to threads</title><link>https://atharva.page/post/rust/rust-conc-move-closures/</link><pubDate>Sat, 09 Nov 2024 14:20:00 +0000</pubDate><guid>https://atharva.page/post/rust/rust-conc-move-closures/</guid><description>&lt;p&gt;When I first started writing threaded Rust code, I hit the same compiler error about forty times in one afternoon. Something about closures borrowing values that might be dropped. I kept slapping &lt;code&gt;move&lt;/code&gt; on closures until things compiled, without really understanding what was happening.&lt;/p&gt;
&lt;p&gt;That&amp;rsquo;s a terrible way to learn. So here&amp;rsquo;s the actual explanation I wish I&amp;rsquo;d had.&lt;/p&gt;
&lt;h2 id="the-problem-closures-and-thread-lifetimes"&gt;The Problem: Closures and Thread Lifetimes&lt;/h2&gt;
&lt;p&gt;When you pass a closure to &lt;code&gt;thread::spawn&lt;/code&gt;, the new thread might run for an arbitrary amount of time. It could outlive the function that spawned it. It could outlive the variables it references.&lt;/p&gt;</description></item><item><title>Lesson 2: std::thread — Spawning and joining</title><link>https://atharva.page/post/rust/rust-conc-threads/</link><pubDate>Thu, 07 Nov 2024 11:45:00 +0000</pubDate><guid>https://atharva.page/post/rust/rust-conc-threads/</guid><description>&lt;p&gt;A few years back I was reviewing a Go service that spawned goroutines like confetti at a parade — hundreds of them, no tracking, no lifecycle management. When the service shut down, half those goroutines just vanished mid-work. Orphaned database connections everywhere. The Go runtime made it &lt;em&gt;so easy&lt;/em&gt; to fire off concurrent work that nobody stopped to think about cleanup.&lt;/p&gt;
&lt;p&gt;Rust&amp;rsquo;s threading model forces you to think about it. Every thread gives you a handle. Every handle demands acknowledgment. You can ignore it — but you have to do so explicitly.&lt;/p&gt;</description></item><item><title>Lesson 1: Why Rust Concurrency Is "Fearless" — The compiler has your back</title><link>https://atharva.page/post/rust/rust-conc-why-fearless/</link><pubDate>Tue, 05 Nov 2024 08:30:00 +0000</pubDate><guid>https://atharva.page/post/rust/rust-conc-why-fearless/</guid><description>&lt;p&gt;I once spent three days chasing a race condition in a Java service that only manifested under production load. The bug? Two threads updating a shared HashMap — no synchronization, no errors at compile time, no warnings. Just silent data corruption that showed up as incorrect billing amounts. Three days of my life, gone, because the language didn&amp;rsquo;t care.&lt;/p&gt;
&lt;p&gt;Then I tried to write the same bug in Rust. The compiler said no.&lt;/p&gt;</description></item></channel></rss>