<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Rust Performance Engineering on Atharva Pandey</title><link>https://atharva.page/series/rust-performance-engineering/</link><description>Recent content in Rust Performance Engineering on Atharva Pandey</description><generator>Hugo</generator><language>en-us</language><copyright>Copyright ©</copyright><lastBuildDate>Sat, 12 Apr 2025 11:30:00 +0000</lastBuildDate><atom:link href="https://atharva.page/series/rust-performance-engineering/index.xml" rel="self" type="application/rss+xml"/><item><title>Lesson 12: Zero-Copy Parsing — bytes, nom, winnow</title><link>https://atharva.page/post/rust/rust-perf-zero-copy/</link><pubDate>Sat, 12 Apr 2025 11:30:00 +0000</pubDate><guid>https://atharva.page/post/rust/rust-perf-zero-copy/</guid><description>&lt;p&gt;I once had to parse 2GB of log files per hour on a machine with 4GB of RAM. The naive approach — read line, split into fields, store as &lt;code&gt;String&lt;/code&gt; — peaked at 6GB memory usage and fell over. The data was being duplicated everywhere: once in the read buffer, once in each split &lt;code&gt;String&lt;/code&gt;, once in the output struct. Three copies of every byte.&lt;/p&gt;
&lt;p&gt;The zero-copy version peaked at 2.1GB — basically the file size plus a thin layer of parsed references. Same output, same correctness, one-third the memory, twice the throughput. Zero-copy parsing is one of the most powerful techniques in Rust&amp;rsquo;s performance toolkit, and the ownership system makes it uniquely natural here.&lt;/p&gt;</description></item><item><title>Lesson 11: Binary Size Reduction — Smaller deployments</title><link>https://atharva.page/post/rust/rust-perf-binary-size/</link><pubDate>Mon, 07 Apr 2025 13:17:00 +0000</pubDate><guid>https://atharva.page/post/rust/rust-perf-binary-size/</guid><description>&lt;p&gt;I shipped a &amp;ldquo;hello world&amp;rdquo; Rust binary to a team once and they came back confused: &amp;ldquo;Why is this 4 megabytes?&amp;rdquo; Fair question. A C hello-world is 16KB. A Go hello-world is about 2MB. A default Rust hello-world with standard linking is 3-4MB.&lt;/p&gt;
&lt;p&gt;That 4MB isn&amp;rsquo;t wasted — it&amp;rsquo;s the Rust standard library, panic handling, formatting machinery, and debug symbols. But when you&amp;rsquo;re building container images, deploying to embedded devices, or targeting WebAssembly, every megabyte counts. Here&amp;rsquo;s how to cut Rust binaries down to size.&lt;/p&gt;</description></item><item><title>Lesson 10: Compile Time Optimization — Strategies that actually work</title><link>https://atharva.page/post/rust/rust-perf-compile-times/</link><pubDate>Thu, 03 Apr 2025 08:55:00 +0000</pubDate><guid>https://atharva.page/post/rust/rust-perf-compile-times/</guid><description>&lt;p&gt;My main Rust project at work took 4 minutes and 38 seconds for a clean build. That was two years ago. Today it takes 52 seconds. Same codebase — more code, actually. Same hardware. The difference is about a dozen targeted changes, none of which involved rewriting application code.&lt;/p&gt;
&lt;p&gt;Rust&amp;rsquo;s compile times are its most legitimate criticism. But &amp;ldquo;Rust is slow to compile&amp;rdquo; is the starting point, not the conclusion. Most projects can cut their build times by 50-80% with the right techniques. Let me show you what actually moves the needle.&lt;/p&gt;</description></item><item><title>Lesson 9: Inlining — #[inline] and LTO</title><link>https://atharva.page/post/rust/rust-perf-inlining/</link><pubDate>Mon, 31 Mar 2025 19:05:00 +0000</pubDate><guid>https://atharva.page/post/rust/rust-perf-inlining/</guid><description>&lt;p&gt;A few years ago I was profiling a JSON parser and noticed something weird. A tiny function — four lines, no allocations — was showing up as a 15% hot spot. Not because it was slow, but because it was called 8 million times per second and the function call overhead (push registers, set up stack frame, call, pop registers, return) was eating 2 nanoseconds each call. That&amp;rsquo;s 16 milliseconds per second just in function prologues and epilogues.&lt;/p&gt;</description></item><item><title>Lesson 8: Cache-Friendly Data Structures — Data-oriented design</title><link>https://atharva.page/post/rust/rust-perf-cache-friendly/</link><pubDate>Sat, 29 Mar 2025 07:40:00 +0000</pubDate><guid>https://atharva.page/post/rust/rust-perf-cache-friendly/</guid><description>&lt;p&gt;Here&amp;rsquo;s a number that should change how you think about data structures: reading from L1 cache takes about 1 nanosecond. Reading from main memory takes about 100 nanoseconds. That&amp;rsquo;s a 100x penalty for a cache miss. On a modern CPU running at 4 GHz, a single cache miss stalls the processor for roughly 400 cycles. Four hundred cycles where your CPU is sitting there, doing nothing, waiting for data to arrive from RAM.&lt;/p&gt;</description></item><item><title>Lesson 7: Choosing the Right Collection — It's not always Vec</title><link>https://atharva.page/post/rust/rust-perf-collections/</link><pubDate>Thu, 27 Mar 2025 15:22:00 +0000</pubDate><guid>https://atharva.page/post/rust/rust-perf-collections/</guid><description>&lt;p&gt;A friend asked me to review their service that was doing &amp;ldquo;thousands of lookups per second&amp;rdquo; against a &lt;code&gt;Vec&lt;/code&gt; of about 50,000 entries. Linear scan every time. They&amp;rsquo;d chosen &lt;code&gt;Vec&lt;/code&gt; because &amp;ldquo;it&amp;rsquo;s the default&amp;rdquo; and hadn&amp;rsquo;t thought about it further. Swapping to a &lt;code&gt;HashMap&lt;/code&gt; took the lookup from 12µs to 40ns. Three hundred times faster. From a one-line change.&lt;/p&gt;
&lt;p&gt;Choosing the right collection is one of the highest-leverage performance decisions you can make, and it requires basically zero cleverness. Just know your access patterns.&lt;/p&gt;</description></item><item><title>Lesson 6: String Performance — SmartString, CompactStr, and when to care</title><link>https://atharva.page/post/rust/rust-perf-string-perf/</link><pubDate>Tue, 25 Mar 2025 11:50:00 +0000</pubDate><guid>https://atharva.page/post/rust/rust-perf-string-perf/</guid><description>&lt;p&gt;I was building an in-memory index that stored about 2 million tag strings. Most were short — &amp;ldquo;rust&amp;rdquo;, &amp;ldquo;go&amp;rdquo;, &amp;ldquo;api&amp;rdquo;, &amp;ldquo;v2&amp;rdquo; — averaging 6 bytes. But each &lt;code&gt;String&lt;/code&gt; carries 24 bytes of overhead (pointer + length + capacity) plus the heap allocation for the actual data. That&amp;rsquo;s 24 bytes of bookkeeping to store 6 bytes of useful information. Plus 2 million separate allocations hammering the allocator.&lt;/p&gt;
&lt;p&gt;Switching to &lt;code&gt;CompactStr&lt;/code&gt; cut memory usage by 60% and index-building time by 40%. Strings matter more than you think.&lt;/p&gt;</description></item><item><title>Lesson 5: Iterators vs Loops — Performance characteristics</title><link>https://atharva.page/post/rust/rust-perf-iterators-vs-loops/</link><pubDate>Sun, 23 Mar 2025 09:12:00 +0000</pubDate><guid>https://atharva.page/post/rust/rust-perf-iterators-vs-loops/</guid><description>&lt;p&gt;When I first started writing Rust, I wrote everything as &lt;code&gt;for&lt;/code&gt; loops. Old habits from C. Then someone on my team rewrote one of my loops as an iterator chain and I got annoyed — it looked &amp;ldquo;slower&amp;rdquo; to me. More function calls, closures, chaining. Obviously that&amp;rsquo;s more overhead, right?&lt;/p&gt;
&lt;p&gt;I benchmarked it. Same performance. Down to the nanosecond. I looked at the assembly. Identical. That was the day I stopped assuming and started measuring.&lt;/p&gt;</description></item><item><title>Lesson 4: Reducing Allocations — Stack, arena, SmallVec</title><link>https://atharva.page/post/rust/rust-perf-allocations/</link><pubDate>Fri, 21 Mar 2025 16:30:00 +0000</pubDate><guid>https://atharva.page/post/rust/rust-perf-allocations/</guid><description>&lt;p&gt;I profiled a Rust web service once and found it was allocating 47,000 times per request. Forty-seven thousand. Most were tiny — 16-byte strings, 3-element vectors, temporary buffers. Each individual allocation was fast (jemalloc is good), but 47,000 of them at ~30ns each is 1.4ms of pure allocator overhead. Per request. At 10K RPS that&amp;rsquo;s 14 seconds of CPU time per second, just asking the allocator for memory.&lt;/p&gt;
&lt;p&gt;The fix took half a day and cut allocations to about 200 per request. Here&amp;rsquo;s everything I know about reducing allocations in Rust.&lt;/p&gt;</description></item><item><title>Lesson 3: Profiling — perf, flamegraph, samply</title><link>https://atharva.page/post/rust/rust-perf-profiling/</link><pubDate>Wed, 19 Mar 2025 10:45:00 +0000</pubDate><guid>https://atharva.page/post/rust/rust-perf-profiling/</guid><description>&lt;p&gt;A colleague once asked me to look at a Rust service that was &amp;ldquo;slow.&amp;rdquo; They&amp;rsquo;d already spent a week trying to optimize the JSON parsing layer because &amp;ldquo;parsing is always the bottleneck.&amp;rdquo; I ran a profiler. Sixty-three percent of CPU time was spent in &lt;code&gt;Drop&lt;/code&gt; implementations, deallocating thousands of small strings that were created and immediately discarded. The JSON parsing was 4% of runtime.&lt;/p&gt;
&lt;p&gt;Profiling would&amp;rsquo;ve found that in five minutes. That&amp;rsquo;s why this lesson exists.&lt;/p&gt;</description></item><item><title>Lesson 2: Benchmarking with criterion and divan — Statistically rigorous benchmarks</title><link>https://atharva.page/post/rust/rust-perf-benchmarking/</link><pubDate>Mon, 17 Mar 2025 14:18:00 +0000</pubDate><guid>https://atharva.page/post/rust/rust-perf-benchmarking/</guid><description>&lt;p&gt;Last year I reviewed a PR where someone claimed their new serialization code was &amp;ldquo;2x faster.&amp;rdquo; Their benchmark? &lt;code&gt;std::time::Instant::now()&lt;/code&gt; called once before and once after. Single run. No warmup. No statistical analysis. The &amp;ldquo;2x speedup&amp;rdquo; was thermal throttling on the first run.&lt;/p&gt;
&lt;p&gt;Benchmarking is harder than it looks. Let&amp;rsquo;s do it properly.&lt;/p&gt;
&lt;h2 id="why-naive-benchmarks-lie"&gt;Why Naive Benchmarks Lie&lt;/h2&gt;
&lt;p&gt;Before we get into the tools, let me show you all the ways a naive benchmark can mislead you:&lt;/p&gt;</description></item><item><title>Lesson 1: Performance Philosophy — Measure, don't guess</title><link>https://atharva.page/post/rust/rust-perf-philosophy/</link><pubDate>Sat, 15 Mar 2025 08:32:00 +0000</pubDate><guid>https://atharva.page/post/rust/rust-perf-philosophy/</guid><description>&lt;p&gt;I once spent three days rewriting a hot loop to avoid a single allocation per iteration. Hand-rolled a custom arena, eliminated two clones, even switched from &lt;code&gt;HashMap&lt;/code&gt; to a hand-tuned open-addressing table. Benchmarked the result: 0.3% improvement. The actual bottleneck? A DNS lookup buried in a library call that I never bothered to profile.&lt;/p&gt;
&lt;p&gt;Three days. Zero meaningful impact. That&amp;rsquo;s the lesson I want to start this entire course with.&lt;/p&gt;</description></item></channel></rss>