<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Performance on</title><link>/tags/performance/</link><description>Recent content in Performance on</description><generator>Hugo</generator><language>en</language><lastBuildDate>Sat, 12 Apr 2025 11:30:00 +0000</lastBuildDate><atom:link href="/tags/performance/index.xml" rel="self" type="application/rss+xml"/><item><title>Lesson 12: Zero-Copy Parsing — bytes, nom, winnow</title><link>/post/rust/rust-perf-zero-copy/</link><pubDate>Sat, 12 Apr 2025 11:30:00 +0000</pubDate><guid>/post/rust/rust-perf-zero-copy/</guid><description>&lt;p&gt;I once had to parse 2GB of log files per hour on a machine with 4GB of RAM. The naive approach — read line, split into fields, store as &lt;code&gt;String&lt;/code&gt; — peaked at 6GB memory usage and fell over. The data was being duplicated everywhere: once in the read buffer, once in each split &lt;code&gt;String&lt;/code&gt;, once in the output struct. Three copies of every byte.&lt;/p&gt;
&lt;p&gt;The zero-copy version peaked at 2.1GB — basically the file size plus a thin layer of parsed references. Same output, same correctness, one-third the memory, twice the throughput. Zero-copy parsing is one of the most powerful techniques in Rust&amp;rsquo;s performance toolkit, and the ownership system makes it uniquely natural here.&lt;/p&gt;</description></item><item><title>Lesson 8: Avoiding Premature Optimization — Measure first, optimize never (usually)</title><link>/post/go/go-perf-premature-optimization/</link><pubDate>Thu, 10 Apr 2025 00:00:00 +0000</pubDate><guid>/post/go/go-perf-premature-optimization/</guid><description>&lt;p&gt;I&amp;rsquo;ve spent time in this series showing you how to make Go programs faster — escape analysis, stack allocation, pre-sizing data structures, zero-copy string handling, benchmarking discipline, pprof profiling, CPU vs memory tradeoffs. Every technique is real and useful. And every single one of them has been misapplied, including by me, when applied before understanding whether they were needed. The last lesson isn&amp;rsquo;t another technique. It&amp;rsquo;s the discipline that makes all the other techniques worth using: measure first, understand where the actual cost is, and optimize only there.&lt;/p&gt;</description></item><item><title>Lesson 11: Binary Size Reduction — Smaller deployments</title><link>/post/rust/rust-perf-binary-size/</link><pubDate>Mon, 07 Apr 2025 13:17:00 +0000</pubDate><guid>/post/rust/rust-perf-binary-size/</guid><description>&lt;p&gt;I shipped a &amp;ldquo;hello world&amp;rdquo; Rust binary to a team once and they came back confused: &amp;ldquo;Why is this 4 megabytes?&amp;rdquo; Fair question. A C hello-world is 16KB. A Go hello-world is about 2MB. A default Rust hello-world with standard linking is 3-4MB.&lt;/p&gt;
&lt;p&gt;That 4MB isn&amp;rsquo;t wasted — it&amp;rsquo;s the Rust standard library, panic handling, formatting machinery, and debug symbols. But when you&amp;rsquo;re building container images, deploying to embedded devices, or targeting WebAssembly, every megabyte counts. Here&amp;rsquo;s how to cut Rust binaries down to size.&lt;/p&gt;</description></item><item><title>Lesson 10: Compile Time Optimization — Strategies that actually work</title><link>/post/rust/rust-perf-compile-times/</link><pubDate>Thu, 03 Apr 2025 08:55:00 +0000</pubDate><guid>/post/rust/rust-perf-compile-times/</guid><description>&lt;p&gt;My main Rust project at work took 4 minutes and 38 seconds for a clean build. That was two years ago. Today it takes 52 seconds. Same codebase — more code, actually. Same hardware. The difference is about a dozen targeted changes, none of which involved rewriting application code.&lt;/p&gt;
&lt;p&gt;Rust&amp;rsquo;s compile times are its most legitimate criticism. But &amp;ldquo;Rust is slow to compile&amp;rdquo; is the starting point, not the conclusion. Most projects can cut their build times by 50-80% with the right techniques. Let me show you what actually moves the needle.&lt;/p&gt;</description></item><item><title>Lesson 9: Inlining — #[inline] and LTO</title><link>/post/rust/rust-perf-inlining/</link><pubDate>Mon, 31 Mar 2025 19:05:00 +0000</pubDate><guid>/post/rust/rust-perf-inlining/</guid><description>&lt;p&gt;A few years ago I was profiling a JSON parser and noticed something weird. A tiny function — four lines, no allocations — was showing up as a 15% hot spot. Not because it was slow, but because it was called 8 million times per second and the function call overhead (push registers, set up stack frame, call, pop registers, return) was eating 2 nanoseconds each call. That&amp;rsquo;s 16 milliseconds per second just in function prologues and epilogues.&lt;/p&gt;</description></item><item><title>Lesson 8: Cache-Friendly Data Structures — Data-oriented design</title><link>/post/rust/rust-perf-cache-friendly/</link><pubDate>Sat, 29 Mar 2025 07:40:00 +0000</pubDate><guid>/post/rust/rust-perf-cache-friendly/</guid><description>&lt;p&gt;Here&amp;rsquo;s a number that should change how you think about data structures: reading from L1 cache takes about 1 nanosecond. Reading from main memory takes about 100 nanoseconds. That&amp;rsquo;s a 100x penalty for a cache miss. On a modern CPU running at 4 GHz, a single cache miss stalls the processor for roughly 400 cycles. Four hundred cycles where your CPU is sitting there, doing nothing, waiting for data to arrive from RAM.&lt;/p&gt;</description></item><item><title>Lesson 7: Choosing the Right Collection — It's not always Vec</title><link>/post/rust/rust-perf-collections/</link><pubDate>Thu, 27 Mar 2025 15:22:00 +0000</pubDate><guid>/post/rust/rust-perf-collections/</guid><description>&lt;p&gt;A friend asked me to review their service that was doing &amp;ldquo;thousands of lookups per second&amp;rdquo; against a &lt;code&gt;Vec&lt;/code&gt; of about 50,000 entries. Linear scan every time. They&amp;rsquo;d chosen &lt;code&gt;Vec&lt;/code&gt; because &amp;ldquo;it&amp;rsquo;s the default&amp;rdquo; and hadn&amp;rsquo;t thought about it further. Swapping to a &lt;code&gt;HashMap&lt;/code&gt; took the lookup from 12µs to 40ns. Three hundred times faster. From a one-line change.&lt;/p&gt;
&lt;p&gt;Choosing the right collection is one of the highest-leverage performance decisions you can make, and it requires basically zero cleverness. Just know your access patterns.&lt;/p&gt;</description></item><item><title>Lesson 6: String Performance — SmartString, CompactStr, and when to care</title><link>/post/rust/rust-perf-string-perf/</link><pubDate>Tue, 25 Mar 2025 11:50:00 +0000</pubDate><guid>/post/rust/rust-perf-string-perf/</guid><description>&lt;p&gt;I was building an in-memory index that stored about 2 million tag strings. Most were short — &amp;ldquo;rust&amp;rdquo;, &amp;ldquo;go&amp;rdquo;, &amp;ldquo;api&amp;rdquo;, &amp;ldquo;v2&amp;rdquo; — averaging 6 bytes. But each &lt;code&gt;String&lt;/code&gt; carries 24 bytes of overhead (pointer + length + capacity) plus the heap allocation for the actual data. That&amp;rsquo;s 24 bytes of bookkeeping to store 6 bytes of useful information. Plus 2 million separate allocations hammering the allocator.&lt;/p&gt;
&lt;p&gt;Switching to &lt;code&gt;CompactStr&lt;/code&gt; cut memory usage by 60% and index-building time by 40%. Strings matter more than you think.&lt;/p&gt;</description></item><item><title>Lesson 5: Iterators vs Loops — Performance characteristics</title><link>/post/rust/rust-perf-iterators-vs-loops/</link><pubDate>Sun, 23 Mar 2025 09:12:00 +0000</pubDate><guid>/post/rust/rust-perf-iterators-vs-loops/</guid><description>&lt;p&gt;When I first started writing Rust, I wrote everything as &lt;code&gt;for&lt;/code&gt; loops. Old habits from C. Then someone on my team rewrote one of my loops as an iterator chain and I got annoyed — it looked &amp;ldquo;slower&amp;rdquo; to me. More function calls, closures, chaining. Obviously that&amp;rsquo;s more overhead, right?&lt;/p&gt;
&lt;p&gt;I benchmarked it. Same performance. Down to the nanosecond. I looked at the assembly. Identical. That was the day I stopped assuming and started measuring.&lt;/p&gt;</description></item><item><title>Lesson 4: Reducing Allocations — Stack, arena, SmallVec</title><link>/post/rust/rust-perf-allocations/</link><pubDate>Fri, 21 Mar 2025 16:30:00 +0000</pubDate><guid>/post/rust/rust-perf-allocations/</guid><description>&lt;p&gt;I profiled a Rust web service once and found it was allocating 47,000 times per request. Forty-seven thousand. Most were tiny — 16-byte strings, 3-element vectors, temporary buffers. Each individual allocation was fast (jemalloc is good), but 47,000 of them at ~30ns each is 1.4ms of pure allocator overhead. Per request. At 10K RPS that&amp;rsquo;s 14 seconds of CPU time per second, just asking the allocator for memory.&lt;/p&gt;
&lt;p&gt;The fix took half a day and cut allocations to about 200 per request. Here&amp;rsquo;s everything I know about reducing allocations in Rust.&lt;/p&gt;</description></item><item><title>Lesson 3: Profiling — perf, flamegraph, samply</title><link>/post/rust/rust-perf-profiling/</link><pubDate>Wed, 19 Mar 2025 10:45:00 +0000</pubDate><guid>/post/rust/rust-perf-profiling/</guid><description>&lt;p&gt;A colleague once asked me to look at a Rust service that was &amp;ldquo;slow.&amp;rdquo; They&amp;rsquo;d already spent a week trying to optimize the JSON parsing layer because &amp;ldquo;parsing is always the bottleneck.&amp;rdquo; I ran a profiler. Sixty-three percent of CPU time was spent in &lt;code&gt;Drop&lt;/code&gt; implementations, deallocating thousands of small strings that were created and immediately discarded. The JSON parsing was 4% of runtime.&lt;/p&gt;
&lt;p&gt;Profiling would&amp;rsquo;ve found that in five minutes. That&amp;rsquo;s why this lesson exists.&lt;/p&gt;</description></item><item><title>Lesson 2: Benchmarking with criterion and divan — Statistically rigorous benchmarks</title><link>/post/rust/rust-perf-benchmarking/</link><pubDate>Mon, 17 Mar 2025 14:18:00 +0000</pubDate><guid>/post/rust/rust-perf-benchmarking/</guid><description>&lt;p&gt;Last year I reviewed a PR where someone claimed their new serialization code was &amp;ldquo;2x faster.&amp;rdquo; Their benchmark? &lt;code&gt;std::time::Instant::now()&lt;/code&gt; called once before and once after. Single run. No warmup. No statistical analysis. The &amp;ldquo;2x speedup&amp;rdquo; was thermal throttling on the first run.&lt;/p&gt;
&lt;p&gt;Benchmarking is harder than it looks. Let&amp;rsquo;s do it properly.&lt;/p&gt;
&lt;h2 id="why-naive-benchmarks-lie"&gt;Why Naive Benchmarks Lie&lt;/h2&gt;
&lt;p&gt;Before we get into the tools, let me show you all the ways a naive benchmark can mislead you:&lt;/p&gt;</description></item><item><title>Lesson 1: Performance Philosophy — Measure, don't guess</title><link>/post/rust/rust-perf-philosophy/</link><pubDate>Sat, 15 Mar 2025 08:32:00 +0000</pubDate><guid>/post/rust/rust-perf-philosophy/</guid><description>&lt;p&gt;I once spent three days rewriting a hot loop to avoid a single allocation per iteration. Hand-rolled a custom arena, eliminated two clones, even switched from &lt;code&gt;HashMap&lt;/code&gt; to a hand-tuned open-addressing table. Benchmarked the result: 0.3% improvement. The actual bottleneck? A DNS lookup buried in a library call that I never bothered to profile.&lt;/p&gt;
&lt;p&gt;Three days. Zero meaningful impact. That&amp;rsquo;s the lesson I want to start this entire course with.&lt;/p&gt;</description></item><item><title>Lesson 7: CPU vs Memory Tradeoffs — Cache it or compute it, pick one</title><link>/post/go/go-perf-cpu-memory/</link><pubDate>Thu, 20 Feb 2025 00:00:00 +0000</pubDate><guid>/post/go/go-perf-cpu-memory/</guid><description>&lt;p&gt;Every performance optimization ultimately makes the same trade: you&amp;rsquo;re giving up memory to gain CPU time, or giving up CPU time to reduce memory usage. There&amp;rsquo;s no free lunch. The skill isn&amp;rsquo;t knowing that the tradeoff exists — it&amp;rsquo;s knowing which side of it you&amp;rsquo;re on, and making the choice consciously rather than accidentally. I&amp;rsquo;ve optimized for memory when the bottleneck was CPU, and optimized for CPU when memory was the problem. Both directions are wrong when you pick them without data.&lt;/p&gt;</description></item><item><title>Lesson 6: pprof Deep Dive — CPU, memory, goroutine — read all three</title><link>/post/go/go-perf-pprof/</link><pubDate>Sun, 05 Jan 2025 00:00:00 +0000</pubDate><guid>/post/go/go-perf-pprof/</guid><description>&lt;p&gt;When I first learned about &lt;code&gt;pprof&lt;/code&gt;, I thought it was a single tool that told you &amp;ldquo;what&amp;rsquo;s slow.&amp;rdquo; It took me an embarrassingly long time to understand that it&amp;rsquo;s actually a family of profilers — CPU, heap, allocs, goroutine, mutex, block — each answering a completely different question. Reading one profile and ignoring the others is like diagnosing engine trouble by only checking the oil. You might find something. You&amp;rsquo;ll definitely miss something. The profilers are designed to work together, and once you start reading all three, performance diagnosis goes from guesswork to diagnosis.&lt;/p&gt;</description></item><item><title>Lesson 5: Benchmarking Done Right — testing.B is not what you think</title><link>/post/go/go-perf-benchmarking/</link><pubDate>Sun, 10 Nov 2024 00:00:00 +0000</pubDate><guid>/post/go/go-perf-benchmarking/</guid><description>&lt;p&gt;Writing a Go benchmark feels simple. You drop &lt;code&gt;Benchmark&lt;/code&gt; in front of a function name, loop from 0 to &lt;code&gt;b.N&lt;/code&gt;, run &lt;code&gt;go test -bench=.&lt;/code&gt;, and get a number. The number feels authoritative. I spent about a year trusting benchmark numbers that were wrong — not wrong because of bugs, but wrong because of how the benchmark was written. The Go benchmark framework is excellent, but it has sharp edges that will mislead you until you learn to see them.&lt;/p&gt;</description></item><item><title>Lesson 4: String and Byte Conversions — The copy nobody sees</title><link>/post/go/go-perf-string-bytes/</link><pubDate>Wed, 25 Sep 2024 00:00:00 +0000</pubDate><guid>/post/go/go-perf-string-bytes/</guid><description>&lt;p&gt;Go strings are immutable. Byte slices are mutable. Converting between them requires copying the data — every time, without exception, unless you use unsafe tricks you almost certainly shouldn&amp;rsquo;t. That sounds like a minor footnote, but it becomes a significant issue the moment you start handling large volumes of text: parsing HTTP requests, processing log lines, building JSON responses. I&amp;rsquo;ve watched a single &lt;code&gt;string(b)&lt;/code&gt; call inside a tight loop add measurable latency to a production API, and the fix was two lines of code once I knew what to look for.&lt;/p&gt;</description></item><item><title>Lesson 3: Slice and Map Performance — The data structure tax</title><link>/post/go/go-perf-slice-map/</link><pubDate>Tue, 20 Aug 2024 00:00:00 +0000</pubDate><guid>/post/go/go-perf-slice-map/</guid><description>&lt;p&gt;Slices and maps are so convenient in Go that it&amp;rsquo;s easy to forget they&amp;rsquo;re not free. They have hidden costs — in allocations, in CPU cache misses, in GC scanning time — that only become visible when you push them into a hot path and watch your benchmarks light up. I&amp;rsquo;ve been burned by both, sometimes in embarrassing ways, and building intuition for when those costs matter has saved me more than one production incident.&lt;/p&gt;</description></item><item><title>Lesson 2: Stack vs Heap Intuition — The allocation you didn't know you made</title><link>/post/go/go-perf-stack-heap/</link><pubDate>Tue, 16 Jul 2024 00:00:00 +0000</pubDate><guid>/post/go/go-perf-stack-heap/</guid><description>&lt;p&gt;There&amp;rsquo;s a class of performance bugs in Go that doesn&amp;rsquo;t show up in code review, doesn&amp;rsquo;t trigger the race detector, and doesn&amp;rsquo;t cause test failures. It just makes your service slowly worse under load. The source is almost always the same: heap allocations happening in places you didn&amp;rsquo;t intend, turning what should be fast stack operations into GC-visible objects that pile up until the collector has to stop and clean them up. Building an intuition for when Go allocates on the heap versus the stack was one of the single highest-leverage things I did to improve the services I work on.&lt;/p&gt;</description></item><item><title>Lesson 1: Escape Analysis — Where your data lives matters more than how you write it</title><link>/post/go/go-perf-escape-analysis/</link><pubDate>Thu, 20 Jun 2024 00:00:00 +0000</pubDate><guid>/post/go/go-perf-escape-analysis/</guid><description>&lt;p&gt;I spent the first two years of writing Go completely unaware that the compiler was making decisions about my code that I never asked for — and that those decisions were quietly shaping the performance profile of everything I shipped. Escape analysis is the mechanism behind all of it. Once I understood it, I started reading code differently. Not just &amp;ldquo;does this work?&amp;rdquo; but &amp;ldquo;where does this data live, and did I give the compiler a chance to put it somewhere fast?&amp;rdquo;&lt;/p&gt;</description></item></channel></rss>