<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Observability on</title><link>/tags/observability/</link><description>Recent content in Observability on</description><generator>Hugo</generator><language>en</language><lastBuildDate>Wed, 05 Mar 2025 00:00:00 +0000</lastBuildDate><atom:link href="/tags/observability/index.xml" rel="self" type="application/rss+xml"/><item><title>Lesson 7: The Complete Observability Stack — Logs, metrics, traces, profiles — wired together</title><link>/post/go/go-obs-complete-stack/</link><pubDate>Wed, 05 Mar 2025 00:00:00 +0000</pubDate><guid>/post/go/go-obs-complete-stack/</guid><description>&lt;p&gt;We have covered each signal in isolation — structured logs, Prometheus metrics, OpenTelemetry traces, correlation IDs, pprof profiles, and latency distributions. The preceding six lessons described the individual instruments. This one is about wiring them together into a system that actually works during an incident, not just in a demo.&lt;/p&gt;
&lt;p&gt;The goal is this: when something breaks in production, you should be able to answer four questions within five minutes: Is it broken? Who is affected? Where in the system did it break? What is the root cause? Logs, metrics, traces, and profiles each answer one of those questions. The wiring between them — shared trace IDs, consistent service names, deployment markers on dashboards — is what lets you move between signals without losing context.&lt;/p&gt;</description></item><item><title>Lesson 6: Debugging Latency and Leaks — Your p99 is lying to you</title><link>/post/go/go-obs-latency-leaks/</link><pubDate>Wed, 15 Jan 2025 00:00:00 +0000</pubDate><guid>/post/go/go-obs-latency-leaks/</guid><description>&lt;p&gt;We had a service whose p99 latency was 800ms. The p50 was 12ms. The SLA was 200ms at p99. All our dashboards showed the p99 breaching during peak traffic, but the service felt fine to most users. We added caching. We optimized the slow database query. We bumped the CPU allocation. The p99 barely moved.&lt;/p&gt;
&lt;p&gt;Three weeks of tuning later, a colleague asked me to show him the raw histogram buckets. We looked at the actual distribution: 98% of requests were under 15ms. The remaining 2% were all clustered around 900ms, nearly a bimodal distribution. This was not a &amp;ldquo;slow&amp;rdquo; service — it was a service with two distinct response time populations, and the p99 was sampling from the slow population. The fix had nothing to do with the code on the hot path.&lt;/p&gt;</description></item><item><title>Lesson 5: Profiling in Production — pprof is not just for development</title><link>/post/go/go-obs-profiling/</link><pubDate>Sun, 01 Dec 2024 00:00:00 +0000</pubDate><guid>/post/go/go-obs-profiling/</guid><description>&lt;p&gt;I used to think profiling was something you did when you had a performance problem: reproduce it locally, run pprof, stare at the flame graph, fix the hot path. A fire-fighting tool, not an always-on system.&lt;/p&gt;
&lt;p&gt;Then we had a memory leak in production that we couldn&amp;rsquo;t reproduce locally. The heap grew by about 50 MB per hour under real traffic patterns, causing an OOM every eight hours and a rolling restart across the fleet. Our staging environment used synthetic load that didn&amp;rsquo;t trigger the leak. We had no profile data from when the leak was building — only from after the crash, when the heap had already been cleared.&lt;/p&gt;</description></item><item><title>Lesson 4: Correlation IDs — Connect the logs to the trace to the user</title><link>/post/go/go-obs-correlation-ids/</link><pubDate>Fri, 18 Oct 2024 00:00:00 +0000</pubDate><guid>/post/go/go-obs-correlation-ids/</guid><description>&lt;p&gt;A support ticket lands: &amp;ldquo;User 8842 says their order failed at 2:47 PM yesterday.&amp;rdquo; You open your log aggregator. You search for &lt;code&gt;user_id = 8842&lt;/code&gt;. You get 4,000 log lines — the user made 80 requests that afternoon. You filter to the 2:43–2:51 PM window. You get 300 lines. They interleave with log lines from 12 concurrent requests from other users because your log output is not partitioned by request. The error message, when you find it, says &lt;code&gt;internal server error&lt;/code&gt;. No stack trace, no underlying cause, no request that produced it.&lt;/p&gt;</description></item><item><title>Lesson 3: Distributed Tracing with OpenTelemetry — Follow the request across services</title><link>/post/go/go-obs-tracing/</link><pubDate>Sun, 15 Sep 2024 00:00:00 +0000</pubDate><guid>/post/go/go-obs-tracing/</guid><description>&lt;p&gt;We had an incident where checkout was timing out intermittently. The logs showed the API gateway receiving the request and returning a 504 after 30 seconds. The payment service logged nothing. The inventory service logged nothing. Something was hanging somewhere in the middle, and we had no way to see where.&lt;/p&gt;
&lt;p&gt;I spent four hours bisecting the call graph by adding temporary log lines, redeploying, and re-triggering the error. We eventually found a database query in the inventory service that was waiting on a lock — a lock held by a background job nobody had thought to instrument. Logs told me what each service did in isolation. They told me nothing about the shape of a single request as it flowed across all of them.&lt;/p&gt;</description></item><item><title>Lesson 2: Metrics That Matter — Count, measure, alert</title><link>/post/go/go-obs-metrics/</link><pubDate>Sun, 28 Jul 2024 00:00:00 +0000</pubDate><guid>/post/go/go-obs-metrics/</guid><description>&lt;p&gt;The first metrics dashboard I built for a Go service had forty-two graphs. CPU, memory, goroutine count, heap allocations, GC pause duration, request rate, error rate, and about thirty-five other things that felt important when I added them. Six months later I was on call at 3 AM and the service was degraded. I opened that dashboard, looked at forty-two graphs, and had no idea where to start.&lt;/p&gt;
&lt;p&gt;Metrics are only useful when you know what to alert on, and you can only alert on things you understand. More metrics is not the same as better observability. The question is not &amp;ldquo;what can I measure?&amp;rdquo; but &amp;ldquo;what breaks, how do I know it&amp;rsquo;s broken, and how quickly can I narrow down why?&amp;rdquo;&lt;/p&gt;</description></item><item><title>Lesson 1: Structured Logging with slog — Printf is not observability</title><link>/post/go/go-obs-slog/</link><pubDate>Tue, 25 Jun 2024 00:00:00 +0000</pubDate><guid>/post/go/go-obs-slog/</guid><description>&lt;p&gt;I spent two years writing Go services that logged like this:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-go" data-lang="go"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#a6e22e"&gt;log&lt;/span&gt;.&lt;span style="color:#a6e22e"&gt;Printf&lt;/span&gt;(&lt;span style="color:#e6db74"&gt;&amp;#34;processing order %s for user %d, amount %.2f&amp;#34;&lt;/span&gt;, &lt;span style="color:#a6e22e"&gt;orderID&lt;/span&gt;, &lt;span style="color:#a6e22e"&gt;userID&lt;/span&gt;, &lt;span style="color:#a6e22e"&gt;amount&lt;/span&gt;)
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;It worked fine — until the day I needed to find every failed payment over $500 from a specific user across three weeks of logs in a production incident at 2 AM. I was grepping through gigabytes of text, trying to parse freeform strings with jq, getting nowhere. The logs existed. The information was technically there. But I couldn&amp;rsquo;t &lt;em&gt;query&lt;/em&gt; it in any meaningful way.&lt;/p&gt;</description></item></channel></rss>