<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Engineering Practices on</title><link>/tags/engineering-practices/</link><description>Recent content in Engineering Practices on</description><generator>Hugo</generator><language>en</language><lastBuildDate>Sun, 11 Aug 2024 00:00:00 +0000</lastBuildDate><atom:link href="/tags/engineering-practices/index.xml" rel="self" type="application/rss+xml"/><item><title>Lesson 7: On-Call Engineering — Reducing toil, improving reliability</title><link>/post/fundamentals/eng-oncall/</link><pubDate>Sun, 11 Aug 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/eng-oncall/</guid><description>&lt;p&gt;I did 12 months of on-call on a team that hadn&amp;rsquo;t invested in reliability. The rotation was weekly. In a bad week, I&amp;rsquo;d get 15-20 pages. A good week was 5. I was exhausted by the end of my shift, and the paging frequency had barely changed over those 12 months. We were fixing incidents, not fixing the causes. The next team I joined approached on-call differently: on-call was treated as a reliability sensor, not a firefighting rotation. Pages were tracked, patterns identified, and root causes fixed. By month 6 I was averaging 2 pages per week on-call. The work we did during on-call made future on-call better.&lt;/p&gt;</description></item><item><title>Lesson 6: Feature Flags — Progressive rollout and kill switches</title><link>/post/fundamentals/eng-feature-flags/</link><pubDate>Sun, 28 Jul 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/eng-feature-flags/</guid><description>&lt;p&gt;We deployed a new checkout flow on a Friday afternoon. It had passed code review, passed testing, passed staging load tests. By Friday evening, the error rate was climbing. The new flow had a race condition that only manifested under specific mobile browser timing that we hadn&amp;rsquo;t tested. Without a feature flag, the fix would have required an emergency deployment — 25 minutes of build time, deployment, and validation. With a feature flag, the rollback was turning off a switch. Thirty seconds.&lt;/p&gt;</description></item><item><title>Lesson 5: Load Testing — k6, vegeta, realistic patterns</title><link>/post/fundamentals/eng-load-testing/</link><pubDate>Sat, 13 Jul 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/eng-load-testing/</guid><description>&lt;p&gt;We load-tested our new checkout service before launch. 1,000 virtual users, 10 minutes, all hitting &lt;code&gt;/v1/checkout&lt;/code&gt; sequentially. It passed with excellent numbers. Launch day: real traffic hit the service, and it fell over at 200 concurrent users. The problem was our test. Real users don&amp;rsquo;t all call the same endpoint in sequence. They browse, add to cart, apply discount codes, fill in addresses, and then checkout — a session that touches 8 different endpoints over 4 minutes. Our test didn&amp;rsquo;t model this. Our test was measuring the wrong thing.&lt;/p&gt;</description></item><item><title>Lesson 4: Monitoring and Alerting — SLOs and alert fatigue</title><link>/post/fundamentals/eng-monitoring/</link><pubDate>Mon, 24 Jun 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/eng-monitoring/</guid><description>&lt;p&gt;The on-call rotation I inherited had 47 active alerts. On a bad week, the on-call engineer got paged 30 times. Most pages were &amp;ldquo;something might be wrong&amp;rdquo; noise — high CPU on one instance, a spike in error rate that self-resolved in 30 seconds, disk space at 70% on a server with months of capacity remaining. Engineers stopped taking the pages seriously. Then the one real incident got buried in the noise, and we had a 4-hour outage because no one treated the first alert seriously. Alert fatigue is not a monitoring problem. It&amp;rsquo;s an architecture-of-trust problem.&lt;/p&gt;</description></item><item><title>Lesson 3: Incident Response — Postmortems and blameless culture</title><link>/post/fundamentals/eng-incidents/</link><pubDate>Sat, 08 Jun 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/eng-incidents/</guid><description>&lt;p&gt;The first major incident I was on-call for, I spent 90 minutes trying to fix the problem and 30 minutes frantically communicating to stakeholders in a panic. The second one, I followed a runbook and spent the 90 minutes coordinating, communicating clearly, and delegating diagnosis — while the problem was resolved in 40 minutes. The difference wasn&amp;rsquo;t technical skill. It was process. Incident response is a skill you can learn and practice, and it makes a measurable difference in how quickly you restore service and how well the team learns from failures.&lt;/p&gt;</description></item><item><title>Lesson 2: Code Review That Works — What to look for</title><link>/post/fundamentals/eng-code-review/</link><pubDate>Wed, 22 May 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/eng-code-review/</guid><description>&lt;p&gt;I&amp;rsquo;ve been on the wrong end of bad code reviews in both directions. Reviews that were nitpick sessions about variable naming while missing a race condition. Reviews that rubber-stamped everything because the reviewer was busy. And I&amp;rsquo;ve given both kinds myself. It took a few years, a few incidents, and a few honest retrospectives to develop a framework for reviews that actually improve code quality without burning out reviewers or demoralizing authors.&lt;/p&gt;</description></item><item><title>Lesson 1: Git Beyond Basics — Rebase, bisect, worktrees</title><link>/post/fundamentals/eng-git/</link><pubDate>Tue, 07 May 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/eng-git/</guid><description>&lt;p&gt;Most engineers use maybe 15% of Git. &lt;code&gt;add&lt;/code&gt;, &lt;code&gt;commit&lt;/code&gt;, &lt;code&gt;push&lt;/code&gt;, &lt;code&gt;pull&lt;/code&gt;, &lt;code&gt;branch&lt;/code&gt;, &lt;code&gt;merge&lt;/code&gt;, and &lt;code&gt;status&lt;/code&gt; covers daily work. That&amp;rsquo;s fine until you need to find which commit introduced a regression across 300 commits, or you need to untangle a messy history before merging, or you want to work on three features simultaneously without context-switching overhead. The commands I&amp;rsquo;m covering here don&amp;rsquo;t come up every day. When they do, they save hours.&lt;/p&gt;</description></item></channel></rss>