<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>ML on</title><link>/tags/ml/</link><description>Recent content in ML on</description><generator>Hugo</generator><language>en</language><lastBuildDate>Fri, 13 Dec 2024 00:00:00 +0000</lastBuildDate><atom:link href="/tags/ml/index.xml" rel="self" type="application/rss+xml"/><item><title>Lesson 4: Recommendation Systems — Collaborative filtering, embeddings, and the cold start problem</title><link>/post/fundamentals/ml-recommendations/</link><pubDate>Fri, 13 Dec 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/ml-recommendations/</guid><description>&lt;p&gt;Recommendation systems are the most economically consequential ML systems most engineers will ever build. Netflix estimates that its recommendation system saves over a billion dollars per year in avoided cancellations. Amazon&amp;rsquo;s &amp;ldquo;customers who bought this also bought&amp;rdquo; drives a significant fraction of its revenue. Spotify&amp;rsquo;s Discover Weekly has become a user retention flywheel. The stakes are real, and the engineering is genuinely interesting.&lt;/p&gt;
&lt;p&gt;I got deep into recommendation systems when I was working on a content platform that had about 500,000 items and needed to surface relevant content for each user without overwhelming the ranking team with engineering requests. What I found was that the architecture of a production recommender is almost always the same shape, regardless of the domain — and the hardest problem is not the ML, it&amp;rsquo;s the cold start.&lt;/p&gt;</description></item><item><title>Lesson 3: Model Serving — Latency, batching, A/B testing in production</title><link>/post/fundamentals/ml-model-serving/</link><pubDate>Mon, 14 Oct 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/ml-model-serving/</guid><description>&lt;p&gt;Training a model is satisfying. Deploying it to serve real traffic is humbling. I remember the first time I pushed a model to production and watched the p99 latency hover at 800ms on what was supposed to be a &amp;ldquo;fast&amp;rdquo; model. The benchmark had shown 12ms inference time. What happened? The benchmark ran the model on pre-loaded batches; production served one request at a time, loaded the model fresh on cold starts, and had no GPU batching. The gap between &amp;ldquo;the model is accurate&amp;rdquo; and &amp;ldquo;the model is fast enough to be useful&amp;rdquo; is where model serving engineering lives.&lt;/p&gt;</description></item><item><title>Lesson 2: Feature Stores — Why feature engineering is 80% of ML</title><link>/post/fundamentals/ml-feature-stores/</link><pubDate>Fri, 02 Aug 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/ml-feature-stores/</guid><description>&lt;p&gt;There is a saying in machine learning that is so universally acknowledged it has become a cliché: 80% of the work in any ML project is feature engineering. I spent a long time thinking this referred to the cognitive labor — the domain expertise required to craft meaningful features. It does, in part. But the deeper meaning is operational. The 80% is not just about &lt;em&gt;what&lt;/em&gt; features to build; it&amp;rsquo;s about &lt;em&gt;how&lt;/em&gt; to compute them consistently, store them efficiently, retrieve them with sub-millisecond latency at serving time, and keep them synchronized between the training pipeline and the production system.&lt;/p&gt;</description></item><item><title>Lesson 1: ML Pipelines — From raw data to deployed model</title><link>/post/fundamentals/ml-pipelines/</link><pubDate>Sun, 02 Jun 2024 00:00:00 +0000</pubDate><guid>/post/fundamentals/ml-pipelines/</guid><description>&lt;p&gt;The first time I built a &amp;ldquo;production ML pipeline,&amp;rdquo; it was a cron job that ran a Python script, trained a model, and saved the artifact to disk. It worked, until it didn&amp;rsquo;t. The model silently degraded over three weeks because the training data schema changed and nobody noticed. There were no tests, no validation, no monitoring. That experience, embarrassing as it was, taught me more about ML system design than any paper on model architecture.&lt;/p&gt;</description></item></channel></rss>