<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Rust and AI/LLM Integration on Atharva Pandey</title><link>https://atharva.page/series/rust-and-ai/llm-integration/</link><description>Recent content in Rust and AI/LLM Integration on Atharva Pandey</description><generator>Hugo</generator><language>en-us</language><copyright>Copyright ©</copyright><lastBuildDate>Sat, 30 Aug 2025 14:15:00 +0000</lastBuildDate><atom:link href="https://atharva.page/series/rust-and-ai/llm-integration/index.xml" rel="self" type="application/rss+xml"/><item><title>Lesson 8: ML Data Pipelines — polars and processing at speed</title><link>https://atharva.page/post/rust/rust-ai-ml-pipelines/</link><pubDate>Sat, 30 Aug 2025 14:15:00 +0000</pubDate><guid>https://atharva.page/post/rust/rust-ai-ml-pipelines/</guid><description>&lt;p&gt;Last quarter I inherited a Python data pipeline that prepared training data for our recommendation model. It processed 50 million rows. Took 3 hours. Used 64GB of RAM. Everyone accepted this as normal — &amp;ldquo;big data is slow.&amp;rdquo; I rewrote it in Rust with polars. Same data. 4 minutes. 6GB of RAM. My teammates thought I was lying until they ran it themselves.&lt;/p&gt;
&lt;p&gt;Polars isn&amp;rsquo;t just &amp;ldquo;pandas but faster.&amp;rdquo; It&amp;rsquo;s a fundamentally different approach to DataFrame operations — lazy evaluation, query optimization, true parallelism, and a Rust-native API that makes data pipeline code genuinely pleasant to write.&lt;/p&gt;</description></item><item><title>Lesson 7: On-Device Inference — ONNX Runtime and candle</title><link>https://atharva.page/post/rust/rust-ai-onnx-inference/</link><pubDate>Tue, 26 Aug 2025 07:49:00 +0000</pubDate><guid>https://atharva.page/post/rust/rust-ai-onnx-inference/</guid><description>&lt;p&gt;I run a sentiment analysis model on every support ticket that comes in. At first I used the OpenAI API — about 2 cents per ticket. Sounds cheap until you do the math: 10,000 tickets a day, $200/day, $6,000/month. For a model that classifies text into &amp;ldquo;positive,&amp;rdquo; &amp;ldquo;negative,&amp;rdquo; and &amp;ldquo;neutral.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;Switched to a local ONNX model running on a $50/month VM. Same accuracy. Latency dropped from 300ms to 8ms. Cost dropped to roughly zero. Not every task needs GPT-4 — and Rust is arguably the best language for running models locally because you get C++ performance without the C++ pain.&lt;/p&gt;</description></item><item><title>Lesson 6: Building MCP Servers in Rust — Model Context Protocol</title><link>https://atharva.page/post/rust/rust-ai-mcp-servers/</link><pubDate>Fri, 22 Aug 2025 10:23:00 +0000</pubDate><guid>https://atharva.page/post/rust/rust-ai-mcp-servers/</guid><description>&lt;p&gt;The first time I heard about MCP, I dismissed it as yet another protocol nobody would adopt. Then Claude Desktop shipped with MCP support, then Cursor, then Windsurf, then half the AI tools I use daily. Turns out when Anthropic publishes a spec and immediately supports it in their flagship products, adoption happens fast.&lt;/p&gt;
&lt;p&gt;MCP — Model Context Protocol — is a standardized way for AI models to discover and use tools, access data sources, and interact with external systems. Think of it as USB for AI: a universal interface so models don&amp;rsquo;t need custom integrations for every data source. And Rust is a fantastic language for building MCP servers because they need to be fast, reliable, and run for a long time without leaking memory.&lt;/p&gt;</description></item><item><title>Lesson 5: Agent Architectures in Rust — ReAct, planning, and loops</title><link>https://atharva.page/post/rust/rust-ai-agent-architectures/</link><pubDate>Wed, 20 Aug 2025 13:08:00 +0000</pubDate><guid>https://atharva.page/post/rust/rust-ai-agent-architectures/</guid><description>&lt;p&gt;I built my first &amp;ldquo;AI agent&amp;rdquo; by stuffing a system prompt into a while loop and hoping for the best. It worked — sometimes. Other times it&amp;rsquo;d get stuck in infinite loops, burn through $50 of API credits hallucinating tool calls that didn&amp;rsquo;t exist, or confidently produce completely wrong answers after three rounds of &amp;ldquo;reasoning.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;The problem wasn&amp;rsquo;t the LLM. The problem was me treating agent design as an afterthought. Good agents need structure — clear state machines, well-defined stopping conditions, and guardrails that prevent runaway behavior. This is where Rust&amp;rsquo;s type system pays massive dividends, because you can encode these constraints at the type level.&lt;/p&gt;</description></item><item><title>Lesson 4: Embeddings and Vector Search — Semantic search in Rust</title><link>https://atharva.page/post/rust/rust-ai-embeddings/</link><pubDate>Mon, 18 Aug 2025 08:55:00 +0000</pubDate><guid>https://atharva.page/post/rust/rust-ai-embeddings/</guid><description>&lt;p&gt;I spent a week building a keyword search system for internal documentation. Regex patterns, stemming, tf-idf scoring — the whole nine yards. Then someone searched &amp;ldquo;how do I deploy&amp;rdquo; and got zero results because every doc said &amp;ldquo;deployment process&amp;rdquo; instead of &amp;ldquo;deploy.&amp;rdquo; That&amp;rsquo;s when I switched to embeddings.&lt;/p&gt;
&lt;p&gt;Embeddings map text into high-dimensional vectors where semantically similar content lives close together. &amp;ldquo;Deploy&amp;rdquo; and &amp;ldquo;deployment process&amp;rdquo; end up near each other in vector space even though they share almost no characters. It&amp;rsquo;s a fundamentally different approach to search, and once you&amp;rsquo;ve used it, keyword search feels like the dark ages.&lt;/p&gt;</description></item><item><title>Lesson 3: Tool Calling / Function Calling Patterns — Agents need tools</title><link>https://atharva.page/post/rust/rust-ai-tool-calling/</link><pubDate>Sat, 16 Aug 2025 16:42:00 +0000</pubDate><guid>https://atharva.page/post/rust/rust-ai-tool-calling/</guid><description>&lt;p&gt;Here&amp;rsquo;s something that took me embarrassingly long to internalize: LLMs don&amp;rsquo;t &lt;em&gt;do&lt;/em&gt; things. They generate text that &lt;em&gt;describes&lt;/em&gt; doing things. The tool calling protocol is just the model saying &amp;ldquo;hey, I&amp;rsquo;d like you to call this function with these arguments&amp;rdquo; — and then your code actually does it.&lt;/p&gt;
&lt;p&gt;This distinction matters because the entire tool calling system is essentially a serialization contract. The model generates JSON conforming to a schema you provided, you execute the function, and you send the result back. Get the schema wrong, and the model hallucinates arguments. Get the execution wrong, and you&amp;rsquo;ve got a broken agent. Get the result format wrong, and the model can&amp;rsquo;t make sense of what happened.&lt;/p&gt;</description></item><item><title>Lesson 2: Streaming LLM Responses — SSE and WebSockets</title><link>https://atharva.page/post/rust/rust-ai-streaming/</link><pubDate>Thu, 14 Aug 2025 11:37:00 +0000</pubDate><guid>https://atharva.page/post/rust/rust-ai-streaming/</guid><description>&lt;p&gt;The first time I demoed an LLM-powered feature to stakeholders, I made the rookie mistake of using non-streaming responses. The CEO asked a question, hit enter, and stared at a blank screen for eight seconds. &amp;ldquo;Is it broken?&amp;rdquo; No — it was thinking. But by the time the response appeared, she&amp;rsquo;d already mentally moved on to the next agenda item.&lt;/p&gt;
&lt;p&gt;Streaming changes everything. Users see tokens appearing in real-time, which feels responsive even when the total generation time is identical. But implementing streaming in Rust? It&amp;rsquo;s one of those things that&amp;rsquo;s surprisingly nuanced once you get past the happy path.&lt;/p&gt;</description></item><item><title>Lesson 1: Building LLM API Clients in Rust — Type-safe AI calls</title><link>https://atharva.page/post/rust/rust-ai-llm-clients/</link><pubDate>Tue, 12 Aug 2025 09:14:00 +0000</pubDate><guid>https://atharva.page/post/rust/rust-ai-llm-clients/</guid><description>&lt;p&gt;Last month I watched a coworker&amp;rsquo;s Python script silently swallow a malformed response from the OpenAI API. The &lt;code&gt;choices&lt;/code&gt; field came back empty, the code plowed ahead with &lt;code&gt;choices[0]&lt;/code&gt;, and the whole pipeline crashed at 2 AM. Nobody got paged because the error handler was also broken. Classic.&lt;/p&gt;
&lt;p&gt;That&amp;rsquo;s the moment I decided to rebuild our LLM integration layer in Rust. Not because I&amp;rsquo;m some Rust evangelist who thinks Python is evil — I use Python daily. But when you&amp;rsquo;re making API calls that cost real money and feed into production systems, maybe you want a type system that actually catches things before runtime.&lt;/p&gt;</description></item></channel></rss>