<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Llm on</title><link>/tags/llm/</link><description>Recent content in Llm on</description><generator>Hugo</generator><language>en</language><lastBuildDate>Sat, 30 Aug 2025 14:15:00 +0000</lastBuildDate><atom:link href="/tags/llm/index.xml" rel="self" type="application/rss+xml"/><item><title>Lesson 8: ML Data Pipelines — polars and processing at speed</title><link>/post/rust/rust-ai-ml-pipelines/</link><pubDate>Sat, 30 Aug 2025 14:15:00 +0000</pubDate><guid>/post/rust/rust-ai-ml-pipelines/</guid><description>&lt;p&gt;Last quarter I inherited a Python data pipeline that prepared training data for our recommendation model. It processed 50 million rows. Took 3 hours. Used 64GB of RAM. Everyone accepted this as normal — &amp;ldquo;big data is slow.&amp;rdquo; I rewrote it in Rust with polars. Same data. 4 minutes. 6GB of RAM. My teammates thought I was lying until they ran it themselves.&lt;/p&gt;
&lt;p&gt;Polars isn&amp;rsquo;t just &amp;ldquo;pandas but faster.&amp;rdquo; It&amp;rsquo;s a fundamentally different approach to DataFrame operations — lazy evaluation, query optimization, true parallelism, and a Rust-native API that makes data pipeline code genuinely pleasant to write.&lt;/p&gt;</description></item><item><title>Lesson 7: On-Device Inference — ONNX Runtime and candle</title><link>/post/rust/rust-ai-onnx-inference/</link><pubDate>Tue, 26 Aug 2025 07:49:00 +0000</pubDate><guid>/post/rust/rust-ai-onnx-inference/</guid><description>&lt;p&gt;I run a sentiment analysis model on every support ticket that comes in. At first I used the OpenAI API — about 2 cents per ticket. Sounds cheap until you do the math: 10,000 tickets a day, $200/day, $6,000/month. For a model that classifies text into &amp;ldquo;positive,&amp;rdquo; &amp;ldquo;negative,&amp;rdquo; and &amp;ldquo;neutral.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;Switched to a local ONNX model running on a $50/month VM. Same accuracy. Latency dropped from 300ms to 8ms. Cost dropped to roughly zero. Not every task needs GPT-4 — and Rust is arguably the best language for running models locally because you get C++ performance without the C++ pain.&lt;/p&gt;</description></item><item><title>Lesson 6: Building MCP Servers in Rust — Model Context Protocol</title><link>/post/rust/rust-ai-mcp-servers/</link><pubDate>Fri, 22 Aug 2025 10:23:00 +0000</pubDate><guid>/post/rust/rust-ai-mcp-servers/</guid><description>&lt;p&gt;The first time I heard about MCP, I dismissed it as yet another protocol nobody would adopt. Then Claude Desktop shipped with MCP support, then Cursor, then Windsurf, then half the AI tools I use daily. Turns out when Anthropic publishes a spec and immediately supports it in their flagship products, adoption happens fast.&lt;/p&gt;
&lt;p&gt;MCP — Model Context Protocol — is a standardized way for AI models to discover and use tools, access data sources, and interact with external systems. Think of it as USB for AI: a universal interface so models don&amp;rsquo;t need custom integrations for every data source. And Rust is a fantastic language for building MCP servers because they need to be fast, reliable, and run for a long time without leaking memory.&lt;/p&gt;</description></item><item><title>Lesson 5: Agent Architectures in Rust — ReAct, planning, and loops</title><link>/post/rust/rust-ai-agent-architectures/</link><pubDate>Wed, 20 Aug 2025 13:08:00 +0000</pubDate><guid>/post/rust/rust-ai-agent-architectures/</guid><description>&lt;p&gt;I built my first &amp;ldquo;AI agent&amp;rdquo; by stuffing a system prompt into a while loop and hoping for the best. It worked — sometimes. Other times it&amp;rsquo;d get stuck in infinite loops, burn through $50 of API credits hallucinating tool calls that didn&amp;rsquo;t exist, or confidently produce completely wrong answers after three rounds of &amp;ldquo;reasoning.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;The problem wasn&amp;rsquo;t the LLM. The problem was me treating agent design as an afterthought. Good agents need structure — clear state machines, well-defined stopping conditions, and guardrails that prevent runaway behavior. This is where Rust&amp;rsquo;s type system pays massive dividends, because you can encode these constraints at the type level.&lt;/p&gt;</description></item><item><title>Lesson 4: Embeddings and Vector Search — Semantic search in Rust</title><link>/post/rust/rust-ai-embeddings/</link><pubDate>Mon, 18 Aug 2025 08:55:00 +0000</pubDate><guid>/post/rust/rust-ai-embeddings/</guid><description>&lt;p&gt;I spent a week building a keyword search system for internal documentation. Regex patterns, stemming, tf-idf scoring — the whole nine yards. Then someone searched &amp;ldquo;how do I deploy&amp;rdquo; and got zero results because every doc said &amp;ldquo;deployment process&amp;rdquo; instead of &amp;ldquo;deploy.&amp;rdquo; That&amp;rsquo;s when I switched to embeddings.&lt;/p&gt;
&lt;p&gt;Embeddings map text into high-dimensional vectors where semantically similar content lives close together. &amp;ldquo;Deploy&amp;rdquo; and &amp;ldquo;deployment process&amp;rdquo; end up near each other in vector space even though they share almost no characters. It&amp;rsquo;s a fundamentally different approach to search, and once you&amp;rsquo;ve used it, keyword search feels like the dark ages.&lt;/p&gt;</description></item><item><title>Lesson 3: Tool Calling / Function Calling Patterns — Agents need tools</title><link>/post/rust/rust-ai-tool-calling/</link><pubDate>Sat, 16 Aug 2025 16:42:00 +0000</pubDate><guid>/post/rust/rust-ai-tool-calling/</guid><description>&lt;p&gt;Here&amp;rsquo;s something that took me embarrassingly long to internalize: LLMs don&amp;rsquo;t &lt;em&gt;do&lt;/em&gt; things. They generate text that &lt;em&gt;describes&lt;/em&gt; doing things. The tool calling protocol is just the model saying &amp;ldquo;hey, I&amp;rsquo;d like you to call this function with these arguments&amp;rdquo; — and then your code actually does it.&lt;/p&gt;
&lt;p&gt;This distinction matters because the entire tool calling system is essentially a serialization contract. The model generates JSON conforming to a schema you provided, you execute the function, and you send the result back. Get the schema wrong, and the model hallucinates arguments. Get the execution wrong, and you&amp;rsquo;ve got a broken agent. Get the result format wrong, and the model can&amp;rsquo;t make sense of what happened.&lt;/p&gt;</description></item><item><title>Lesson 2: Streaming LLM Responses — SSE and WebSockets</title><link>/post/rust/rust-ai-streaming/</link><pubDate>Thu, 14 Aug 2025 11:37:00 +0000</pubDate><guid>/post/rust/rust-ai-streaming/</guid><description>&lt;p&gt;The first time I demoed an LLM-powered feature to stakeholders, I made the rookie mistake of using non-streaming responses. The CEO asked a question, hit enter, and stared at a blank screen for eight seconds. &amp;ldquo;Is it broken?&amp;rdquo; No — it was thinking. But by the time the response appeared, she&amp;rsquo;d already mentally moved on to the next agenda item.&lt;/p&gt;
&lt;p&gt;Streaming changes everything. Users see tokens appearing in real-time, which feels responsive even when the total generation time is identical. But implementing streaming in Rust? It&amp;rsquo;s one of those things that&amp;rsquo;s surprisingly nuanced once you get past the happy path.&lt;/p&gt;</description></item><item><title>Lesson 1: Building LLM API Clients in Rust — Type-safe AI calls</title><link>/post/rust/rust-ai-llm-clients/</link><pubDate>Tue, 12 Aug 2025 09:14:00 +0000</pubDate><guid>/post/rust/rust-ai-llm-clients/</guid><description>&lt;p&gt;Last month I watched a coworker&amp;rsquo;s Python script silently swallow a malformed response from the OpenAI API. The &lt;code&gt;choices&lt;/code&gt; field came back empty, the code plowed ahead with &lt;code&gt;choices[0]&lt;/code&gt;, and the whole pipeline crashed at 2 AM. Nobody got paged because the error handler was also broken. Classic.&lt;/p&gt;
&lt;p&gt;That&amp;rsquo;s the moment I decided to rebuild our LLM integration layer in Rust. Not because I&amp;rsquo;m some Rust evangelist who thinks Python is evil — I use Python daily. But when you&amp;rsquo;re making API calls that cost real money and feed into production systems, maybe you want a type system that actually catches things before runtime.&lt;/p&gt;</description></item><item><title>Lesson 6: Agent Architectures — Building autonomous agents in Go</title><link>/post/go/go-ai-agent-architectures/</link><pubDate>Thu, 10 Jul 2025 00:00:00 +0000</pubDate><guid>/post/go/go-ai-agent-architectures/</guid><description>&lt;p&gt;An agent is a loop: observe, think, act, repeat. The LLM is the &amp;ldquo;think&amp;rdquo; step — it decides what to do next given the current state. Your Go code handles &amp;ldquo;observe&amp;rdquo; (gathering context), &amp;ldquo;act&amp;rdquo; (executing tool calls), and the loop control that keeps everything running. I&amp;rsquo;ve built agents that write and execute code, agents that browse the web, and agents that orchestrate multi-step data pipelines. The underlying architecture is always the same few patterns, and Go&amp;rsquo;s concurrency makes the execution layer clean and fast.&lt;/p&gt;</description></item><item><title>Lesson 5: Embedding and Vector Search — Semantic search in Go without Python</title><link>/post/go/go-ai-embeddings/</link><pubDate>Sun, 18 May 2025 00:00:00 +0000</pubDate><guid>/post/go/go-ai-embeddings/</guid><description>&lt;p&gt;For a long time, embedding-based semantic search felt like Python territory. The tutorials all pointed to LangChain, FAISS, and numpy. But the actual operations — generate an embedding vector, store it in a database, query for nearest neighbors — map directly onto Go&amp;rsquo;s strengths: clean HTTP client code for the embedding API, &lt;code&gt;pgx&lt;/code&gt; for PostgreSQL with pgvector, and fast concurrent query pipelines. I&amp;rsquo;ve built production semantic search systems entirely in Go and they&amp;rsquo;re fast, maintainable, and don&amp;rsquo;t require a Python sidecar.&lt;/p&gt;</description></item><item><title>Lesson 4: Tool Calling Patterns — Letting the LLM invoke your Go functions</title><link>/post/go/go-ai-tool-calling/</link><pubDate>Wed, 12 Mar 2025 00:00:00 +0000</pubDate><guid>/post/go/go-ai-tool-calling/</guid><description>&lt;p&gt;Tool calling is where LLM integrations get genuinely powerful. Without tool calling, you&amp;rsquo;re limited to asking the model to generate text. With tool calling, you can build a conversational interface where the model decides which of your functions to call, calls them, incorporates the results into its reasoning, and decides whether to call more tools or return a final answer. I&amp;rsquo;ve used this to build support assistants that query databases, coding assistants that run test suites, and research tools that fetch live web content — all driven by the model&amp;rsquo;s judgment about which tools to invoke.&lt;/p&gt;</description></item><item><title>Lesson 3: Streaming Responses — Token-by-token output without buffering the whole response</title><link>/post/go/go-ai-streaming/</link><pubDate>Wed, 18 Dec 2024 00:00:00 +0000</pubDate><guid>/post/go/go-ai-streaming/</guid><description>&lt;p&gt;If you&amp;rsquo;ve ever used a non-streaming LLM endpoint in a user-facing feature, you know the experience it creates: the user submits a question, watches a spinner for 8 seconds, then suddenly gets a wall of text. Streaming changes this completely — the user sees output appearing word by word, which feels responsive and alive even when the total time-to-complete is identical. In Go, streaming LLM responses means reading Server-Sent Events (SSE) from the API and piping them to the client in real time. It&amp;rsquo;s genuinely one of the nicer concurrency patterns I&amp;rsquo;ve implemented.&lt;/p&gt;</description></item><item><title>Lesson 2: LLM API Clients — Calling Claude, GPT, and Groq from Go</title><link>/post/go/go-ai-llm-clients/</link><pubDate>Sun, 06 Oct 2024 00:00:00 +0000</pubDate><guid>/post/go/go-ai-llm-clients/</guid><description>&lt;p&gt;Most Go developers approach LLM APIs the same way they approach any REST API — write an HTTP client, handle errors, parse JSON. That instinct is correct, but LLM APIs have a few characteristics that require specific handling: they&amp;rsquo;re slow (seconds, not milliseconds), they have complex nested response structures, they support streaming, and the model selection and token management have real cost implications. This lesson is about building Go clients that handle all of this properly.&lt;/p&gt;</description></item><item><title>Lesson 1: Building MCP Servers in Go — Give AI agents tools with the Model Context Protocol</title><link>/post/go/go-ai-mcp-servers/</link><pubDate>Tue, 30 Jul 2024 00:00:00 +0000</pubDate><guid>/post/go/go-ai-mcp-servers/</guid><description>&lt;p&gt;When Claude or another AI assistant needs to look up a database record, call an internal API, or read a file from your filesystem, it can&amp;rsquo;t do that on its own — it needs tools. The Model Context Protocol (MCP) is Anthropic&amp;rsquo;s open standard for giving AI agents exactly those tools. An MCP server is a small program you write that exposes tools via a JSON-RPC protocol; the AI client calls your server to invoke them. I find this genuinely exciting as a Go developer: Go&amp;rsquo;s concurrency model and fast startup time make it a natural fit for MCP servers.&lt;/p&gt;</description></item></channel></rss>