Frontier results, on device - RL Nabors, Arize
Description
Most of use reach for a frontier model by default and pay for it on every call, in latency, in energy, in cash, and in everything that leaves their stack. For most of those calls, a small local model would do the job. RL Nabors, former Meta/React core team member and AWS alum, covers the vocabulary you need to reason about model performance (capability evals, golden datasets, LLM-as-judge) and walks through real cases: a local agentic harness replacing a frontier call, an in-browser moderation classifier defended with production-trace evals, and a generative summarization feature where the rubric turns out to be harder than the model. You'll leave with a framework for deciding when to choose large and off-prem or small and local models, and how to measure your way to the answer instead of guessing. You will learn: - The vocabulary to reason about model performance (capability evals, golden datasets, LLM-as-judge). - A framework for deciding when a small or local model can replace a frontier one and when it can't. - A repeatable process for building capability evals from your own production traces, not someone else's benchmark. - Working examples of using eval results to iterate on prompts and ship with confidence instead of vibes. Speakers: - RL Nabors (Arize): RL Nabors builds developer tools and the communities that make them stick. Previously React and MDN, currently developer experience at Arize, perpetually building Mima. X/Twitter: https://x.com/rachelnabors LinkedIn: https://linkedin.com/in/nearestnabors GitHub: https://linkedin.com/in/nearestnabors
Summary
Generated by claude-sonnet-4-5At-a-Glance
- Verdict: Watch fully
- Core thesis: You can replace expensive frontier LLM API calls with on-device small language models (SLMs) by systematically evaluating tasks, selecting the smallest model that meets success criteria, and optimizing with prompt engineering—saving costs, improving latency, and preserving privacy.
- Why it matters: The talk provides a concrete, repeatable framework for right-sizing AI inference that Ken can apply to reduce API spend, improve UX, and enable offline-capable agent systems.
- Best use: Use as a tactical playbook for agent system cost reduction and model selection; extract the 4-step framework (prove, define, test, select) and the prompt engineering tactics for closing performance gaps.
Executive Summary
Rachel Lee Nabors, now at Arize, presents a systematic framework for replacing expensive frontier model API calls (GPT, Claude) with on-device small language models (SLMs). She argues that most tasks don't require the full parameter space of frontier models and demonstrates how to prove feasibility with a large model, collect golden datasets, benchmark SLMs (Gemma, Llama, Qwen) against success criteria, and close performance gaps via prompt engineering. The talk centers on a real case study: converting a social media thread summarization feature from Claude Sonnet to Llama 3.2 3B, achieving comparable accuracy while eliminating $1/day in inference costs and reducing latency from ~3 seconds to ~1 second.
Nabors emphasizes that token costs are falling, but total inference spend is rising due to agentic and reasoning workloads consuming tokens faster than prices drop. She advocates 'prototype big, deploy small'—prove the task is possible with a frontier model, then test progressively smaller models until you find the SAGE (Small And Good Enough) model. She demonstrates using Arize Phoenix for eval workflows, testing Llama 3.2, Gemma 4, and Qwen models on a 28-example golden dataset measuring JSON validity, reference accuracy, factual consistency, length compliance, and latency (P50/P95).
Key tactical findings: few-shot prompting was the most effective optimization (improving reference accuracy and factual consistency by ~5% while adding only 200ms latency), explicit negative constraints backfired, and chain-of-thought added 600ms without meaningful gains. Post-processing (truncation, reference validation) closed the final gap. She also notes Chrome's Prompt API exposes Gemini Nano natively, eliminating model shipping overhead for browser-based features.
The broader implication is that security (no third-party data retention), offline capability, latency (P50 <1s vs 3s for cloud models), and zero marginal inference cost make SLMs viable for production agentic systems. Nabors stresses continuous regression evals to prevent prompt/model updates from degrading output quality, treating them like CI/CD tests. She recommends isolating one variable per prompt variant when optimizing and opening eval results to inspect judge bias (Claude favored its own outputs in factual consistency checks).
Key Takeaways
- Claim: Frontier model inference costs are rising despite falling token prices because agentic and reasoning workloads consume tokens faster than prices drop. | Evidence: Nabors cites that her Mima app (social media client) was consuming ~$1/day in Claude API costs for thread summarization alone; extrapolated to many users, this becomes unsustainable. | Caveat: The $1/day figure is specific to her use case (50-comment threads, multiple summaries); your mileage will vary based on workload volume and complexity. | Implication: For agent systems with high inference volume, marginal API costs can explode—Ken should audit current LLM spend by task and identify high-frequency, low-complexity calls as SLM candidates. | Timestamp: timestamp unavailable
- Claim: SLMs (1B-5B parameters) can match frontier model performance on narrow tasks while consuming 25% of the energy and delivering sub-1-second P50 latency on-device. | Evidence: Nabors references a 2025 NVIDIA research paper finding SLMs are 'sufficiently powerful for agentic task loads' and consume ~25% of LLM energy; her Llama 3.2 3B deployment achieved P50 latency of ~1s vs Claude Sonnet's ~3s. | Caveat: Performance depends on task specificity—SLMs lack the broad knowledge and multimodal capabilities of frontier models; they excel at narrow, well-defined tasks like summarization or classification. | Implication: Ken should map agent workflows to identify tasks that don't require general knowledge (e.g., formatting, extraction, classification) and test SLMs for those steps. | Timestamp: timestamp unavailable
- Claim: The 'prototype big, deploy small' framework: prove feasibility with a frontier model, collect a golden dataset, test SLMs from smallest to largest, select the SAGE model (Small And Good Enough). | Evidence: Nabors tested Claude Sonnet baseline, then Qwen 2.5 (1.5B), Qwen 3 (1.7B), Llama 3.2 (3B), and Gemma 4 (5B) on 28 thread summarization examples; Llama 3.2 achieved 90% accuracy at ~1s latency, beating Gemma 4's 8s latency. | Caveat: Requires upfront investment in golden dataset curation (human-labeled or LLM-labeled ground truth); also assumes you can control model deployment (not always true in enterprise or mobile environments). | Implication: Ken can operationalize this by creating a template golden dataset for common agent tasks (classification, extraction, summarization) and running comparative evals before committing to a model. | Timestamp: timestamp unavailable
- Claim: Few-shot prompting is the most effective optimization for closing the SLM-to-frontier gap; explicit negative constraints ('don't do X') backfire on SLMs. | Evidence: Nabors tested 5 prompt variants: baseline, numbered input, few-shot, strict rules, chain-of-thought. Few-shot improved reference accuracy to 91.7% and factual consistency to 92.9% with only +200ms latency; strict rules ('no preamble, don't count words') degraded performance ('naughty child didn't like instructions'). | Caveat: Few-shot requires curating representative examples, which may not generalize to edge cases; also, different SLMs respond differently to prompt styles (Llama vs Gemma vs Qwen). | Implication: Ken should invest in few-shot prompt libraries for critical tasks and avoid negative instructions; test prompt variants in isolation (one variable at a time) to measure impact. | Timestamp: timestamp unavailable
- Claim: Post-processing (truncation, validation) can close the final performance gap without retraining; LLM-as-judge evals can be biased toward the judge's own model family. | Evidence: Nabors added post-processing to enforce length limits and validate references, bringing JSON validity to 100% and structural validity to 100%; Claude Opus as judge favored Claude Sonnet outputs ('she was cross, not angsty'). | Caveat: Post-processing only works for structured outputs (JSON, references); unstructured tasks (creative writing, reasoning) may require model tuning or distillation. | Implication: Ken should layer deterministic post-processing into agent pipelines (schema validation, range checks) and use multiple judges or human spot-checks to catch eval bias. | Timestamp: timestamp unavailable
- Claim: On-device models (e.g., Chrome Prompt API with Gemini Nano) eliminate model shipping overhead and enable zero-marginal-cost inference at the cost of user device energy. | Evidence: Nabors notes Chrome's Prompt API ships Gemini Nano natively; her Llama 3.2 deployment shifted inference cost to users' devices, saving $1/day in API costs with no ongoing fees. | Caveat: User device variability (CPU/GPU, battery life) affects performance; may not be suitable for low-end devices or battery-constrained use cases. | Implication: Ken should explore browser-native SLMs for web-based agent features (e.g., in-browser summarization, classification) and monitor client-side performance/battery impact. | Timestamp: timestamp unavailable
- Claim: Continuous regression evals (run like CI/CD tests) prevent prompt or model updates from degrading output quality; this is critical for multi-agent systems where one change can cascade. | Evidence: Nabors recounts a founder friend whose CTO 'blew away the agentic experience by accident one morning' via a prompt change; she emphasizes running evals on every prompt/model update. | Caveat: Requires tooling (Phoenix, custom harness) and discipline to integrate evals into deployment pipelines; can slow iteration if evals are expensive or slow. | Implication: Ken should treat agent prompt changes as code changes—version control prompts, run automated evals pre-deployment, and track performance over time. | Timestamp: timestamp unavailable
Detailed Brief
The Case Against Frontier Models: Security, Latency, Cost, Offline
- Claims: Remote LLM calls expose data to third-party servers, risking interception and retention; Latency >4 seconds breaks believability in VR/AI chat experiences (research-backed threshold); Agentic workloads compound inference costs—even if tokens are cheaper, total spend rises; Offline/secure environments cannot use cloud models, limiting productivity
- Evidence: Nabors cites cases of AI chatbot data breaches leaking sensitive business data; Research on VR LLM chats sets 4s as believability limit; many frontier calls exceed this; Mima app example: $1/day for one user's thread summaries → unsustainable at scale; Pixel 10 Pro ships with on-device SLM; Chrome exposes Gemini Nano via Prompt API
- Caveats: Security depends on deployment—self-hosted frontier models can be private; Latency varies by model size, quantization, and hardware; P95 can spike; Cost savings assume high inference volume; low-volume use cases may not benefit
- Implications: Ken should audit current LLM calls for PII exposure and offline failure modes; Agent systems with sub-second latency requirements should prioritize on-device SLMs; High-frequency, low-complexity tasks (classification, extraction) are prime SLM candidates
Task-Specific Models vs SLMs: When to Use What
- Claims: Vision tasks (camera input) → use MobileNet, YOLO, MediaPipe; Audio tasks (microphone) → use Whisper, Wave2Vec; Chat/translation/analysis → use SLMs like Gemma, Qwen, Llama (millions-to-billions of parameters); SLMs contain 1B-5B parameters vs LLMs (billions-to-trillions); 1B params ≈ 2GB in FP16
- Evidence: Nabors provides a 'cheat sheet' mapping input type to model family; Size comparison chart: largest SLM (small dot) << smallest LLM (large dot); Grid shows 1B params fits in 2GB FP16; 5B in 3.1GB (Gemma 4 E2B)
- Caveats: Task-specific models excel at narrow tasks but lack generalization; SLM definition is 'up for debate'—no hard parameter cutoff; Quantization (8-bit, 4-bit) reduces size but may degrade accuracy
- Implications: Ken should match model family to task type (don't use SLMs for vision if MobileNet suffices); For language tasks requiring reasoning but not broad knowledge, SLMs are optimal; Device constraints (RAM, storage) dictate max model size—test on target hardware
The 4-Step Right-Sizing Framework: Prove, Define, Test, Select
- Claims: Step 1: Prove feasibility with the largest capable model (frontier or task-specific); Step 2: Collect golden dataset (curated input-output pairs) and define success criteria (JSON validity, accuracy, latency); Step 3: Test from smallest to largest SLMs, comparing outputs to baseline; Step 4: Select SAGE model (Small And Good Enough) that meets success criteria
- Evidence: Nabors tested Claude Sonnet (baseline), then Qwen 2.5 (1.5B), Qwen 3 (1.7B), Llama 3.2 (3B), Gemma 4 (5B); Golden dataset: 14 threads, 28 examples (summaries + annotations), ~6.7k words total; Success criteria: JSON validity, reference structural validity, factual consistency, length compliance, P50/P95 latency; Llama 3.2 won: 90% accuracy, ~1s P50 latency, vs Gemma 4's 8s latency
- Caveats: Golden dataset quality is critical—garbage in, garbage out; Success criteria must be task-specific (what's 'good enough' for summarization ≠ creative writing); Model selection may change as new SLMs release (Llama 4, Gemma 5, etc.)
- Implications: Ken should template this workflow: baseline → golden set → comparative evals → SAGE selection; Invest in human-labeled golden sets for high-stakes tasks; LLM-labeled for prototyping; Re-run evals quarterly as new SLMs release to avoid overpaying for outdated models
Prompt Engineering Tactics: Few-Shot Wins, Negative Constraints Fail
- Claims: Few-shot prompting (add 2-3 examples) improved accuracy +5% with only +200ms latency; Numbered input (replacing JSON with '1. message') didn't move the needle; Strict rules ('don't do X') backfired—SLMs responded negatively like 'naughty children'; Chain-of-thought added 600ms latency for minimal accuracy gain
- Evidence: Baseline: 91.2% ref accuracy, 87.1% factual consistency, 1s latency; Few-shot: 91.7% ref accuracy, 92.9% factual consistency, 1.2s latency; Strict rules: worse performance across all metrics; Chain-of-thought: slight length improvement, +600ms latency
- Caveats: Few-shot requires curating representative examples (added prompt engineering time); SLM behavior varies by model family—Llama ≠ Gemma ≠ Qwen; Latency impact depends on prompt length and model size
- Implications: Ken should build few-shot prompt libraries for critical agent tasks; Avoid negative instructions in SLM prompts—prefer positive constraints; Test chain-of-thought only if accuracy is paramount and latency is flexible
Post-Processing and Judge Bias: Closing the Final Gap
- Claims: Post-processing (truncation, reference validation) closed the gap to 100% JSON/structural validity; LLM-as-judge can be biased—Claude Opus favored Claude Sonnet outputs by ~5-10%; Final Llama 3.2 + few-shot + post-processing matched/beat Claude Sonnet at 1/365th the cost
- Evidence: Added truncation for length compliance, reference count validation for structural validity; Claude Opus judged Claude Sonnet responses more favorably ('cross' vs 'angsty' interpretation); Final metrics: 100% JSON validity, 100% structural validity, 92.9% factual consistency, P50 ~1s, P95 <3.5s
- Caveats: Post-processing only works for structured outputs—doesn't help with unstructured tasks; Judge bias requires manual spot-checking or using multiple judges; Cost savings assume high inference volume—low-volume apps may not justify optimization effort
- Implications: Ken should layer deterministic validation into agent pipelines (schema checks, range validation); Use multiple LLM judges or human spot-checks to catch eval bias; Document post-processing steps in regression tests to prevent future breakage
Continuous Evals as CI/CD: Preventing Accidental Degradation
- Claims: Treat evals like CI/CD tests—run on every prompt/model update to prevent regression; One prompt change can 'blow away' agentic workflows (real founder story); Phoenix (Arize open-source tool) enables local eval workflows with experiment comparison
- Evidence: Nabors recounts a founder whose CTO accidentally degraded agent performance via a prompt change; Demonstrated Phoenix UI showing raw responses, expected outputs, and side-by-side comparison; Ran 3 trials per prompt variant, averaged results to account for variance
- Caveats: Requires tooling investment (Phoenix, custom harness, golden datasets); Slow/expensive evals can bottleneck iteration—balance coverage vs speed; Golden datasets must evolve with product features to stay relevant
- Implications: Ken should integrate agent evals into deployment pipelines (GitHub Actions, CI/CD); Version control prompts alongside code; treat prompt changes as breaking changes; Track eval performance over time to detect drift and model degradation
Notable Concepts & Terms
- SAGE model (Small And Good Enough): Nabors' framework for selecting the smallest SLM that meets success criteria—optimize for cost/latency while preserving acceptable accuracy.
- Golden dataset: Curated, high-quality input-output pairs (human- or LLM-labeled) used as ground truth for eval benchmarking; critical for right-sizing models.
- Prototype big, deploy small: Core workflow: prove feasibility with frontier model, then convert to SLM for production—balances innovation with cost/latency optimization.
- P50/P95 latency: Median (P50) and worst-case (P95) response times; Nabors set P50 <1.5s, P95 <3.5s based on 4s believability threshold from VR research.
- LLM-as-judge bias: Tendency for LLM judges to favor outputs from their own model family; Nabors caught Claude Opus favoring Claude Sonnet by ~5-10% in factual consistency.
- Few-shot prompting: Adding 2-3 example input-output pairs to the prompt; most effective optimization for SLMs in Nabors' tests (+5% accuracy, +200ms latency).
- Quantization (8-bit, 4-bit): Technique to reduce model size/memory by lowering parameter precision; 1B params ≈ 2GB in FP16, smaller in 8-bit/4-bit.
- Phoenix (Arize): Open-source observability/eval tool for LLMs and agents; Nabors used it to run comparative evals and inspect raw outputs.
- Chrome Prompt API + Gemini Nano: Chrome ships Gemini Nano natively, accessible via Prompt API—eliminates model shipping overhead for browser-based SLM features.
Operator Notes / Why Ken Should Care
- The 4-step framework (prove, define, test, select) is immediately actionable—Ken can apply it to any agent workflow with high inference costs.
- Few-shot prompting delivers the best ROI for SLM optimization; avoid negative constraints and chain-of-thought unless latency is flexible.
- Golden datasets are the linchpin—invest in human-labeled sets for high-stakes tasks, LLM-labeled for prototyping; treat them as version-controlled assets.
- LLM-as-judge bias is real—use multiple judges or human spot-checks; consider fine-tuning a small judge model on your domain.
- Chrome's Prompt API (Gemini Nano) and Pixel's on-device SLMs signal a shift toward native browser/OS inference—Ken should explore for client-side agent features.
- Continuous regression evals prevent 'CTO ruins the workflow' moments—integrate into CI/CD, version control prompts, and track performance over time.
- Post-processing (truncation, validation) is underutilized—layer deterministic checks into agent pipelines to close gaps without model retraining.
- Energy efficiency matters for user experience (battery drain) and sustainability—SLMs consume ~25% of LLM energy per task.
- Latency thresholds are task-specific: 4s for VR/chat believability, <1s for snappy UX; measure P50/P95 to avoid worst-case surprises.
- Model selection is not one-time—re-run evals quarterly as new SLMs release (Llama 4, Gemma 5) to avoid overpaying for outdated models.
Watch Map
- timestamp unavailable: Timestamps not provided in transcript; speaker uses slides/demos but no explicit chapters.
- ~0:00-3:00: Intro: security, latency, cost, and offline problems with frontier models; energy consumption comparison (SLM 25% of LLM).
- ~3:00-8:00: Task-specific models vs SLMs; size comparison chart; definition of 'small' (1B-5B params); on-device deployment (Pixel 10, Chrome Prompt API).
- ~8:00-15:00: 4-step framework walkthrough: prove feasibility (Claude baseline), define success criteria (JSON validity, latency, accuracy), test SLMs (Qwen, Llama, Gemma).
- ~15:00-22:00: Eval results: Llama 3.2 wins (90% accuracy, 1s latency); Gemma 4 loses (8s latency); scatter plot comparison.
- ~22:00-28:00: Prompt engineering: few-shot wins (+5% accuracy, +200ms), strict rules fail, chain-of-thought adds latency; post-processing closes gap to 100%.
- ~28:00-31:00: LLM-as-judge bias (Claude favors Claude); continuous regression evals as CI/CD; closing advice (prototype big, deploy small).
Source/Metadata
- Title: Frontier results, on device - RL Nabors, Arize
- Transcript words: 6786
- Duration seconds: 1851
- Timestamp note: Timestamps/chapters not provided in transcript; estimated from content flow and duration (1851s ≈ 31 min).
Transcript
Hi there, I'm Rachel Lee Neighbors, and today I'm here to talk with you about how to use local models to stop paying for frontier models. Let's dig into it. So I've worked on standards that power today's web with Mozilla on Firefox DevTools and the W3C on web standards, and of course on Microsoft's Edge browser. I've even been on the React team. Now, I've spent the past three years consulting with AI startups and some of our favorite LLM and browser companies on all things web, AI, and UI. And recently I've joined Arise. Have you ever had your CTO ruin your agentic workflow with a slight change of prompt or an LLM migration? Have you ever been that CTO? Well, you probably need Arise's observability platform for models and the agents who love them. Anyway, I'll be actually using one of Arise's open source projects today, Phoenix. We'll talk more about that later. But today, specifically, I'm here to talk with you about how AI is costing you. Every time you reach for foundation models like GPT-5 or CLOD, it's costing you, your users, and the environment. Let's have a look at the costs of one-size-fits-all inference. So first, there's security. Security costs trust. When you use a large LLM that's in the cloud, you're sending data to remote servers, and it always carries the risk of exposure, interception, and retention by third parties. We have cases where the use of remote AI chatbots has led to sensitive business data being stored, breached, or leaked to the public. Latency costs the user experience. Now, there's been research on mitigating response delays in LLM chats in VR and found that four seconds is the limit of believability for users. And many calls that you will make to large models are going to take longer than four seconds, as we will see when we check Phoenix. It costs your business. Third-party inference costs are uncontrollable compared to API costs. Agency compounding levels of inference, and this means even if the tokens are cheaper, you may be using more of them, or if you're using four of them, they may be more expensive. And of course, if you aren't connected, remote models simply aren't going to work, which means that unless your software is connected to the web, nobody can use it. Big inference going offline costs productivity. If there's an outage, if you're in a place where you cannot reach Wi-Fi, or you're in a very secure environment. Now, token costs have been falling as of late, but total inference spend has been rising because agentic and reasoning workloads consume tokens way faster than prices are dropping. But we can completely eliminate most of these costs, and it starts by asking ourselves exactly how much is this costing? Do we really need an LLM to do this job? All right, you can use task-specific models, which are small in size and power consumption compared to an AI foundation. This is my little cheat sheet. If you're looking for an expert model, is a camera pointing at something? Do vision, is there a vision component? You can use a model like MobileNet, YOLO, MediaPipe. Is it a microphone recording something? There are audio models like Whisper and Wave 2, Vectu. Chat, translation, or analysis. This is where you might use something like a small language model like Gemma or Quinn. These small models are called SLMs, or smaller language models. The definition of small is up to debate here. These are great for times when we do need the language power of a generative pre-trained transformer, a GPT, but we probably don't need the sum total of human knowledge in a black box at our disposal. Or we don't need multimodal capabilities. We're not going to be analyzing images and audio at the same time. Now, SLMs, smaller language models, contain millions to billions of parameters. An LLM contains billions to, well, trillions. So you see the big dot there? That's one of the smaller LLMs that's out there. But the little dot actually represents one of the larger SLMs. So you can see that there is a huge difference in parameter size. And this means that there is a vast difference in the size of machine you're going to need to run one of these models. The good news is you don't need most of what's in the big green dot there. You don't need history. You don't need philosophy. You don't need all those Reddit chats. You don't need a lot of what the models have learned and been trained on. Most of us are using our models for things like summarizing a chat thread or detecting if this person's being a jerk right now. And it takes a surprisingly smaller amount of parameters to determine those things. Now, the nice thing about smaller language models is that they consume the same or less energy as large language models to produce correct responses. We have good research that shows this. Small language models come in all sizes and shapes. Most small language models for mobile and web are deployed with quantization. That is to say 8 bit, 4 bit. And that can have reduced disk and memory requirements. One billion parameters fits on about two gigabytes in FP16. This grid, by the way, gives upper bound estimates. So these are just to give you an idea of what will and will not fit on different devices. These are so lightweight that they can be put on devices. Speaking of, my Pixel Pro ships with one. And I, of course, bought the Pixel 10 Pro as soon as I could because I wanted to know what it would be like to work with an on-device model. An SLM of my very own. The good news is that small language models are production ready. NVIDIA called SLMs the future of agentic AI. Once again, great research paper from 2025 that found that SLMs are sufficiently powerful for running agentic task loads. And they consume less energy than language models. Let me give you a bar chart here to take a peek. So let's say this is the total amount of energy consumed to perform a task that an LLM would take. An SLM takes about 25% of that. And a task specific model takes about half of that overall. So as you can see, from an energy perspective, it's much cheaper to run the smaller language models and task specific models. So what are the benefits? Once more, more secure, works offline, no fees, more efficient, and lower latency because it's on-device, no round trips. Now, I've been using local AI for some time. I, for instance, first started using local models. You see on the left, you've got Claude, but on the right, you have Goose, which is an open agent harness that has a really nice interface for chatting with models. And this one here, it's running Gemma, and it's able to answer. It takes a little longer. Smaller models can take a little longer depending on what device you're running and how big the model is. But one of my favorite questions to ask a model is, how likely is it that ichthosaurs, marine reptiles from the dinosaur era, had echolocation? Because if you look at their skeletons, they look a lot like dolphins. I, for instance, first started using local models. You see on the left, you've got Claude, but on the right, you have Goose, which is an open agent harness that has a really nice interface for chatting with models. And this one here, it's running Gemma, and it's able to answer. It takes a little longer. Smaller models can take a little longer depending on what device you're running and how big the model is. But one of my favorite questions to ask a model is, how likely is it that ichthosaurs, marine reptiles from the dinosaur era, had echolocation? Because if you look at their skeletons, they look a lot like dolphins. And it takes reasoning about biology to be able to determine whether or not that's possible. And it's hilarious to me that even, I think this was Gemma 3, was able to come up with a good response that there is no evidence in the fossil record of their having developed a melon or the specialized bones required for conduction that dolphins have. Of course, Claude hedged it by saying, well, we just can't know about one out of every three times, which was interesting. Claude has always been a little nervous about coming forth with an opinion. So how do you pick the best model for the job? This is the hard part. I recently built a framework with Google that I use on my own projects. You can find more about it at web.dev, but I'll run through it here. Now, first off, I like to think of this as prototype big, deploy small. Just repeat this to yourself. Prototype big. Think big, go big. Deploy small. You want to convert the parts of your system over to SLMs and specialized models for production, but you can prototype on a foundation model. No problem. The first step in this process is to determine whether or not what you're trying to do is even possible at all. You use the largest, most capable model and you just see if you can get it to do the thing. See if you can get the model to recognize people by their handwriting or the way they write. If that is possible, then you know that another model can probably handle it too. So you could use a foundation model like Gemini, or you could use a really tough task specific model for handwriting recognition, for instance. Now, this is a feature of the product that I build on the side. It's Mima. It is a client for all your social networks. And I built a feature for it that summarizes long conversation threads because sometimes I go to bed and I wake up and there's 50 comments on something I posted. And I just want to know, are people angry at me? And this is a feature that saves me a lot of heart attacks in the morning. So I first prototyped it out with Claude to prove that it's good enough. You could see the little summaries there. And so that's a pretty good job of saying who's talking about what. So first step for this was to collect a set of inputs and outputs. That is to say, in my case, I wanted to collect a set of threads and then how I would summarize them. Or in this case, how Claude summarized them seemed adequate to me. So I exported a golden data set. Now a golden data set is a curated high quality collection of preferably human labeled input output pairs that you're going to use as the ground truth to evaluate, validate and benchmark your model. So this is what the data looked like. This is all public knowledge, but you can see it's got the author handle, the content created at, whether or not it came from me. And I created a big JSON-L of all of these. So there were 14 threads and I evaluated each one for summaries and annotations, because some of the summaries actually link deeper into the conversation. So it was the same threads, but had two different outcomes for them. One was a short summary and one was a summary with references. All right. So that means 28 examples all together. So these are the things that I was going to be measuring. You need to, before you start doing anything, you need to know what it is that you're measuring. What is success? In this case, was the JSON being output by the summarization process? The easy way to test is to try JSON parse and if it works, yay, if it doesn't, no. The reference structural validity, if it's pointing to different parts of the conversation inside the summary, do those parts actually exist? Factual consistency. That is to say that it's actually able to summarize the content of the threads within reasonableness. So if we're talking about cats, it doesn't give a summary saying we're talking about puppies. This is something that we need an LLM to judge or humans, but the LLMs will be cheaper. Length compliance, making sure it stays within a certain word count. And of course, checking the latency P 50. This is the median across the eval set and the P 95 is the worst case scenario. All right. So you collect these, and now you're going to test through different models and compare how those models rank against the big models. You test from small to large approaches where we compare the performance of the large model against the performance of a selection of smaller models. We say, this is what Claude Opus produced, and this is what Gemma 4 produced. How do these two stack up against each other? Now, I used Claude Sonnet for the baseline, and you can see actually there in the bottom, number one, Approaches. Approaches. Approaches. Approaches. Approaches. Approaches. Approaches. Approaches. Approaches. Approaches. where we compare the performance of the large model against the performance of a selection of smaller models. We say, this is what Claude Opus produced, and this is what Gemma 4 produced. How do these two stack up against each other? Now, I used Claude Sonnet for the baseline, and you can see actually there in the bottom, number one, you can see the baseline. It's looking pretty good. It's got an average latency of 2.9 seconds. It costs about 0.22 cents to run 14 of these tasks. So I actually did some math, and it turns out I'm using about a dollar worth of inference every day using MIMA. I don't have the money to pay for that many teenage girls using MIMA every day. So we're going to have to run this on device. Good news is the total cost column for all these small local models is absolutely zero [SPEAKER_00] because that inference has been pushed to the consumer. [SPEAKER_00] It runs on their device. [SPEAKER_00] They're the one who has to charge the phone [SPEAKER_00] so that it can draw energy from the battery to run the model. [SPEAKER_00] Of course, you want it to not suck the battery dry, [SPEAKER_00] but that is another conversation for something [SPEAKER_00] that we could be testing and evaluating in the future. So let's look at our contestants here. We've got Quen 2.5 Instruct weighing in at 1.5 billion parameters, only a handy one gigabyte on disk, and his sister, Quen 3, 1.7 billion, just a little bit bigger. Llama 3.2 weighing in at 3 billion parameters, and a tidy two gigabytes on disk. And then there was Gemma 4 E2B, [SPEAKER_00] which is 5 billion parameters, a hefty 3.1 gigabytes. [SPEAKER_00] But this was the one that so many engineers I spoke with [SPEAKER_00] when I was picking a model were saying, oh, Gemma 4 is the best. You got to use Gemma 4. And I think that's important here because if I had just gone with what my buddies told me, I would have given the user an extremely different experience, not a good experience. You're going to want to select the smallest model [SPEAKER_00] that gives acceptable responses for your use case, [SPEAKER_00] or as I call it, the SAGE model, [SPEAKER_00] the small and good enough model. [SPEAKER_00] I'm trying to make this a thing. [SPEAKER_00] Bear with me. [SPEAKER_00] I hope we can make SAGE happen. [SPEAKER_00] Now, at first I thought I wanted the Quen 2.5 model [SPEAKER_00] because it was the fastest. [SPEAKER_00] You can see that Quen is all the way here. Let's see where to go. Yeah, 2.5 is the gray circle in the lower left-hand corner. It came in around one second total in latency on the P50. That's amazing. That's really fast for an AI summary. The problem was that its accuracy was pretty low compared to everybody else. The orange square is Gemma 4 E2B, and the blue diamond is Claude Sonnet, which is our ceiling. It's the most accurate, but it's also a little slow. It's weighing in around three seconds in latency. So when we take accuracy into account, the winner was actually Llama 3.2, which is the big green circle that's weighing in right here around the 90% for accuracy. And you can see that both Claude and Llama 3.2 are much faster than Gemma 4. Gemma 4 was coming in around eight seconds. Now, maybe that was because Gemma 4 needed a different kind of prompt, but this was pretty consistent. I was testing each one of these three times and then averaging the results. This is what the results look like. I recommend when you're running evals with something like Phoenix, you open up the experiment. You actually take a look. You can see what the raw responses were, what the expected responses were. You can actually compare them against each other. I found that in many cases, Llama's response was so close to Claude's response as to be indistinguishable from one another. So Llama 3.2 was the ultimate winner, and I decided to move forward with them. And this makes sense because Llama 3.2 is created by Meta, and Meta has a strong interest in creating models that do a good job with human inputs and summarizing human things on a social network. Makes sense, right? MIMA is a social thing, benefits from a social model. So recap, here's how you right-size your model in four steps. Number one, you prove it's possible. You test whether what you're trying to accomplish is possible at all by using the largest possible model. This could be a foundation model like Gemini or a task-specific model. And you set success criteria. You collect a set of inputs and outputs that you want to see coincide with one another. This will be the bar. And then you test from small to large. You compare the outputs of small models against your test criteria, and you work your way up from the smallest model until you get within an acceptable range of that small and good enough model, that SAGE model. And that's when you select your SAGE model, the smallest model that gives acceptable responses for your inputs. But you're probably wondering, what are we going to do about that gap? I mean, 90% accuracy. That sounds like it could be a deal breaker. Well, we can squeeze better performance out of smaller models with prompt engineering. This is important in cases where you can't control which model you're using. Some people might, for instance, create a distilled model that's been trained to do this one task really well. But if you're working, for instance, with a mobile app, you might not want to be using a distilled model, because every time you add capabilities, you'll probably have to train the model slightly, and then you'll be shipping a new one or two gigabyte model every time to your users. 90% accuracy. That sounds like it could be a deal breaker. Well, we can squeeze better performance out of smaller models with prompt engineering. This is important in cases where you can't control which model you're using. Some people might, for instance, create a distilled model that's been trained to do this one task really well. But if you're working, for instance, with a mobile app, you might not want to be using a distilled model, because every time you add capabilities, you'll probably have to train the model slightly, and then you'll be shipping a new one or two gigabyte model every time to your users. So this is a great example of a scenario where I don't think I'll be able to control that model. Once it's on the user's device, I'm not going to ask them to download new ones that would eat up their data plan. You might join a team and they're already committed to Gemma 3 and they're just not going to be implementing Gemma 4 until they've rolled out an update that addresses all of the evals. So let's take a look at closing the gap between SANA and LAMA. Now the green and blue dots in that upper left-hand quadrant. So the measures I honed in on, the ones that they really seem to have different results with, were JSON and reference structural validity, factual consistency, latency, P50 latency, and P95 latency. And I decided that the P50 latency cutoff is one, one and a half seconds for P95, it would be three and a half. Because remember, four seconds is that worst case scenario for people feeling disconnected from an AI powered experience, as per the research. So where do we start? I recommend that you optimize one step at a time. You want to isolate one variable per prompt variant to test whether what you're trying to accomplish is moving the needle when you're using the different prompts. So I created five prompts. Well, I created four prompts and I had the original prompt as the baseline. The original prompt was pretty good. V2 used number input. The same prompt reformatted the thread, the thread as numbered messages instead of using JSON. The hypothesis being that smaller models could track natural language indexing better than array offsets in a heap of JSON. The second one was a few shot. It added a couple of examples and outcomes to the prompt. The hypothesis here being that small models learn format from examples faster than from rules. Then there was the strict rules version. This one was a house of no prompt. It added explicit negative constraints. No preamble. Don't count words before responding. Do count words before responding. And the hypothesis was that small models respond to literal commands and that they like to be bossed around a bit. And then lastly, there was chain of thought, which forced the model to identify key moments before writing the hypothesis being that thinking out loud would improve grounding. So ran the tests again with these new prompts and just to the LAMA 3.2. And we discovered a couple of things. I ran this locally on my machine and was able to compare the results. You can even put the results into something like Claude and have a conversation about the trades if you like. So all right, we're back. So in this case, the baseline wasn't really good at determining how short it should be. So it would fit within a certain section of content. The ref accuracy was 91.2%. It was factually correct 87.1% of the time. And the latency was one second. Reformatted input didn't really make a difference. Explicit rules actually made things worse. The model responded very negatively to being told what it couldn't do. It was a naughty child that didn't like to take instructions. Chain of thought didn't have that big improvement. It did a little bit better on the length. And unfortunately that came at increasing the latency by 600 milliseconds. The best performing one was the few shot one that provided a couple of threads and a couple of examples. It was much better at getting the length right. It was more accurate. And in its references, it agreed with Claude model a little bit more about what was said, and it only increased the latency by 200 milliseconds. So that sounds like a pretty good deal, right? So it was the few shot prompt that one made the biggest improvement. Let's have a look at how it stacks up to the original bump up. So you can see the bar here, Claude Sonnet versus Llama 3.2 3B. I actually did a couple of things at this point because I wanted to close that gap completely. So Llama 3.2 3B with a few shot prompt was actually able to get within a reasonable error margin. The P50 latency went to less than one and a half seconds. So it was totally green. The structural validity was 91.7%. Factual consistency was also at 92.9%. The P95 latency was well under 750 milliseconds, less than Claude in general. Its latency was really good. But let's see about those couple of 10% here between structural validity and factual consistency. This is why it's important to actually open up your evals and take a look at what's inside. When it came to factual consistency, it turned out that Claude was just being a very strict judge. I was using Claude to judge the responses and comparing Claude Opus was comparing Claude Sonnet's response to Llama 3.2's response. And of course, Claude was favoring its little sister, saying, "I don't think your interpretation of what Jenna said is accurate because you said she was being angsty and she was actually being cross." It was that sort of thing. And this is why it's important to crack them open. Now, as for reference consistency and length, those actually could be handled inside the harness and post-processing. So making sure that it's got the right number of references in the thread, that's something very simple to look at. You can just see how long the thread is. And if there are more references than there are members of the thread, that's incorrect. When it comes to how long the actual summary is, if it's too long, you can just truncate it. And so when we added the post-processing, we're actually able to close that gap pretty solidly. Now we've got 100% JSON validity, structural validity was 100%, factual consistency there's only a little bit of a disagreement there. And it turns out that it was the judge being too strict. We've got the P50 latency down to around one second and P95 well, it was under 300 milliseconds. Thread, that's incorrect. When it comes to how long the actual summary is, if it's too long, you can just truncate it. And so when we added the post-processing, we're actually able to close that gap pretty solidly. Now we've got 100% JSON validity, structural validity was 100%, factual consistency. There's only a little bit of disagreement there. And it turns out that the judge was being too strict. We've got the P50 latency down to around one second and P95 was under three and a half seconds. So that's pretty good. It actually ended up meeting and beating Claude Sonnet after doing this little bit of extra effort. And I'm saving about a dollar a day in inference costs. So the important thing is after you've done something like this, the first thing is that you don't want to lose ground. You want to make sure that you can update the prompt in the future or upgrade the model. You don't want to lose ground. You want to ensure that you can upgrade the model or change the prompt in the future. And it won't cause summaries to expand to be a paragraph long, or to start hallucinating things that are untrue. You want to keep your evals running for that. And that would be a regression eval. And you run these like you run CI/CD tests. It's how you keep your CTO from blowing away your agentic experience by accident. One morning, true story happened to a founder friend of mine. So where do you get started? How many Claude calls could be Llama calls? I challenge you to go home today and take a look at what you're sending to LLMs and ask yourself, is this something that a smaller model could handle and how much money would I save if I did that? When working on AI projects, keep an eye out for SLMs and specialized models that might be on device already. For instance, Chrome and the Prompt API, they access Gemini Nano, which ships with Chrome natively. And that can be really useful because that means you don't need to be shipping a model to anyone using the browser. You can take advantage of what's right there. Keep in mind that these models are more efficient and they meet your users where they are. Their information stays on device. You don't have to worry about PII and the energy consumption is a blessing for the environment, so to speak. Remember to prototype big and deploy small. You may want to prototype your system with a foundation model and then convert parts of it to small language models and specialized models for production. Keep in mind you want to prove it, define it, test it, and then select your SLM model that passes the test. You can use prompt engineering to help get better results from SLMs and close that gap with remote models. So I challenge you to consider your current implementation of prototype and convert one feature to use a smaller local model. Go home, try it out for yourself, and you might be surprised. Anyway, I look forward to seeing you out there on the web. If you enjoyed talking about small local models or the authentic web, you can follow me. I'm at nearestneighbors.com. If you want to try running some evals with your current setup, you can start testing with Phoenix at phoenix.arise.com. And lastly, you can sign up for the beta for MIMA at mima.social. I've been Rachel Neighbors, and it's been awesome chatting with you today. Go forth and build your own inference stack. Well, we can squeeze better performance out of smaller models with prompt engineering. This is important in cases where you can't control which model you're using. Some people might, for instance, create a distilled model that's been trained to do this one task really well. But if you're working, for instance, with a mobile app, you might not want to be using a distilled model, because every time you add capabilities, you'll probably have to train the model slightly, and then you'll be shipping a new one or two gigabyte model every time to your users. So this is a great example of like, ooh, I don't think I'll be able to control that model. Once it's on the user's device, I'm not going to ask them to download new ones that would eat up their data plan. You might join a team and they're already committed to Gemma 3 and they're just not going to be implementing Gemma 4 until they've rolled out an update that, you know, addresses all of the evals. So let's take a look at closing the gap between SANA and LAMA. Now the green and blue dots in that upper left-hand quadrant. So the measures I honed in on, the ones that they really seem to have different results with, were JSON and reference structural validity, factual consistency, latency, P50 latency, and P95 latency. And I decided that the P50 latency, the cutoff is, you know, one, one and a half seconds for P95, it would be three and a half. Because remember, four seconds is that worst case scenario for people feeling disconnected from an AI powered experience, as per the research. So where do we start? I recommend that you optimize one step at a time. You want to isolate one variable per prompt variant to test whether what you're trying to accomplish is moving the needle when you're using the different prompts. So I created five prompts. Well, I created four prompts and I had the original prompt as the baseline. The original prompt was pretty good. V2 used number input. The same prompt reformatted the thread, the thread as numbered messages instead of using JSON. The hypothesis being that smaller models could track natural language, indexing better than array offsets in a heap of JSON. The second one was a few shot. It added a couple of examples and outcomes to the prompt. The hypothesis here being that small models learn format from examples faster than from rules. Then there was the strict rules version. This one was a house of no prompt. It added explicit negative constraints. No preamble. Don't count words before responding. I mean, do count words before responding. And the hypothesis was that small models respond to literal commands and that they like to be bossed around a bit. And then lastly, there was chain of thought, which forced the model to identify key moments before writing the hypothesis being that thinking out loud would improve grounding. So ran the tests again with these new prompts and just, just, just to the LAMA 3.2. And we discovered a couple of things. We, I mean, myself, just ran this locally on my machine and was able to compare the results. You can even put the results into something like Claude and have a conversation about the trades if you like. So stop that. All right, we're back. So in this case, uh, the baseline wasn't really good at, at the, the determining how short it should be. So it would fit within, um, a certain section of content. The ref accuracy was 91.2%. It was factually correct 87.1% of the time. And the latency was one second reformatted input didn't really make a difference. Um, explicit rules actually didn't, made things worse. Um, the model responded very negatively to being told what it couldn't do, could and couldn't do it. Um, it was a naughty child didn't like to take instructions. Chain of thought, uh, didn't have that big, uh, it, it did a little bit better on the length. And unfortunately that came at increasing the latency by, uh, 600 milliseconds. The best performing one was the few shot one that provided a couple of threads and a couple of examples. It was much better at, uh, getting the length, right. It was more accurate. Um, and, uh, in its references, it agreed with Claude model a little bit more about what was said, and it only increased the accuracy, uh, the latency by 200 milliseconds. So that sounds like a pretty good deal, right? So it was the few shot Trump, uh, prompt that one made the biggest improvement. Let's have a look at how it stacks up to the original bump up. Um, so you can see the bar here, Claude Sonnet versus Llama 3.23 B. Um, I actually did a couple of things at this point because I wanted to close that gap completely. So Lava 3.2 B with a few shot prompt was actually able to get within a reasonable era area of, um, a reasonable error margin of error. The P 50 latency went to less than one and a half minutes. So it was totally green. Uh, the structural validity, 91.7% factual consistency was also at 92.9%. The P 95 latency was well under 750 milliseconds, uh, less than Claude in general, its latency was really good, but let's see about those, uh, those couple of those 10% here between structural validity and factual consistency. This is why it's important to actually open up your evals and take a look at what's inside. Uh, when it came to factual consistency, it turned out that Claude was just being a very strict judge. I was using Claude to judge the responses and comparing, you know, Claude Opus was comparing Claude sonnets response to, uh, Lama 3.2's response. And of course, Claude was favoring its little sister being like, yeah. Um, I think, I think that, uh, you know, I don't think your interpretation of what Jenna said is accurate because you said she was being angsty and, uh, she was actually being cross. It was that sort of thing. And this is why it's important to crack them open. Now, as for rough consistency and length, those actually could be handled inside the harness and post-processing. So making sure that it's, you know, got the right number of references in the thread, that's something very simple to look at. You can just see how long the thread is. And if there are more roughs than there are members of the thread, that's incorrect. Uh, when it comes to, you know, how long the actual, uh, summary is, if it's too long, you can just truncate it. And so when we added the post-processing, we're actually able to close that gap pretty solidly. Now we've got, uh, 100% JSON validity, structural validity was 100% factual consistency. There's only a little bit of a disagreement there. And it turns out that it was the judge being too darn strict. We've got the P5, uh, P50 latency down to around, uh, one second and P95. Well, it was under 300 and, uh, three, three, three and a half seconds. So that's pretty good. It actually ended up meeting and beating Claude Sonnet after doing this little bit of extra effort. And I'm saving about a dollar a day in inference costs. So the important thing is after you've done something like this, the first thing is that you don't want to lose one ground. You'll want to make sure that, uh, you can update the prompt in the future or upgrade the modal. You don't want to lose one ground. You want to ensure that you can upgrade the, the model or change the prompt in the future. And it won't cause summaries to expand, to be a paragraph long, or to start hallucinating things that are untrue. You'll want to keep your evals running for that. And that would be a regression eval. And you run these sort of like you run, um, CICD tests, you know, uh, it's how you keep your CTO from blowing away your agentic experience by accident. One morning, true story happened to a founder friend of mine. So where do you get started? Uh, how many Claude calls could be llama calls? I challenge you to go home today and take a look at what you're sending to LLMs and ask yourself, is this something that a smaller model could handle and how much money would I save if I did that? When working on AI projects, keep an eye out for SLMs and specialized models that might be on device already. For instance, Chrome and the prompt API, they access Gemini Nano, which ships with Chrome, um, natively. And that can be really useful because that means you don't need to be shipping a model to anyone using the browser. You can take advantage of what's right there. Keep in mind that these models are more efficient and they meet your users where they are. Their information stays on device. You don't have to worry about PII and, you know, the energy consumption is, well, a blessing for the environment, so to speak. Remember to prototype big and deploy small. You may want to prototype your system with a foundation model and then convert parts of it to small language models and specialized models for production. Keep in mind you want to prove it, define it, test it, and then select your SAGE model that passes the test. You can use prompt engineering to help get better results from SLMs and close that gap with remote models. So I challenge you to consider your current implementation of prototype and convert one feature to use a smaller local model. Go home, try it out for yourself, and you might be surprised. Anyway, I look forward to see you out there on the web. If you enjoyed talking about small local models or the authentic web, you can follow me. I'm at nearestneighbors.com. If you want to try running some evals with your current setup, you can start testing with Phoenix at phoenix.arise.com. And lastly, you can sign up for the beta for MIMA at mima.social. I've been Rachel Neighbors, and it's been awesome chatting with you today. Go forth and build your own inference stack.