Open Reader

Voice In, Visuals Out: The Agony and the Ecstasy - Allen Pike, Forestwalk Labs

completed 13:05 Jun 28, 2026 Watch on YouTube

Current Status

completed

Video ID

65X0pQ6Lmbg

RAG / Chat

Enabled
Voice In, Visuals Out: The Agony and the Ecstasy - Allen Pike, Forestwalk Labs
Description

The latest AI models have made it possible to go past text-based chats, and build what Andrej Karpathy argues is the pinnacle of AI UX: voice-in, visuals-out. In this talk, Forestwalk Labs co-founder Allen Pike shares why this approach for LLM-powered product development is so useful, what's necessary to make it actually delight users, and lessons his team has learned building products with highly responsive AI agents like these – with a key focus on techniques for achieving low latency. Speakers: - Allen Pike (Forestwalk Labs): Allen is co-founder of Forestwalk Labs, runs the Infer AI engineering meetup, and hosts the It Shipped That Way podcast. X/Twitter: https://twitter.com/apike LinkedIn: https://www.linkedin.com/in/allenpike/ GitHub: https://github.com/apike

Summary

Generated by claude-sonnet-4-5

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: Voice-in + visual-out AI experiences are now feasible and superior to text-based chat if you solve the latency challenge with fast models, eager inference intervals, and aggressive caching.
  • Why it matters: Ken's operator stack needs low-latency voice agents for real-time workflow—this is a practical engineering brief on how to build them using existing models (Haiku, GPT-4o Mini) and prefix caching to stay under 1s response times.
  • Best use: Reference when designing real-time voice agent architecture; note the specific latency thresholds, model choices, and caching strategy for production voice workflows.

Executive Summary

Allen Pike from Forestwalk Labs argues that Andre Karpathy's thesis—voice is the preferred human input for AI, visuals the preferred output—is now practically achievable. The traditional text-in/text-out paradigm has given way to rich visual outputs (HTML, tool calls, illustrations) and voice input that works in real time. Pike's team built a voice agent that lives in their calls and can file Linear issues within a second of spoken intent, feeling 'perfectly natural' without interrupting conversation.

The central engineering challenge is latency. For seamless voice-to-visual experiences, responses must appear within ~1 second (the human attention threshold), not the 200ms needed for voice-to-voice conversation. Pike rejects slow models like GPT-4o Mini (P95 latencies of 5–10 seconds on OpenAI) in favor of Claude Haiku-class models, which reliably hit sub-second response times. He proposes a two-tier architecture: a fast 'real-time' model (Haiku or small open-source) handles continuous inference every 1–2 seconds as users speak, then hands off heavier tasks asynchronously to larger models.

Pike shares three critical techniques: (1) Use Haiku-class models on latency-optimized inference platforms, not just small models on slow platforms. (2) Send inference requests eagerly every 1–2 seconds during speech, not after waiting for silence, so the system reacts to partial sentences. (3) Implement stable prefix caching—keep the first 90% of the context identical across requests to get 90% cost and latency savings, then append only the new 10% and minimize output tokens. These techniques let Forestwalk deliver real-time voice-driven workflow automation that feels 'instant' rather than 'awkward and confused' like current Siri or ChatGPT voice experiences.

Key Takeaways

  • Claim: Voice-in + visual-out AI experiences are now feasible and preferable to text-in/text-out, per Karpathy's thesis. | Evidence: Forestwalk built a voice agent that lives in calls and files Linear issues within a second of spoken intent ('I said let's file that as a linear issue, and the voice agent within a second responded that it had done so'). Rich visual outputs (HTML, tool calling, illustrations) are now possible with recent model advances. | Caveat: Most current voice interfaces (Siri, ChatGPT voice) are 'slow and dumb'—Pike acknowledges users are down on voice because existing experiences are awkward and confused. Success depends on solving latency. | Implication: Ken should prioritize voice-in + visual-out for agent workflows, not voice-in + voice-out or traditional chat UX. The architecture is proven at Forestwalk and delivers higher bandwidth than typing (more words/min, more meaning per word). | Timestamp: 00:00–03:30
  • Claim: Seamless voice-to-visual experiences require <1s response latency, not the 200ms needed for voice-to-voice conversation. | Evidence: Pike cites human attention thresholds: 100ms feels 'instant,' 1000ms is the limit before people lose train of thought. For voice-to-voice, the limit is 200ms or conversation feels broken. For voice-to-visual, the more forgiving 1s envelope is sufficient ('if something appears on screen within a second of what you've said, it feels seamless'). | Caveat: Achieving even 1s is 'a ridiculous amount of work' when you chain network requests, speech-to-text, model inference, and network return. Most teams will miss this budget without careful architecture. | Implication: Ken's agent stack must target sub-1s end-to-end latency for voice-triggered visual responses. Voice-to-voice is harder (200ms) but voice-to-visual is achievable now with the right model and caching choices. | Timestamp: 03:30–05:30
  • Claim: GPT-4o Mini is too slow for real-time voice; Claude Haiku-class models are required. | Evidence: Pike reports P95 latencies of 5–10 seconds for GPT-4o Mini on OpenAI ('5,000ms, 7,000ms, P95 sometimes 10,000ms latency for the small model'). Haiku is 'much better in terms of that P95 latency.' The fast model handles real-time turns; heavy work is handed off asynchronously to larger models. | Caveat: No specific Haiku latency numbers are given, and Pike does not name the inference platform (likely Anthropic API or a low-latency provider). Open-source small models are mentioned as alternatives but not specified. | Implication: Ken should default to Haiku (or small open-source models on low-latency inference) for the real-time agent layer, not GPT-4o Mini. Use a two-tier architecture: fast model for continuous inference, larger model for heavy async tasks. | Timestamp: 06:00–07:30
  • Claim: Send inference requests every 1–2 seconds as the user speaks, not after waiting for silence. | Evidence: Traditional voice apps wait for seconds of speech + 1s of silence before inference, blowing the latency budget. Pike's system 'eagerly responds as the person is talking' every 1–2 seconds ('you have to have a model and infrastructure where you're able to get these fast turns going'). | Caveat: This requires the model and prompt to handle partial/incomplete sentences gracefully. Sentences like 'oh hey we're going to want to change this and actually also let's do this other thing' must be parsed incrementally. | Implication: Ken's voice agent should use streaming STT (e.g. Deepgram, AssemblyAI) with inference triggers every 1–2s, not end-of-utterance detection. The system must tolerate mid-sentence inference and update visuals as speech continues. | Timestamp: 07:30–09:00
  • Claim: Stable prefix caching is essential: keep the first 90% of context identical across requests to get 90% cost and latency savings. | Evidence: Pike describes 'prefix caching' (available on multiple platforms) where identical context prefixes are cached. 'You can get up to 90% cheaper, faster inference… the first 90% of the context window should be the same from request to request, then just use that final 10%, and minimize output tokens.' | Caveat: This assumes long-running or frequently running agents where context is stable. No guidance on prompt design to ensure prefix stability or how to handle context drift over time. | Implication: Ken's agent prompts must be structured so system instructions and static context (e.g. user profile, tool schemas) are identical across turns. Only append new user speech/events in the final 10%. Minimize output tokens to keep inference fast. | Timestamp: 09:00–10:30

Detailed Brief

Voice-in + visual-out as the new AI UX paradigm

  • Claims: Voice is the human-preferred input (per Karpathy), visuals the preferred output; Humans communicate faster and with more bandwidth via speech than typing (more words/min, more meaning per word—'okay' vs 'okay'); Rich visual outputs (HTML, tool calling, illustrations) are now feasible with recent model advances; Forestwalk built a voice agent that files Linear issues within a second of spoken intent during calls
  • Evidence: Karpathy's recent argument (last month) on voice-in, visuals-out; One-third of the human brain is dedicated to visual processing; Models can now generate 'rich HTML, tool calling, visualizations, interactive controls, beautiful illustrations'; Pike's team mentioned a Slack bug on a call, spoke 'let's file that as a linear issue,' and the agent responded within a second that it had done so; Voice is 'the ultimate way that as humans we communicate'; people 'jump on a call' for high-bandwidth communication
  • Caveats: Most current voice interfaces (Siri, ChatGPT voice) are 'slow and dumb'—users are down on voice because experiences are awkward and confused; Traditional text-in/text-out is still the dominant AI UX; voice-in, visuals-out is emerging but not yet mainstream
  • Implications: Ken should design agent workflows for voice-in + visual-out, not voice-in + voice-out or text chat; Visual outputs allow richer feedback (charts, controls, images) than voice responses; Voice input enables hands-free, high-bandwidth agent interaction during calls, meetings, or while multitasking

The latency challenge: hitting the 1s envelope for voice-to-visual

  • Claims: Voice-to-visual experiences require <1s response latency to feel seamless; Voice-to-voice requires 200ms or less; voice-to-visual is more forgiving (1s); 100ms is the threshold for computer reactions to feel 'instant'; 1000ms is the limit before people lose train of thought; Achieving sub-1s latency across speech-to-text, model inference, and network round-trip is 'a ridiculous amount of work'
  • Evidence: Pike cites research from the 1960s on the 100ms instant-reaction threshold; 1000ms (1 second) is the limit before 'people start to lose their train of thought'; For voice-to-voice conversation, 200ms latency is required for full back-and-forth without interruption; Thinking Machines/Neolab recently demoed a time-sliced architecture that does continuous inference in 200ms chunks for voice-to-voice; Voice-to-visual benefits from the 1s envelope because visuals appearing on screen within a second 'feel seamless'
  • Caveats: Novel architectures (like time-sliced continuous inference) exist for voice-to-voice but are not yet widely adopted; No specific latency numbers are given for Forestwalk's production system—only qualitative descriptions ('within a second'); Achieving 1s end-to-end requires optimization across STT, model, network, and rendering layers
  • Implications: Ken's agent stack must measure and optimize end-to-end latency from speech start to visual update; Voice-to-visual is the pragmatic choice today; voice-to-voice requires bleeding-edge architecture; Latency budget must account for STT, inference, network round-trip, and rendering—each layer must be sub-300ms to stay under 1s total

Three critical techniques for sub-1s latency

  • Claims: 1. Use a Haiku-class model on a latency-optimized inference platform, not just a small model on a slow platform; 2. Send inference requests eagerly every 1–2 seconds as the user speaks, not after waiting for silence; 3. Implement stable prefix caching: keep 90% of context identical, append only new 10%, minimize output tokens
  • Evidence: GPT-4o Mini on OpenAI had P95 latencies of 5–10 seconds despite being small and cheap; Claude Haiku is 'much better in terms of P95 latency' and can respond in hundreds of milliseconds; Fast model handles real-time turns; heavy work is handed off asynchronously to larger models; Traditional voice apps wait for speech + 1s silence before inference, blowing the budget; Forestwalk infers every 1–2 seconds as people talk, handling sentences like 'oh hey we're going to want to change this and actually also let's do this other thing'; Prefix caching on multiple platforms offers 'up to 90% cheaper, faster inference' if the first 90% of context is stable; Pike recommends structuring context so the first 90% is static (system instructions, tools) and only the final 10% changes per turn
  • Caveats: Haiku latency numbers are not specified; no mention of inference provider (Anthropic API vs. self-hosted vs. low-latency vendor); Open-source small models are mentioned as alternatives but not named or benchmarked; No guidance on prompt engineering to ensure prefix stability or how to handle context drift over long conversations; Eager inference every 1–2s requires handling partial/incomplete sentences; no detail on how this is implemented
  • Implications: Ken should benchmark Haiku vs. GPT-4o Mini vs. small open-source models on target inference platforms (Anthropic, Together, Fireworks, etc.) for P95 latency; Use a two-tier architecture: fast model (Haiku) for continuous inference, larger model (Opus, GPT-4) for async heavy tasks; Integrate streaming STT (Deepgram, AssemblyAI) with inference triggers every 1–2s, not end-of-utterance detection; Structure agent prompts with static prefix (system, tools, user profile) and append only new speech/events in the final 10%; Minimize output tokens (short responses, JSON tool calls) to keep inference fast

Notable Concepts & Terms

  • Voice-in, visuals-out (Karpathy thesis): Preferred AI UX paradigm: users speak (high-bandwidth input), AI responds with visual outputs (charts, HTML, tool calls) rather than voice or text
  • Tyranny of latency: The central engineering challenge for real-time voice AI: achieving sub-1s (voice-to-visual) or sub-200ms (voice-to-voice) response times across the full stack
  • 100ms instant-reaction threshold: Human perception threshold (since the 1960s): computer reactions under 100ms feel 'instant'; over 1000ms causes loss of train of thought
  • Prefix caching: LLM inference optimization where identical context prefixes are cached across requests, offering up to 90% cost and latency savings (available on Anthropic, OpenAI, others)
  • Two-tier agent architecture: Fast 'real-time' model (Haiku) handles continuous inference every 1–2s; larger model (Opus, GPT-4) handles heavy async tasks handed off by the fast model
  • Eager inference intervals: Sending inference requests every 1–2 seconds as the user speaks, rather than waiting for silence, to keep the experience feeling real-time
  • Haiku-class model: Pike's term for Claude Haiku or equivalent small models with sub-second P95 latency on low-latency inference platforms (contrasted with GPT-4o Mini on OpenAI)

Operator Notes / Why Ken Should Care

  • Pike's talk is a practical engineering brief, not theory—Forestwalk ships a production voice agent that files Linear issues in <1s. The latency thresholds (100ms instant, 1s attention, 200ms voice-to-voice) are critical for Ken's agent UX design.
  • The GPT-4o Mini latency data (5–10s P95 on OpenAI) is a red flag: even 'small' models can be slow if the inference platform prioritizes throughput over latency. Ken must benchmark on low-latency providers (Anthropic, Together, Fireworks) before choosing a model.
  • The two-tier architecture (fast real-time model + async heavy model) is a repeatable pattern for Ken's operator stack: Haiku for continuous monitoring/response, Opus/GPT-4 for deep reasoning when triggered.
  • Prefix caching (90% static context, 10% dynamic) is a must-have for real-time agents. Ken should audit current prompts to ensure system instructions, tool schemas, and user profiles are identical across turns, with only new events appended.
  • Eager inference every 1–2s requires streaming STT and prompts that tolerate partial sentences. This is a departure from traditional voice apps (wait for silence) and likely requires custom integration work.
  • Pike's call for collaboration ('I'd love to hear from anybody who has been exploring and experimenting') suggests Forestwalk is ahead of the curve but not alone—Ken should monitor the real-time AI agent space (Neolab/Thinking Machines, Fixie, others) for emerging architectures.

Watch Map

  • timestamp unavailable: Timestamps unavailable in transcript
  • 00:00–03:30: Intro: Karpathy's voice-in/visuals-out thesis, why visuals are preferred output, why voice is controversial input (Siri/ChatGPT pain), Forestwalk demo (file Linear issue in 1s)
  • 03:30–05:30: The latency challenge: 100ms instant, 1s attention, 200ms voice-to-voice, why voice-to-visual is more forgiving (1s envelope), time-sliced architectures for voice-to-voice
  • 06:00–07:30: Technique 1: Fast model on latency-optimized platform (GPT-4o Mini too slow at 5–10s P95, Haiku is better, two-tier architecture)
  • 07:30–09:00: Technique 2: Eager inference intervals (send requests every 1–2s as user speaks, not after silence, handle partial sentences)
  • 09:00–10:30: Technique 3: Stable prefix caching (90% static context, 10% dynamic, minimize output tokens, 90% cost/latency savings)
  • 10:30–end: Closing remarks: invitation to share learnings, call to build delightful experiences

Source/Metadata

  • Title: Voice In, Visuals Out: The Agony and the Ecstasy - Allen Pike, Forestwalk Labs
  • Transcript words: 1943
  • Duration seconds: 785
  • Timestamp note: Timestamps unavailable—transcript does not include chapter markers or MM:SS annotations. Watch_map estimated from flow of argument.

Transcript

1858 words en Processed in 134.2s

All right, I'm Alan Pike, and today I'm going to be sharing some of what we've learned building voice in, visuals out, experiences using AI. Here we go. So this is Andre Karpathy, and he made an argument last month that voice is the human preferred input for AIs, but that we prefer visuals as the output. And he knows a thing or two about this stuff, but that's not how we have been, for the most part, building or using AI. We've been typing to it, it's been typing back maybe with some markdown. But there's some breakthroughs over the last few months where both visuals in and audio, or audio in and visuals out experience is now really feasible. And we can create really delightful experiences with it. The visuals out is pretty intuitive, right? Of course, a third of our brain is dedicated to processing visual information. We love looking at things. And models have recently got to the point where they can generate rich HTML, tool calling. So we can have these experiences where there's visualizations that come back, explain things, help us understand, communicate the responses from these models. They can give us interactive controls that allow us to explore and understand and modify and change and direct the models. And they can even respond with beautiful illustrations and images, right? And so the ceiling on the visuals out piece has really lifted and we have a lot more capability of what we can do in terms of responding from these models. The more controversial half of Karpathy's argument though is this idea of voice as the preferred input. We have long idealized and fantasized about speaking to an AI and having this real time conversation where it understands what we need and it reacts appropriately in real time. But the experiences that most people have had so far with voice interfaces have been more like trying to get Siri to turn the lights on and it's not working. Or this guy that is trying to get ChatGPT voice mode to do things, but it keeps being awkward and confused, right? The models that we have so far, the experiences that most people have seen are both slow and dumb, which is not a great combination. And so a lot of people are down on voice as input. But speaking is the ultimate way that as humans we communicate. When we're speaking we have more words per minute than when we're typing. But we also convey more with each word. There's a huge difference between if you say to me, okay, versus if you say, okay. And that's why when we have something that is really important that we need to communicate or get through to somebody, we will jump on a call. And we will speak in person so that we can get through that high bandwidth communication. And so what we've done at Forest Walk is we have built an agent that is in our calls that can help us out in real time. And so the other day I mentioned to my co-founder on a call that I'd seen a bug with the Slack integration, or at least I thought I had. And she responded that she had seen the same bug. And so I just said, okay, well, let's file that as a linear issue. And the voice agent within a second responded that it had done so. And that feels perfectly natural when you get it really dialed in, that you are speaking either incidentally or with a purpose, intentionally speaking to the AI. And it responds in a way that is not interrupted view of you. It doesn't need to be voice. It is just taking action on your intent. And so this is, we're going to see this experience more and more in the coming months and years. But there's a huge barrier, a huge challenge to making this actually work and feel good. And that is the tyranny of latency. It's really difficult to get a response through that whole chain fast enough. We've known since the 60s that to have a computer react to us fast enough that it feels instant, it needs to react in about 100 milliseconds, a tenth of a second. And so, of course, we are always aspiring to get our products to react that fast. And sometimes we can achieve that, but with networking and everything, it can be challenging, depending on what work needs to get done. And so sometimes we might flex up to a thousand milliseconds, right? You get a full second. It's about the limit before people start to lose their train of thought. They ask a computer to do something, it takes more than a second, we're off to something else mentally, right? And so we're always in that play between those two kind of human limits of trying to respond within a second, ideally within 100 milliseconds. But if we want to have a seamless voice conversation with an AI or another human, the limit is much more aggressive. We need to have a 200 millisecond latency or less if we want to have a fully conversive, people are verbalizing something and they are interrupting or interjecting or agreeing and forming that sort of connection for a full voice in, voice out conversation. And so imagine trying to do that where you have a network request, maybe you turn speech to text, and then you do some sort of model inference on that text and then send it back on the network. It's a ridiculous amount of work to do in 200 milliseconds. There are some clever approaches of trying to work around that. Thinking Machines in Neolab has a really thoughtful architecture where they were demoing recently, just a few weeks ago, a time sliced into 200 millisecond chunks like that 200 millisecond goal and have a model that's doing continuous inference in and out of 200 millisecond slices. So there are ways around it for voice in voice out. But there's also we don't need to wait for novel architectures. We can just switch to having voice in visuals out and then we benefit from the more forgiving visual response envelope that people have. And if you have something that appears on screen within a second of what you've said, it feels like it has meant that it's keeping within your attention span, it's reacting to you in a way that feels seamless. And so we've been building for that over the last few months. And it's a lot of fun and we've learned a lot doing that. And so I want to finish up just by sharing three of the things that we found to be really important for getting within that latency envelope and making the experience feel really delightful for people who are using it. So the first one is that in order to respond in a way that feels seamless, you have to use a really fast model. So that's not just a model that is so small that it can respond fast enough in a few hundred milliseconds. Of course, it does need to be. But it also has to be on an inference platform that prioritizes latency. So when GPT-4o Mini came out, we were excited that, okay, well, yes, we get this more intelligent model and presumably it'll be fast. But then in practice, we are seeing latencies of 5,000 milliseconds, 7,000, P95, sometimes 10,000 millisecond latency for the small model. It was cheaper than the bigger model, but we weren't consistently or ever seeing it respond fast enough. Haiku is much better in terms of that P95 latency. So you really need a Haiku class model as the one that's responding in this real time or one of the smaller open source ones. And then if there is a larger chunk of work that needs to get done, then that model then hands off or sends off an asynchronous message to a larger model that can think. And then the continuous real time model or soft real time model that's responding quickly, then can we interleave in those responses if it's doing something heavier. So that's number one, you need to have a fast model, obviously with a short enough context that you're providing it that it can respond in a few hundred milliseconds. The second thing that's really critical if we're going to have this feeling of it being really responsive is to have short intervals of how quickly we're sending out for inference. Traditionally for a voice application, you might listen to the user for a few seconds and then they stop. You listen for a second of silence and then you have some sort of inference goes on and then you've blown your budget by a pretty wide margin just for waiting for silence. So if we're going to have something up on screen within a second, then when we want to have our inference pretty eagerly responding as the person is talking, even if we're not entirely sure they've stopped talking, being willing to send inference every one or two seconds as they speak, because we sometimes have a sentence where we ask, oh, hey, we're going to want to change this and actually also let's do this other thing. It feels much more seamless if you're doing those things as people talk. So you have to have a model and an infrastructure where you're able to get these fast turns going. And then finally, in order to actually make all of that work, you need to have a stable caching regimen. So we have this huge improvement in the way that we work with LLMs over the last year, where we have on the different platforms, different ideas of this prefix caching, where if the beginning of the context you send to the model is the same each time, then you can get up to 90% cheaper, faster inference, depending on the conditions. And so you need to lean heavily into this architecture. Really, I think for most applications, we're moving towards this, whether it's a long running agent or a frequently running agent. It's the same principle applies that we want the first 90% of the context window, if we can, to be the same from request to request, and then just use that final 10%, and then also, of course, minimize the number of output tokens, so that you can get these really fast and relatively affordable inference turns to create that delightful experience. So these are some of the techniques we've found really useful. I would love to hear from anybody who has been exploring and experimenting, whether it's with real-time, or any way that people are pushing the boundaries of creating delightful experiences with these models. I'd love to chat, share what we've been learning. And I hope some of this inspires some of you to go out and build something great. Thanks. so that you can get these really fast and relatively affordable inference turns to create that delightful experience. So these are some of the techniques we've found really useful. I would love to hear from anybody who has been exploring and experimenting, whether it's with real-time, or any way that people are pushing the boundaries of creating delightful experiences with these models. I'd love to chat, share what we've been learning. And I hope some of this inspires some of you to go out and build something great. Thanks.