The Prompt Is Still a Punch Card - Ted Johnson, JoinIn AI
Description
Interfaces outlive their constraints. The keyboard, command line, mouse, menus, forms, voice assistants, and even prompts were all brilliant compromises with the machines of their time. But AI gives us a chance to renegotiate that bargain. This talk reframes AI as an interface technology, not only an intelligence technology. We will trace a pattern across computing history: humans repeatedly learn the machine’s protocol, from punching cards to writing commands to engineering prompts. Then we will ask what changes when computers can reason, listen, speak, infer, clarify, and adapt. The next frontier is not just better models or better voices. It is more human-compatible interfaces: systems that understand timing, attention, interruption, ambiguity, repair, shared context, and when to stay silent. Speakers: - Ted Johnson (JoinIn AI): Ted Johnson is an executive, enterprise architect, and co-founder of JoinIn.AI, focused on AI-powered collaboration, enterprise architecture, cloud strategy, and technology transformation. LinkedIn: https://linkedin.com/in/johnsontedm GitHub: https://github.com/JoinIn-AI/
Summary
Generated by claude-sonnet-4-5At-a-Glance
- Verdict: Watch fully
- Core thesis: Current AI prompting is fundamentally batch protocol (like punch cards), mismatched with AI's capability for real-time conversational interaction—the interface itself, not users, is the bottleneck to adoption.
- Why it matters: This reframes AI UX design from 'better prompts' to 'real-time participatory protocols'—directly impacts agent system design, enterprise AI adoption, and the next generation of conversational AI products.
- Best use: Watch fully to internalize the channel/expression/protocol framework, then reference detailed sections when designing agent interactions or evaluating conversational AI systems. The punch-card metaphor is a powerful framing device for operator architecture.
Executive Summary
Ted Johnson (JoinIn AI co-founder, 25 years in enterprise software/collaboration) argues that despite AI's intelligence leap, we're still using 1860s interaction protocols. His core framework: Channel (physical medium: keyboard, voice), Expression (richness of meaning carried), and Protocol (rules of exchange). While LLMs exploded expression—natural language vs. command syntax—the protocol remains batch: assemble complete request, submit, wait, read output, retry. This is identical to punch cards, which inherited batch from weaving looms. Speed fooled us into thinking prompting is interactive; it isn't.
The mismatch is acute: model capacity (reasoning, memory, multimodality) curves upward while interface protocol stays flat. Humans still do all orchestration work—choosing context, timing, repairing output, engineering prompts. When it fails, users blame themselves ('I'm bad at AI'), but the fault is the interface design. Prompt engineering is rebranded punch-card operator skill: magic words, format tricks, incantations to package batch jobs correctly.
Johnson demos JoinIn's approach: real-time conversational AI that participates in group meetings, understands turn-taking, knows who's speaking to whom, back-channels ('mm-hmm'), yields when interrupted, and takes action only when contextually appropriate (e.g., capturing requirements mid-discussion without explicit prompts). He contrasts this with current voice-mode failures (models answering 'Hey Ted' as if addressed to them) and emerging real-time systems (OpenAI's back-channeling, NVIDIA Personaplex interruption handling).
The call to action: stop treating AI as 'smarter machine behind old interfaces' and start designing intelligence as interface technology itself. The question shifts from 'how do we encode intent better' to 'what burden are we putting on humans only because machines used to be limited?' The answer isn't always chat or voice—it's human conversational affordances (pauses, sketches, silence, modality-switching) handled by AI, removing translation/precision/context/repair taxes humans have paid for 75 years.
Key Takeaways
- Claim: Prompting is batch protocol identical to punch cards: assemble complete request, submit, wait, read output, retry. | Evidence: Punch-card workflow: encode request away from machine, carry deck to operator, submit job, wait hours/overnight, read printout, fix one thing, resubmit. Prompting: assemble request in box, submit, wait seconds/minutes, read response, rephrase, resubmit. Speed (seconds vs. overnight) creates illusion of interactivity, but structure is identical—machine engages only after complete packaged turn. | Caveat: Streaming responses and features like 'ask for updates' add interactivity, but at the core it remains batch with 'interactive sprinkles'—the fundamental protocol hasn't changed because the machine still can't participate during the human's thinking process. | Implication: For agent systems, this means current orchestration models (prompt → execute → return) are architecturally misaligned with AI's capability. Operators should explore real-time protocols where agents can interject, clarify, or yield mid-task rather than waiting for complete instructions. | Timestamp: 04:45-07:30
- Claim: Prompt engineering is rebranded punch-card operator mastery—learning magic words to package batch jobs correctly, not actual human-AI fluency. | Evidence: Examples of prompt engineering tricks: 'think step-by-step,' provide examples, 'be an expert' (or don't), pace context differently, use markdown documents, specific phrasings. These are incantations for packaging requests so jobs don't fail—identical to punch-card operator knowing exact deck assembly to avoid errors. | Caveat: Prompt engineering is useful and currently necessary; Johnson doesn't say it's useless, but rather that its necessity reveals the interface's limitations—it's a workaround, not a feature. | Implication: Organizations investing heavily in prompt-engineering training are essentially teaching punch-card skills. The real ROI is in interfaces that eliminate the need for such training, lowering adoption friction and expanding AI accessibility beyond power users. | Timestamp: 07:30-08:15
- Claim: Channel (physical medium), Expression (meaning richness), and Protocol (interaction rules) are the three design layers—AI advanced Expression massively but Protocol stayed stuck. | Evidence: Channel evolution: keyboard patent from 1860, Hansen writing ball, Dvorak/Colemak layouts—arbitrary legacy devices. Expression progression: assembly (op codes) → shell commands → programming languages (primitives) → natural language (ocean of meaning). Protocol: batch from weaving looms → punch cards → prompts, unchanged. Model capacity (reasoning, vision, memory) shoots up; protocol flat. | Caveat: The framework is conceptual/analytical rather than empirically validated—it's a lens for understanding the mismatch, not a measured phenomenon. The 'flatness' of protocol is somewhat subjective (streaming, tool use, memory features do add protocol complexity). | Implication: For AI product design, this framework suggests auditing where investment goes: most resources go to model capability (expression), little to protocol innovation. Rethinking protocol could unlock adoption without needing more capable models. For Ken's work, this is a strategic framework for evaluating conversational AI products and agent architectures. | Timestamp: 02:45-10:00
- Claim: Real-time conversational AI requires understanding multi-party turn-taking, speaker identification, and contextual action-triggering—not just back-channeling sounds. | Evidence: JoinIn demo: AI in group meeting labels statements (question, proposal, answer), only acts when appropriate, knows 'AI, capture that' vs. side conversation, resolves scope without explicit prompt ('Expense Approvals. 5,000 Threshold'—AI synthesizes decision from discussion). Contrast: co-founder's example where voice mode answered 'Hey Ted, come on in' because it has no speaker-awareness. OpenAI adding 'mm-hmm' back-channels; NVIDIA Personaplex yields when interrupted. | Caveat: JoinIn's demos are controlled/designed scenarios, not arbitrary real-world chaos. The hard problems—noisy rooms, ambiguous pronouns, implicit context, hallucination when inferring intent—are acknowledged but not solved on-screen. 'Making listening noises is not the same as knowing who's in the room' is Johnson's own caveat. | Implication: For agent operators, the lesson is that true conversational AI needs explicit models of discourse structure, turn-taking norms, and context-tracking—not just better prompts or faster responses. This points toward building agent systems with explicit conversational state machines, speaker tracking, and action-gating based on discourse cues, not just semantic content. | Timestamp: 11:30-17:45
- Claim: The interface mismatch causes user blame—people think they're 'bad at AI' when the protocol is at fault, not their skill. | Evidence: Quote: 'When the magic words don't land, people blame themselves. They decide they're bad at this, not specific enough, don't get AI. I want to say as clearly as I can: it is not our fault. We are being asked to operate brand new intelligence through a protocol of a punch card. The mismatch isn't the user, it's the interface.' | Caveat: This is a normative/advocacy statement rather than cited user research data—it's Johnson's interpretation of adoption friction, not empirical study results. However, it's consistent with known UX principles (blame attribution, learned helplessness). | Implication: For GTM and adoption strategies, this reframes the problem from 'users need training' to 'interface needs redesign.' It suggests that AI companies blaming low adoption on user education are missing the structural issue—echoes Clayton Christensen's 'jobs to be done' thinking where users 'hire' products that don't require learning new protocols. For Ken's investing lens, this is a bet on interface-first AI startups over model-first. | Timestamp: 09:30-10:45
- Claim: AI should be seen as interface technology, not just intelligence technology—the design question becomes 'what burdens are we putting on humans only because machines used to be limited?' | Evidence: Historical taxes: translation tax (encode intent in machine syntax), precision tax (exact commands), context tax (re-explaining), repair tax (fixing output). Johnson: 'Stop picturing AI as smarter machine hiding behind prompts/agents/loops. Start seeing intelligence itself as thing that can remove interface constraints.' Examples: AI choosing modality/timing, not human; AI knowing when to speak vs. stay silent; AI understanding 'a pause, a sketch, a checklist, a quiet aside.' | Caveat: The vision is aspirational and directional rather than prescriptive—Johnson doesn't specify technical implementation beyond conversational examples. 'The answer isn't always chat, isn't always voice' acknowledges modality diversity but doesn't resolve when/how to choose or blend them. | Implication: For AI product strategy, this suggests designing from 'remove human burden' rather than 'add AI feature.' For Ken's operator systems, it means agent workflows should infer context, self-correct, and adapt modality—not require humans to pre-configure, monitor, and repair. It's also a contrarian bet: if true, the next AI winners aren't LLM trainers but interface innovators—companies rethinking the entire interaction layer. | Timestamp: 17:45-19:30
Detailed Brief
The Punch-Card Metaphor: Channel, Expression, Protocol Framework
- Claims: The keyboard is arbitrary legacy tech (1860 QWERTY patent, could've been Hansen writing ball), yet we put it between humans and super-intelligence; Channel = physical medium (keyboard, mic, screen, punch card, prompt box); Expression = richness of meaning carried (op codes → commands → primitives → natural language); Protocol = interaction rules (batch: assemble, submit, wait); AI's leap is in Expression (natural language carries context/nuance/intent), not Channel (still typing in box) or Protocol (still batch); Batch protocol inherited from weaving looms (set pattern, run cloth) → punch cards → prompts; we shrank wait time but kept the structure
- Evidence: Keyboard enthusiast examples: Dvorak/Colemak layouts, multi-key thumb clusters to avoid 'wasting' thumbs on spacebar; Expression timeline: assembly (few dozen op codes) → shell (commands/flags) → modern languages (composable primitives) → LLMs (full human language); Punch-card workflow detailed: sit away from machine, encode request, carry deck to operator, submit, wait hours/overnight, read printout, fix, resubmit; Current prompting: assemble request, submit, wait seconds/minutes, read, rephrase, resubmit—'speed fooled us into thinking it's interactive, it isn't'
- Caveats: Not all keyboards are QWERTY legacy—mobile/touch changed input, but point is about protocol persistence, not specific hardware; Streaming responses, tool use, memory features do add interactivity beyond pure batch—Johnson calls these 'interactive sprinkles' but acknowledges them; Framework is conceptual lens, not empirical measurement—'flat' protocol is relative to model capability growth, not absolute lack of evolution
- Implications: For operators: audit whether agent orchestration is still 'prompt → execute → return' batch loops vs. real-time participatory design; For investors: interface-protocol innovation may be higher leverage than marginal model improvements—look for startups rethinking interaction layer, not just fine-tuning; For product teams: the 'speed = interactivity' illusion means users may tolerate current UX until they see true real-time alternative—adoption curve could be steep once better protocol ships
Prompt Engineering as Punch-Card Operator Skill
- Claims: Prompt engineering is learning magic words to package batch jobs correctly, not mastery of AI itself; Examples: 'think step-by-step,' provide examples, 'be an expert' (or don't), pace context, use markdown, specific phrasings—all incantations to avoid job failure; Getting good at prompting should bother us, not reassure us—it means we've learned the black box's quirks, not that the interface is good; When prompts fail, users blame themselves ('I'm bad at AI, not specific enough'), but the fault is the protocol mismatch, not user inadequacy
- Evidence: Johnson's personal observation: 'We trade incantations. We've learned the magic words. That's the illusion—feels like mastery, but it's same mastery a punch-card operator had.'; Quote: 'It is not our fault. We are not bad at using AI. We are being asked to operate brand new intelligence through protocol of punch card. Mismatch isn't the user, it's the interface.'; Historical parallel: punch-card operators knew exactly how to assemble deck so job wouldn't fail—same as knowing exact prompt format today
- Caveats: Prompt engineering is currently useful and necessary—Johnson doesn't advocate abandoning it, but rather sees its necessity as symptom of deeper design problem; User blame claim is interpretive (Johnson's view from career in usability) rather than cited research—though consistent with UX principles on learned helplessness and attribution; Some prompt techniques (chain-of-thought, few-shot) do leverage model architecture, not just 'magic words'—but point is they're workarounds for protocol limits
- Implications: For enterprise adoption: investing in prompt-engineering training may be necessary short-term but strategically dead-end—like teaching punch-card skills in 1990; For product differentiation: interfaces that eliminate need for prompt engineering have massive adoption/accessibility advantage—this is the 'remove friction' bet; For Ken's content/GTM: framing current AI literacy efforts as 'temporary workaround' vs. 'durable skill' helps position where to invest in education vs. where to wait for better tools
Real-Time Conversational AI: Protocol Evolution in Practice
- Claims: True conversational AI requires turn-taking, speaker identification, contextual action-gating—not just faster responses or back-channeling sounds; Current voice modes fail at multi-party context: co-founder's 'Hey Ted' example where AI answered because it has no speaker-awareness; Emerging systems (OpenAI back-channels, NVIDIA Personaplex interruption handling) show field converging on participatory protocols; JoinIn's approach: AI in group meetings labels discourse acts (question, proposal, answer), only takes turn when appropriate, resolves context without explicit prompts
- Evidence: Voice mode failure: 'When is next Timberwolves game?' (answered correctly) → 'Hey Ted, come on in' (AI responds 'Sure, I'm here, what's on your mind?') because protocol has one slot: your turn, its turn; OpenAI GPT real-time: now back-channels ('mm-hmm,' 'right') to show active listening; NVIDIA Personaplex demo: interrupted during diet advice, yields, switches to marathon topic, picks thread back up—'real turn-taking, listening and speaking at once'; JoinIn demo: group meeting about expense approvals, AI tracks 'REQ-142,' understands '$5,000 threshold' decision emerged from discussion, captures requirement without 'AI, write this down' prompt, later corrects threshold to $10k on request, answers unrelated question ('is room free after?') only when directly addressed
- Caveats: Demos are designed scenarios, not arbitrary real-world chaos—noisy rooms, overlapping speech, implicit context, hallucination risks acknowledged ('making listening noises not same as knowing who's in room'); Speaker tracking 'easy with direct reference (AI, do X), will happen again without direct reference' but no detail on how that's solved; Group conversation understanding is hard problem—JoinIn working on it, but not claiming it's solved at scale
- Implications: For agent architecture: real-time participation requires explicit conversational state (who's speaking, what's the discourse structure, what's the current goal), not just semantic parsing; For product roadmaps: back-channeling/interruption-handling are table stakes; next frontier is multi-party context and action-gating based on discourse cues; For Ken's operator systems: consider building agents with turn-taking models (when to act vs. listen vs. clarify) rather than just LLM calls—this is protocol-level design, not prompt-level
AI as Interface Technology: The Burden-Removal Design Question
- Claims: AI is not just intelligence technology but interface technology—book-smart models alone aren't enough; For 75 years, humans adapted to machines (syntax, forms, timing, batch); now intelligence should meet humans partway or all the way; Design question shifts: 'What burden are we putting on humans only because machine used to be too limited to carry that burden itself?'; Answer isn't always chat/voice/markdown—it's human conversational affordances (question, pause, sketch, checklist, silence) with AI choosing modality/timing
- Evidence: Historical burdens: translation tax (encode intent in syntax), precision tax (exact commands), context tax (re-explain), repair tax (fix output)—'every step carried old constraint forward'; Quote: 'Stop picturing AI as smarter machine hiding behind prompts/agents/loops. Start seeing intelligence itself as thing that can remove interface constraints and amplify human potential.'; Johnson as usability person: 'When you take that burden off people, friction disappears and adoption follows.'; Final quote: 'If machine can finally understand more of what we mean, we can and should stop reshaping ourselves to be understood by it.'
- Caveats: Vision is aspirational/directional, not technically prescriptive—doesn't specify how to implement 'AI chooses modality' or when to use sketch vs. text; Sherry Turkle (MIT) quote about conversation as 'most human/humanizing thing' is philosophical framing, not empirical claim about interface design; Removing all burden risks over-automation—some human control/choice may be desirable, especially in high-stakes domains (Johnson doesn't address this tension)
- Implications: For AI product strategy: invert the build sequence—start from 'what human burden can we remove?' rather than 'what AI capability can we add?'—this is different from feature-driven roadmaps; For Ken's investing thesis: bet on companies innovating interface protocol, not just model capability—if Johnson's right, next AI winners are interface-first (JoinIn, real-time conversational platforms), not model-trainers; For agent operators: design agents to infer context, self-correct, adapt modality rather than requiring humans to configure/monitor/repair—this is 'burden removal' in practice; For GTM: pitch AI as 'finally fluent with humans' vs. 'smarter tool you need to learn'—this reframes adoption from training problem to interface problem
Notable Concepts & Terms
- Channel, Expression, Protocol: Johnson's three-layer framework for interface design. Channel = physical medium (keyboard, voice, screen). Expression = richness of meaning the channel can carry (op codes → natural language). Protocol = rules of interaction (batch, conversational, real-time). Key insight: AI advanced Expression massively but Protocol stayed stuck at batch.
- Batch protocol: Interaction model where human assembles complete request, submits to machine, waits for response, then acts—inherited from weaving looms → punch cards → prompts. Contrasted with real-time/conversational protocols where machine can participate during human's thinking.
- Prompt engineering as punch-card operator skill: Reframing prompt tips/tricks (think step-by-step, provide examples, etc.) as workarounds for batch protocol limits, not durable AI fluency—analogous to knowing exact punch-card deck assembly to avoid job failure.
- Interactive sprinkles: Johnson's term for features like streaming, tool use, memory that add interactivity to batch protocol but don't fundamentally change it—still submit-and-wait structure underneath.
- Translation tax, precision tax, context tax, repair tax: Historical burdens humans paid to use computers: encoding intent in machine syntax (translation), giving exact commands (precision), re-explaining context (context), fixing wrong output (repair). AI can remove these by meeting humans in natural language and conversational flow.
- Real-time conversational AI: Systems that participate during conversation, not just respond after complete turns—includes turn-taking, back-channeling (mm-hmm), yielding when interrupted, speaker identification, discourse act labeling (question, proposal, answer), contextual action-gating. Examples: OpenAI real-time mode, NVIDIA Personaplex, JoinIn's group-meeting AI.
- Speaker identification and discourse acts: In JoinIn's system, the AI tracks who's speaking, labels their statements (question, proposal, answer), and only takes action when contextually appropriate (e.g., when directly addressed or when a decision resolves). Critical for multi-party conversation where not all speech is directed at AI.
- AI as interface technology: Johnson's core reframing: AI isn't just making machines smarter, it's enabling entirely new interaction protocols—intelligence itself removes interface constraints rather than hiding behind old ones (prompts, forms, batch).
Operator Notes / Why Ken Should Care
- For agent orchestration: current 'prompt → execute → return' loops are batch protocol—consider designing agents with real-time participation: ability to interject questions, clarify ambiguity, yield when not relevant, choose modality dynamically
- For conversational agents: implement explicit turn-taking models (when to act vs. listen vs. clarify), speaker tracking, and discourse structure awareness—don't rely solely on semantic parsing of content
- For product differentiation: interfaces that eliminate prompt engineering have massive adoption advantage—this is 'remove friction' strategy, not 'add features'
- For GTM/positioning: frame AI as 'finally fluent with humans' vs. 'smarter tool requiring training'—shifts adoption from education problem to interface problem, potentially faster/broader uptake
- For investing: Johnson's bet is interface-first AI (JoinIn, real-time conversational platforms) will outcompete model-first companies—marginal model improvements less leverage than protocol innovation
- For content strategy: the punch-card metaphor is powerful framing device—use it to explain why current AI feels awkward despite being smart (protocol mismatch), not user inadequacy
- For workflow design: apply 'burden removal' question to every human task in AI workflow—what are we asking humans to do only because machines used to be limited? (choose context, repair output, engineer prompts, remember to ask, decide timing)
- For enterprise adoption: prompt-engineering training may be necessary short-term but strategically dead-end—like teaching punch-card skills in 1990—bet on interfaces that eliminate the need
- For multi-agent systems: JoinIn's group-conversation model (discourse acts, contextual action-gating) is blueprint for agents in shared spaces—not all agent output should be broadcast, agents should understand conversational floor
- For Ken's operator work: this talk is strategic blueprint for rethinking agent interaction layer—channels/expression/protocol framework is analytical tool for auditing current systems and designing next-gen ones
Watch Map
- 00:00-02:30: Intro: making prompting feel strange again—keyboard as arbitrary legacy (QWERTY, Hansen writing ball)
- 02:30-05:00: Framework: Channel (medium), Expression (meaning richness), Protocol (interaction rules)
- 05:00-07:30: Punch-card metaphor: batch protocol (assemble, submit, wait) unchanged from looms → punch cards → prompts
- 07:30-09:30: Prompt engineering as magic words—rebranded punch-card operator skill, not AI fluency
- 09:30-10:45: User blame: people think they're bad at AI when protocol is at fault—'it is not our fault'
- 10:45-12:00: Voice mode failure example: 'Hey Ted' answered by AI because protocol has no speaker-awareness
- 12:00-13:30: Emerging systems: OpenAI back-channels, NVIDIA Personaplex interruption handling—field converging on participatory protocols
- 13:30-17:45: JoinIn group-meeting demo: AI tracks discourse acts, resolves context without prompts, only acts when appropriate (expense approval example, threshold correction, room availability)
- 17:45-19:30: AI as interface technology: burden-removal design question—what are we asking humans only because machines used to be limited?
- 19:30-20:10: Conclusion: if machine can understand us, we should stop reshaping ourselves to be understood by it
Source/Metadata
- Title: The Prompt Is Still a Punch Card - Ted Johnson, JoinIn AI
- Transcript words: 4909
- Duration seconds: 1213
- Timestamp note: Timestamps estimated from 1213-second (20:13) video; transcript includes repeated sections but watch_map reflects logical flow
Transcript
I'm sure that sometime in the last few hours most of you did this. You typed a request into a small box to a super intelligence and then you waited. You watched the cursor blink, maybe a little throbber cycled through clever gerunds like hullabalooing, tomfoolooing, and philosophizing to hide the weight. Maybe it gave you what you wanted, maybe you rephrased it and tried again. It all felt completely normal. I want to spend the next 20 minutes making prompting feel unfamiliar and strange again. I'm Ted Johnson, co-founder of Join in AI. During my 25-year career building enterprise software, collaboration systems, and AI-enabled interfaces, I've always focused on human interaction. I've also been following AI for two decades, including back to the far less impressive GPT-1 and 2. And when ChatGPT arrived, I felt two things at once. First, an unsurprisingly amazement, knowing the world would never be the same, followed by actually surprising disappointment I couldn't shake. This disappointment turned into an observation that started a company, Join in AI, and that I keep coming back to, which is why do we still have to learn AI? Why does something this powerful so often feel unnatural to use? Here's the path we'll take to answer that. We'll start with the most familiar computer interface and make it strange again. Then I'll give you three key concepts. The channel, the physical transport that carries your intent, expression, the range and richness of meaning the channel can carry, and the protocol, the shape or rules of this exchange. I'll use those three concepts to show you that the prompt is our present day punch card. We'll share examples of the ways interfaces could progress and we'll wrap with some practical advice for AI and human centered design. Everyone knows what this is, the keyboard. It's everywhere and it feels completely normal or natural, it feels like it's a real problem. But it isn't. We all had to take lessons. We all had to practice. And I say this as someone who loves keyboards. But it seems that we've been trying to fix them as long as they've been around. People have tried more efficient layouts like Dvorak or Colemak to save their fingers some work. A more extreme example, some keyboard enthusiasts refuse to squander their two digits on the space bar, giving them four or eight or ten keys to press with that efficient thumb of theirs. And what do we even mean by the keyboard? Here's the patent drawing for the layout we use every day. This patent's from about 1860. And my personal favorite, it just as could have easily been the Hansen writing ball, which looks anything but like a way you want to talk to a super intelligence. We carry these legacies of an arbitrary input device designed under constraints that haven't existed for a century. And we put it between ourselves and the most capable machines ever built. Nobody alive chose it. We all inherited it. And then we stopped noticing. So that's the first idea. The channel. The medium an interface gives you to work in. The keyboard is a channel. A microphone, channel. A screen, a punch card, a prompt box, all channels. And channels matter because each one can physically carry a different kind of signal. For example, text is a stream of discrete symbols. Voice can carry timing, pitch, hesitation, and words. A diagram can carry spatial relationship all at once. But these are differences in what the medium can transmit. It's bandwidth, not the differences in the meaning. And carrying more signal isn't the same as the machine understanding any of it. That's the separate question. That's the next idea. Humans use all these channels constantly without thinking. We never pick one channel and force everything through it. That would be absurd. And yet, that's what we ask people to do with machines over and over. Hold on to the word channel because here's the plot twist. With AI, the channel never really changed. You're still typing into a box, but what you were allowed to push through it was about to. For the first time, the computer channels carry rich, complete human language. Notice I said what it carries, not what it is. You're still typing into a box. You're still hitting the submit button. The keyboard didn't change. What improved is the range of what you're permitted to express through it. That's the second idea. Expression. How much of what you actually communicate or mean will the interface let through. Here's an example of expression progress over time with computers. Starting with assembly, which gave you an instruction set, a few dozen op codes, then came the commands with the shell, inputs, flags, then modern programming languages gave you primitives that you could compose. While powerful, step by step, each one of these is a fixed vocabulary requiring you to express your intent by choosing from a menu the machine will accept. Natural language blew that menu open. For the first time you can say almost anything the way you'd say it to another person. And on the expression axis the leap is real and enormous. There's an ocean of meaning in an ordinary human request. Context, nuance, intent. All things we've never had to spell out to each other. For the first time you can say almost anything the way you'd say it to another person. For the first time a machine can take it in. So here's what should bother us as engineers and designers, as it's bothered and inspired me, the channels for computers have been the same for 150 years, some 180 years. Now with AI we've poured an ocean of expression into it. So why does it feel like we're still sipping through a straw and struggling to learn how the AI thinks? Because there's a third idea underpinning the other two. And it's really the one that hasn't kept up. The protocol. The rules you follow and the shape of the interaction itself. Channel stayed the same. Expression exploded in the last three years with LLMs. But the protocol, prompting, is the protocol of a punch card. And the punch card's protocol is good old batch. Here's what punch card batch meant. You sat down, away from the machine, carefully encoded your entire request in advance, carried your deck to the operator, you submitted the job, and then you waited. Sometimes hours, sometimes overnight. Then you read the printout, found one thing that was wrong, fixed it, resubmitted it, and waited again. The machine never engaged with you while you were thinking. It engaged with the finished package after the fact. Now let's look at the prompt. Assemble the whole request, submit it, wait, read what comes back. Something's off, assemble it again, submit again, and wait. We have to acknowledge that there are features, interactive features improving this. You can ask for updates, you can ask for summaries of what was done. But at the end, it's still batch with interactive sprinkles. It's the same protocol. We shrank the wait time from overnight to a few seconds or a few minutes, and the speed fooled us into thinking that it'd become interactive. It hasn't. It's still batch. You still package a complete turn before the machine is allowed to participate. We learn tricks, send tips to use code skills or rewrite prompts a certain way to manage Something's off, assemble it again, submit again, and wait. We have to acknowledge that there are features, interactive features improving this. You can ask for updates, you can ask for summaries of what was done. But at the end, it's still batch with interactive sprinkles. It's the same protocol. We shrank the wait time from overnight to a few seconds or a few minutes, and the speed fooled us into thinking that it'd become interactive. It hasn't. It's still batch. You still package a complete turn before the machine is allowed to participate. We learn tricks, send tips to use code skills, or rewrite prompts a certain way to manage this. And speaking doesn't change it. Your voice just gets transcribed into the box and submitted. Shorter batch is still batch. Because the protocol is the part that did not advance. The protocol is the part we've had to learn. We just gave it a flattering name. We call it prompt engineering and treat it like it's a power user skill. Strip the label off and it's a set of rules for packaging up good old batch. For example, tell it to think step by step. Give it examples. Ask it to be an expert. Don't ask it to be an expert. Don't ask it that way. Pace more context. Pace less context. Only talk to it through markdown documents. We trade incantations. We've learned the magic words. That's the illusion. It feels like mastery, but it's the same sort of mastery a punch card operator had. Knowing exactly how to assemble the deck so the job wouldn't fail. Moreover, we've gotten good at prompting these black boxes, and that's the part that should bother us, not reassure us. None of this means prompts are bad. Punch cards weren't bad. Command lines aren't bad. They're brilliant solutions for constraints of their time. But that's the whole question. Is batch still the right protocol? Are we still prepackaging our intent for a machine that no longer needs us to? Because it shouldn't need us anymore. It can ask a follow-up. It can clarify mid-thought. It can notice it's missing something and say so. It should be human conversational. Sherry Turkle of MIT puts it very well. Conversation is the most human and humanizing thing we do. It's one of humanity's superpowers. The capacity to engage and think is right there. And yet, we're still making people submit the deck and wait for the run. Even the punch card inherited its protocol, in fact. Batch came from the weaving loom. You set the whole pattern in advance, then ran the cloth. The punch card got reused on computers by default. We're still standing at the same moment again. AI could finally meet us in the middle of a thought, got handed a protocol of a loom. That's what I mean. The prompt is still a punch card, not because of how you encode it. The encoding is powerful and awesome. Because when the LLM is allowed to engage only after you've packaged a complete turn and submitted it. And this is where the mismatch bites. Model capacity is shooting straight up. Reasoning, speech, vision, memory, planning, all curving upwards. The interface protocol, flat. Still a box. Still a submit button. Still the human doing all the work around the LLM. The human still decides what context matters. Still remembers what to ask. Still chooses the timing. Still notices the ambiguity. Still repairs the output. Still has to carefully engineer a prompt. But the intelligence feels magical. It's the interface that still feels like work. And when it feels like work, when the output's wrong, when the magic words don't land, people blame themselves. They decide they're bad at this. They're not specific enough. They don't get AI. I want to say as clearly as I can, it is not our fault. We are not bad at using AI. We are being asked to operate a brand new kind of intelligence through a protocol of a punch card. The mismatch isn't the user, it's the interface. In the race to enable AI, we shortcut the interface. Okay, let's make this concrete and familiar. A few weeks ago, my co-founder was using a frontier company's voice mode. These are known as speech-to-speech models. He asked it a normal question. When is the next Timberwolves game? Fine. It answered it quickly. Then, he pretended I showed up as if to speak to me and said, "Hey Ted, come on in." He wasn't talking to the AI, but these models have no way to know that. So the AI did the only thing a prompt box can do. It took his speech as a turn and answered it. "Sure, I'm here. What's on your mind?" That's not a good answer, but it's not a dumb model. It answered the first question perfectly, but it's a protocol with exactly one slot. Your message, then its reply. It has no concept of who's speaking, whether the words were even meant for it. And the frontier companies want to make strides as well. OpenAI released GPT real-time too in late May and started trying it for their voice mode as well. It back channels now. It goes, "mm-hmm," and "right." The little sounds we make to show we're listening actively. We're seeing the field is converging on the same conclusion we built our company on. The interface has to stop being batch and start participating. Others are working on real-time conversation as well. This is NVIDIA's Personaplex, a research model, not ours. Watch what happens when it gets interrupted. "I've been thinking about starting a diet." "Yeah, starting a diet can feel a bit daunting, but you could keep it simple. Focus on eating more veggies and fruits." "Before I forget, I signed up for a marathon." "All right, congrats on signing up for the marathon. That's a big challenge. You've got a lot of time. Focus on building a solid base with regular long runs. Stay hydrated. Make sure you fuel right before and after. And don't forget to stretch and take care of your feet." Personaplex stops. It yields. It picks the thread back up. That's real turn-taking. Listening and speaking at once. In real time. "You need to come visit me." "Oh, okay." "Because you can go into the city." "Okay." "Because that's the thing. There's the random spray paint, but then there's also..." "I'm not sure... People must commission them." "Yeah, I think they do." "These massive mural spray paint pieces." "Yeah, I think they do." And Personaplex's back channels listens and lands where a person would. Beyond conversational flow, there are lots of challenges, and it's a complex problem. Making listening noises is not really the same as knowing who's in the room. These are not trained to tell that, "Hey, Ted" wasn't meant for it. But we are. We're working on improving the protocol to the models by giving it a better understanding of human and group conversation. "Good afternoon, everyone." "Good afternoon, Sam." "Hi, Jordan." "Good afternoon. Good to see you both." "Quick one. Which requirement is this? Do we have an ID?" "This is REQ 142. Expense Approvals." Beyond conversational flow, there are lots of challenges, and it's a complex problem. Making listening noises is not really the same as knowing who's in the room. These are not trained to tell that, hey, Ted wasn't meant for it. But we are. We're working on improving the protocol to the models. By giving it a better understanding of human and group conversation. Good afternoon, everyone. Good afternoon, Sam. Hi, Jordan. Good afternoon. Good to see you both. Quick one. Which requirement is this? Do we have an ID? This is REQ 142. Expense Approvals. There. It answered a question. It only takes actions based on the utility-driven model. So it creates goals to fulfill as it labels each of the participants' statements as a question, a proposal, an answer, and then only takes a turn when no one else is speaking or holding the floor. Right. We need users to approve requests faster. Yeah. The approval flows too slow. What kind of requests, though? Expense Approvals first. Access Requests eventually. AI. Hold that. Actually, let's pause. Expense Approvals. Expense Approvals. Or a General Approval workflow. Expense Approvals first release. Access Requests are future scope. That changes the data model. Good to know. AI. Pull that up for everyone. Tracking and determining who the speaker is referring to is critical. In this case, it was easy with a direct reference to the AI. But it will happen again without a direct reference. I'd forgotten that was a rule. So over 5,000 needs a second approver? Yep. Manager plus finance. So a big one can't be a single tap. Right. Over the limit, it routes to a second approver. Agreed. Under 5, one tap's fine. Works for me. Okay. Agreed. Expense Approvals. 5,000 Threshold. Right there. The AI resolved the scope objective. No one wrote the prompt. No one packaged the turn and hit submit. The system was in the conversation, following it, understanding, and choosing its moment. AI. Capture that for us. First release supports expense approvals only. Access requests are out of scope. Managers can approve or reject an expense right from a notification. And per the finance controls policy, anything over $5,000 routes to a second approver. Actually, make the threshold 10,000, not 5. Want me to update the requirement to a 10,000 threshold? Yes. AI, is this room free after the meeting? Let me check. The room looks free after this, but yes. It's yours until 3 o'clock. And that's the difference between a smart machine behind the same old prompt and an interface that finally participates. Here's the mindset shift I want to leave you with. AI is not just an intelligence technology. It's increasingly becoming an interface technology. And if so, then book smart models alone are not enough. Stop picturing AI as a smarter machine hiding behind prompts, agents, loops, and all the old paradigms. We have to start seeing intelligence itself as a thing that can finally remove interface constraints and amplify human potential. For 75 years, humans adapted to the machine. Its syntax, its forms, its timing, its batch. A system that can reason, listen, infer, adapt, should be able to meet us part way, if not all the way. Instead, if AI is for users, then we should obsess about maximizing the interface. So then the design question changes. What burden are we still putting on humans only because the machine used to be too limited to carry that burden itself? Ask that question and the whole interface space opens up. The answer isn't always chat. It isn't always voice and not a wall of markdown. It's definitely not a decade-old set of digital constructs. The right answer is the affordance humans already use with each other. Communication, a question, a pause, a sketch, a checklist, a quiet aside, or saying nothing at all. An interface where timing and modality aren't the human's job anymore, where choosing the right channel at the right moment is done by the AI. And as a usability person, this is the part that excites me the most. When you take that burden off people, the friction disappears and adoption follows. Computing has mostly been about improving how humans encode their intent for machines. The punch card, type a command, click a menu, use your thumb on an iPhone, write a prompt. Every step was progress and every step carried the old constraint forward into the next era. A translation tax, a precision tax, context tax, repair tax. AI is our chance to put those down. Not by making everything magical, not by making everything voice, not by replacing human judgment, but by making computers for once more fluent with us. Most talks and videos cover how to use or adopt AI. The deeper question is how AI intelligence changes the interface. Human conversation is the most human thing we do because if a machine can finally understand more of what we mean, then we can and should stop reshaping ourselves to be understood by it. Thank you. humanizing thing we do. It's one of humanity's superpowers. The capacity to engage and think is right there. And yet, we're still making people submit the deck and wait for the run. Even the punch card inherited its protocol, in fact. Batch came from the weaving loom. You set the whole pattern in advance, then ran the cloth. The punch card got reused on computers by default. We're still standing at the same moment again. AI could finally meet us in the middle of a thought, got handed a protocol of a loom. That's what I mean. The prompt is still a punch card, not because of how you encode it. The encoding is powerful and awesome. Because when the LLM is allowed to engage only after you've packaged a complete turn and submitted it. And this is where the mismatch bites. Model capacity is shooting straight up. Reasoning, speech, vision, memory, planning, all curving upwards. The interface protocol, flat. Still a box. Still a submit button. Still the human doing all the work around the LLM. The human still decides what context matters. Still remembers what to ask. Still chooses the timing. Still notices the ambiguity. Still repairs the output. Still has to carefully engineer a prompt. But the intelligence feels magical. It's the interface that still feels like work. And when it feels like work, when the output's wrong, when the magic words don't land, people blame themselves. They decide they're bad at this. They're not specific enough. They don't get AI. I want to say as clearly as I can, it is not our fault. We are not bad at using AI. We are being asked to operate a brand new kind of intelligence through a protocol of a punch card. The mismatch isn't the user, it's the interface. In the race to enable AI, we shortcut the interface. Okay, let's make this concrete and familiar. A few weeks ago, my co-founder was using a frontier company's voice mode. These are known as speech-to-speech models. He asked it a normal question. When is the next Timberwolves game? Fine. It answered it quickly. Then, he pretended I showed up as if to speak to me and said, Hey Ted, come on in. He wasn't talking to the AI, but these models have no way to know that. So the AI did the only thing a prompt box can do. It took his speech as a turn and answered it. Sure, I'm here. What's on your mind? That's not a good answer, but it's not a dumb model. It answered the first question perfectly, but it's a protocol with exactly one slot. Your message, then its reply. It has no concept of who's speaking, whether the words were even meant for it. And the frontier companies want to make strides as well. OpenAI released GPT real-time too in late May and started trying it for their voice mode as well. It back channels now. It goes, mm-hmm, and right. The little sounds we make to show we're listening actively. We're seeing the field is converging on the same conclusion we built our company on. The interface has to stop being batch and start participating. Others are working on real-time conversation as well. This is NVIDIA's Personaplex, a research model, not ours. Watch what happens when it gets interrupted. I've been thinking about starting a diet. Yeah, starting a diet can feel a bit daunting, but you could keep it simple. Focus on eating more veggies and fruits. Before I forget, I signed up for a marathon. All right, congrats on signing up for the marathon. That's a big challenge. You've got a lot of time. Focus on building a solid base with regular long runs. Stay hydrated. Make sure you fuel right before and after. And don't forget to stretch and take care of your feet. Personaplex stops. It yields. It picks the thread back up. That's real turn-taking. Listening and speaking at once. In real time. You need to come visit me. Oh, okay. Because you can go into the city. Okay. Because that's the thing. There's the random spray paint, but then there's also... I'm not sure... People must commission them. Yeah, I think they do. These massive mural spray paint pieces. Yeah, I think they do. And Personaplex's back channels listens and lands where a person would. Beyond conversational flow, there are lots of challenges, and it's a complex problem. Making listening noises is not really the same as knowing who's in the room. These are not trained to tell that, hey, Ted wasn't meant for it. But we are. We're working on improving the protocol to the models. By giving it a better understanding of human and group conversation. Good afternoon, everyone. Good afternoon, Sam. Hi, Jordan. Good afternoon. Good to see you both. Quick one. Which requirement is this? Do we have an ID? This is REQ 142. Expense Approvals. There. It answered a question. It only takes actions based on the utility-driven model. So it creates goals to fulfill as it labels each of the participants' statements as a question, a proposal, an answer, and then only takes a turn when no one else is speaking or holding the floor. Right. We need users to approve requests faster. Yeah. The approval flows too slow. What kind of requests, though? Expense Approvals first. Access Requests eventually. AI. Hold that. Actually, let's pause. Expense Approvals. Expense Approvals. Or a General Approval workflow. Expense Approvals first release. Access Requests are future scope. That changes the data model. Good to know. AI. Pull that up for everyone. Tracking and determining who the speaker is referring to is critical. In this case, it was easy with a direct reference to the AI. But it will happen again without a direct reference. I'd forgotten that was a rule. So over 5,000 needs a second approver? Yep. Manager plus finance. So a big one can't be a single tap. Right. Over the limit, it routes to a second approver. Agreed. Under 5, one tap's fine. Works for me. Okay. Agreed. Expense Approvals. 5,000 Threshold. Right there. The AI resolved the scope objective. No one wrote the prompt. No one packaged the turn and hit submit. The system was in the conversation, following it, understanding, and choosing its moment. AI. Capture that for us. First release supports expense approvals only. Access requests are out of scope. Managers can approve or reject an expense right from a notification. And per the finance controls policy, anything over $5,000 routes to a second approver of the time. Actually, make the threshold 10,000, not 5. Want me to update the requirement to a 10,000 threshold? Yes. AI, is this room free after the meeting? Let me check. The room looks free after this, but... Yes. It's yours until 3 o'clock. And that's the difference between a smart machine behind the same old prompt and an interface that finally participates. Here's the mindset shift I want to leave you with. AI is not just an intelligence technology. It's increasingly becoming an interface technology. And if so, then book smart models alone are not enough. Stop picturing AI as a smarter machine hiding behind prompts, agents, loops, and all the old paradigms. We have to start seeing intelligence itself as a thing that can finally remove interface constraints and amplify human potential. For 75 years, humans adapted to the machine. Its syntax, its forms, its timing, its batch. A system that can reason, listen, infer, adapt, should be able to meet us part way, if not all the way. Instead, if AI is for users, then we should obsess about maximizing the interface. So then the design question changes. What burden are we still putting on humans only because the machine used to be too limited to carry that burden itself? Ask that question and the whole interface space opens up. The answer isn't always chat. It isn't always voice and not a wall of markdown. It's definitely not a decade-old set of digital constructs. The right answer is the affordance humans already use with each other. Communication, a question, a pause, a sketch, a checklist, a quiet aside, or saying nothing at all. An interface where timing and modality aren't the human's job anymore, where choosing the right channel at the right moment is done by the AI. And as a usability person, this is the part that excites me the most. When you take that burden off people, the friction disappears and adoption follows. Computing has mostly been about to date improving how humans encode their intent for machines. The punch card, type a command, click a menu, use your thumb on an iPhone, write a prompt. Every step was progress and every step carried the old constraint forward into the next era. A translation tax, a precision tax, context tax, repair tax. AI is our chance to put those down. Not by making everything magical, not by making everything voice, not by replacing human judgment, but by making computers for once more fluent with us. Most talks and videos cover how to use or adopt AI. The deeper question is how AI intelligence changes the interface. Human conversation is the most human thing we do because if a machine can finally understand more of what we mean, than we can and should stop reshaping ourselves to be understood by it. Thank you.