Field Guide to Fable — Thariq Shihipar, Anthropic
Description
Ask a chat model which Pokemon names end in aw and it fails, even though it knows every Pokemon by heart. Ask Claude Code and it writes a script, fetches the list, and filters for the answer in seconds. Thariq Shihipar, who works on Claude Code at Anthropic, calls that gap capability overhang: models get smarter in spiky ways, and the tools you give them decide which spikes you can reach. Shihipar covers what it takes to work with Fable, Anthropic's newest model. Claude Code cut 80 percent of its system prompt, since heavy instructions now constrain a model more imaginative than the examples it's given. The ask user question tool went from barely working under Opus 4 to generating embedded HTML questionnaires under Fable. He built a full keynote deck in four hours with it, and argues teams should stop picking two of good, fast, and cheap and start demanding all three. Speaker info: - https://x.com/trq212/ Timestamps:
Summary
Generated by claude-sonnet-4-5At-a-Glance
- Verdict: Watch fully
- Core thesis: Fable (Anthropic's new agentic coding model) represents a capability jump requiring operators to 'unhobble' both the model and themselves by finding unknowns, embracing spiky capability overhang, and rejecting traditional tradeoffs to do unreasonably ambitious work.
- Why it matters: First practitioner-level field guide from Anthropic insiders on how to actually work with frontier agentic models; reveals specific prompting evolutions, capability discovery patterns, and operational mindset shifts required for 2025 AI-native workflows.
- Best use: Watch for concrete Fable/Claude Code prompting techniques (blind spot passes, interview loops, reference-based maps), system prompt evolution patterns, and the philosophical frame on 'being unreasonable' about productivity tradeoffs—then apply these patterns to your own agent workflows immediately.
Executive Summary
Thariq Shihipar, a member of Anthropic's Claude Code team, delivered a practitioner's guide to working with Fable—Anthropic's new agentic coding model being rolled out at the conference. His central metaphor: Fable is like an RPG tutorial ending and the open world beginning—enormous possibility but also confusion about how to navigate it. He structures the talk around four concepts: unhobbling Claude (understanding how models get spiky-smart through capability overhang, not linear improvements), finding your unknowns (using structured techniques to map the territory vs. your mental map), dealing with the grief (acknowledging what's lost when hand-coding becomes obsolete while embracing what's gained), and being unreasonable (rejecting implicit tradeoffs and forcing reality to show you constraints rather than pre-compromising).
On unhobbling: Shihipar explains that models are 'grown not designed'—they don't linearly improve but gain capabilities in spiky ways (e.g., Claude Code couldn't list Pokémon ending in 'AW' via reasoning but could via code execution). He documents three major capability jumps: chat models gained context → Claude Code gained 'arms' (bash/tools to build its own context) → Claude Tag gained proactive/multiplayer capabilities. Critically, he reveals Anthropic removed 80% of Claude Code's system prompt recently because newer models are more imaginative than constraining examples, now preferring 'context not constraints.' The ask_user_question tool evolved from barely callable in Opus 4 to interviewing users with 40 questions in 4.5 to embedding questions in full HTML reports in Fable. Markdown/HTML output similarly evolved from simple formatting → planning visualization → rich interactive reports.
On finding unknowns: Shihipar introduces a 2x2 matrix (known knowns, known unknowns, unknown knowns, unknown unknowns) and argues Fable is 'bottlenecked by your ability to match map and territory.' He provides six concrete techniques: (1) blind spot passes—ask Claude to audit your prompt for gotchas in unfamiliar domains, pointing it to git diffs/Slack for context; (2) brainstorms/prototypes—'I have no visual taste, make 4 wildly different design versions so I can react'; (3) interviews—let Claude question you, prioritizing architecture-changing questions; (4) references—give Claude existing code/mockups as a 'map' rather than writing specs from scratch; (5) implementation notes—log deviations when Claude hits unknowns; (6) quizzes—have Claude test your understanding before merging. These techniques help operators stay 'in the loop' despite Fable's expanded autonomy.
On grief and being unreasonable: Shihipar shares personal reflection from revisiting his YC startup codebase—tasks that took weeks now take hours, inducing both joy and mourning for the craft of hand-coding. His resolution: 'the only way out is through.' He advocates Anthropic's internal culture of rejecting tradeoffs ('tradeoffs are not real')—rather than reasonably prioritizing, attempt everything and force reality to reveal actual constraints. The new math is 'good, fast, cheap—pick three.' He frames AI Engineer attendees as needing to prove agents work by 'doing the best work of our lives faster than ever before' while actually working less. He created his entire conference deck in 4 hours with Fable. Final caveat: building is easier but generating value is still hard—requires many swings to find what matters.
Key Takeaways
- Claim: Anthropic removed 80% of Claude Code's system prompt between Sonnet 3.5 and recent models because newer models are more imaginative than the constraining examples they were given. | Evidence: Original best practice was small system prompt + few tools + many examples. Then larger prompts + many tools + many examples. Now smaller prompts with context (not constraints), fewer examples, because 'examples tend to constrain it' and models surpass the example quality. | Caveat: This evolution is model-specific and empirical—'closer to biology than physics'—so what works for Fable may not apply to other model families or earlier Claude versions. | Implication: Ken should audit his agent system prompts: if using extensive few-shot examples or 'do not' constraints with frontier models, he's likely limiting performance. Shift to contextual framing and let the model be more creative than your examples. | Timestamp: 04:45
- Claim: Claude Code can now answer 'which Pokémon end in AW' (Croconaw, Dreadnaw) via code execution despite base models failing pure reasoning, illustrating 'capability overhang'—models get smarter in spiky ways when given tools. | Evidence: Viral tweet showed chat models couldn't list the two Pokémon ending in AW out of 1000 despite 'knowing' all names. Claude Code fetches the list and filters programmatically, succeeding immediately. | Caveat: This only works when the tool (code execution) is available and the model is prompted to use it; capability overhang requires discovering which tools unlock which latent abilities. | Implication: Ken should instrument his agents with diverse tools (bash, search, API calls) and test for surprising capability unlocks rather than assuming linear intelligence increases from model upgrades alone. The harness matters as much as the weights. | Timestamp: 03:20
- Claim: Fable is 'bottlenecked by your ability to match the map (your mental spec/prompt) and the territory (actual codebase constraints)' when it hits unknowns not specified in your prompt. | Evidence: Shihipar introduces a 2x2 matrix: known knowns (what you prompt), known unknowns (acknowledged gaps), unknown knowns (unspoken 'I know it when I see it' design taste), unknown unknowns (unconsidered factors). Fable traverses more territory, hitting more unknowns. | Caveat: No explicit failure mode described—unclear what happens when Fable hits an unknown without guidance (does it guess poorly, stop and ask, hallucinate a solution?). | Implication: Ken should front-load discovery work using Fable itself via blind spot passes, interviews, and brainstorming before full execution. Treat prompting as iterative territory mapping, not one-shot instruction. Budget time for unknown discovery in agent workflows. | Timestamp: 09:15
- Claim: The ask_user_question tool in Claude Code evolved from barely callable in Opus 4 → interviewing users with 40 questions in 4.5 → embedding questions in full HTML reports in Fable, showing spiky capability progression. | Evidence: Shihipar worked on the feature and had to 'tweak the tool' heavily for Opus 4 to call it at all. By 4.5, prompting 'ask me 40 questions about the spec' worked reliably. In Fable, it generates structured HTML with embedded questions. | Caveat: This is a single feature's evolution; not all capabilities follow this trajectory. Shihipar doesn't specify whether this required tool schema changes or just model improvements. | Implication: Ken should revisit tools/features that were marginal in earlier models—they may now be core capabilities. Don't assume a tool that 'didn't work well' six months ago is still limited. Re-test systematically with each model generation. | Timestamp: 06:30
- Claim: Six concrete techniques for finding unknowns with Fable: blind spot passes, brainstorms/prototypes, interviews, references, implementation notes, and quizzes. | Evidence: Blind spot pass example: 'I know nothing about this auth provider, do a blind spot pass to find my relevant unknowns and search my git diff/Slack for gotchas.' Brainstorm: 'I have no visual taste, make 4 wildly different dashboard designs so I can react.' Interview: 'Ask me questions prioritizing ones that would change the architecture.' References: 'Here's HTML mockup code, use it as a map for the React component.' Implementation notes: 'Log deviations when you hit unknowns.' Quiz: 'Test my understanding before I merge this.' | Caveat: These are templates, not rigid scripts; effectiveness depends on giving Claude 'context about you, the work, and the stage you're at.' No quantitative data on time savings or error reduction provided. | Implication: Ken should codify these six patterns as reusable agent prompts/workflows for discovery phases. Especially valuable when entering unfamiliar domains or when delegation scope is large. Treat these as 'pre-flight checks' before agentic execution. | Timestamp: 11:00
- Claim: Anthropic's internal culture is 'tradeoffs are not real'—rather than reasonably prioritizing among competing goals, attempt everything and force reality to show you actual constraints. | Evidence: Shihipar contrasts his YC startup habit of 'writing down priorities and reasonably choosing this quarter's focus' with Anthropic's approach: 'what if you just did all of it?' He created his entire conference deck in 4 hours with Fable and frames the new math as 'good, fast, cheap—pick three' instead of two. | Caveat: Shihipar acknowledges 'building is easier but generating value is still hard'—this isn't about making everything easy, just removing false constraints so you can take more swings at value generation. | Implication: Ken should audit his own implicit tradeoffs: what have you pre-ruled-out because 'it would take too long' or 'we can't do both'? Use agents to attempt the theoretically impossible and let reality (not reasonable planning) show you the real limits. This is a strategic posture shift, not just tactical prompting. | Timestamp: 18:45
- Claim: First-time use of Mythos-class models (Fable) induces both gain and grief: tasks that took weeks now take hours, but the craft of hand-coding and 'rotating the codebase in your mind' is lost. | Evidence: Shihipar revisited his YC startup codebase and realized weeks-long tasks are now hours-long. He loved 'seeing the code base in my mind and rotating it' but also 'remembers swimming in failure' and that 'most projects I've worked on have failed.' Concludes: 'I cannot go back.' | Caveat: This is personal reflection, not empirical data. Some engineers may not experience grief or may find new forms of craft in agent orchestration. No discussion of when hand-coding is still preferable. | Implication: Ken should expect psychological adjustment among himself and teams when adopting agentic coding—it's not just learning new tools but mourning old skills. Frame the transition as 'the only way out is through' (Shihipar's words) to acknowledge discomfort while committing to the shift. | Timestamp: 16:00
Detailed Brief
Unhobbling Claude: Understanding Capability Overhang and Model Evolution
- Claims: Models are 'grown not designed'—capabilities emerge organically from data/feedback/compute, not intentional feature design.; What constrains models is often the harness (prompts, tools, system design) imposed by human understanding, not the model itself.; Claude gets smarter in 'spiky ways' (capability overhang), not linearly—e.g., code execution tool suddenly enables complex reasoning tasks pure chat couldn't do.; Three major capability jumps: chat (context pasting) → Claude Code (bash/tools to build own context) → Claude Tag (proactive/multiplayer agents).; System prompt evolution: small prompt + few tools + many examples → large prompt + many tools + many examples → small prompt + few tools + minimal examples (because newer models are more imaginative than examples).; Ask_user_question tool evolved dramatically: barely callable in Opus 4 → 40-question interviews in 4.5 → full HTML report embedding in Fable.; Markdown/HTML output evolved: basic rich text → planning visualization for users → interactive HTML reports.
- Evidence: Pokémon AW example: 1000 Pokémon, only Croconaw and Dreadnaw end in AW. Chat models fail; Claude Code writes filter script and succeeds.; Anthropic removed 80% of Claude Code's system prompt recently because examples were constraining the model's native creativity.; Shihipar had to heavily tweak ask_user_question for Opus 4; by 4.5 it could handle 40-question interviews without changes; Fable embeds them in HTML.; Reference to Anthropic's 'Biology of a Large Language Model' paper as framework for understanding organic capability emergence.
- Caveats: Model behavior is empirical and closer to biology than physics—rules are discovered, not predetermined.; What works for Fable may not work for other models or even earlier Claude versions.; Capability overhang requires discovering which tools unlock which abilities—it's not automatic.; System prompt best practices are moving targets; no guarantee current patterns persist.
- Implications: Ken should treat model updates as opportunities to remove constraints (examples, system prompt bulk) rather than add them.; Instrument agents with diverse tools and test for non-obvious capability unlocks.; Build intuition empirically—spend time testing what the model can/can't do rather than relying on documented capabilities.; Revisit 'failed' features from earlier models; they may now work well.; Don't assume linear progress; expect discontinuous jumps in specific capability domains.
Finding Your Unknowns: Techniques for Mapping Territory vs. Map
- Claims: Map (your mental spec/prompt) vs. territory (actual constraints in codebase/world) mismatch creates unknowns—decision points you haven't specified.; Fable traverses much larger territory autonomously, so unspecified unknowns become bottleneck.; 2x2 matrix: known knowns (what you prompt), known unknowns (acknowledged gaps), unknown knowns (implicit 'know it when I see it' taste), unknown unknowns (unconsidered factors).; Blind spot passes: Ask Claude to audit unfamiliar domains for gotchas, pointing it to git diffs/Slack/docs for context.; Brainstorms/prototypes: Generate wildly different design options to surface unknown knowns (design preferences you can't articulate).; Interviews: Let Claude question you, prioritizing questions that would change architecture or key decisions.; References: Provide existing code/mockups as maps rather than writing specs from scratch—'give Claude another map.'; Implementation notes: Log deviations when Claude hits unknowns during execution so you can review and understand them.; Quizzes: Have Claude test your understanding of what happened before merging/shipping.
- Evidence: Blind spot pass example: 'I know nothing about this auth provider, help me figure out my relevant unknowns' + search git diff/Slack.; Brainstorm example: 'I have no visual taste, make me an HTML page with 4 wildly different design decisions so I can react.'; Interview prompt: 'Prioritize questions that would change the architecture.'; Reference example: 'If I'm making a React component, I might have an HTML mockup that is my map.'; Implementation notes: 'While running Fable, if it runs into an unknown, ask it to log it.'; Quiz: 'Get Fable to quiz me about what happened just to make sure I understand what I'm doing and can represent this work when creating a PR.'
- Caveats: These are templates requiring contextualization—'giving it context about you, the work, and the stage you're at' is critical.; No quantitative data on time savings, error rates, or success metrics provided.; Unclear what happens when Fable hits an unknown without guidance (does it stop, guess, hallucinate?).; These techniques assume Claude is good at meta-cognition (knowing what it doesn't know)—may not always surface true unknowns.
- Implications: Ken should codify these six patterns as reusable prompts/workflows for agent discovery phases.; Treat large delegations as multi-stage: discovery (blind spots, interviews, brainstorms) → execution → review (notes, quizzes).; Build a library of domain-specific 'blind spot' prompts for common areas (auth, payments, data pipelines, etc.).; Use references aggressively—show don't tell—to reduce specification burden and ambiguity.; Instrument agents to log deviations automatically for post-hoc review rather than relying on memory.
Dealing with Grief and Being Unreasonable: Mindset Shifts for Agentic Era
- Claims: First-time use of Mythos-class models induces grief: what took weeks now takes hours, but the craft of hand-coding is lost.; Pre-LLM coding 'feels like a foreign country'—constant forced tradeoffs because coding was hard.; Shihipar loved 'rotating the codebase in my mind' but also 'remembers swimming in failure'—most projects/startups fail.; Conclusion: 'I cannot go back' despite affection for the craft; 'the only way out is through.'; Anthropic's internal culture: 'tradeoffs are not real'—reject reasonable prioritization and attempt everything, forcing reality to show actual constraints.; The new math: 'good, fast, cheap—pick three' (instead of pick two).; Goal is to 'do the best work of our lives faster than ever before' while actually working less and spending more time with people.; Building is easier, but generating value is still hard—takes many swings to find what matters.; AI Engineer attendees are expected to prove agents work, not just talk about them.
- Evidence: Shihipar revisited his YC startup codebase (30 people, constant forced tradeoffs) and realized week-long tasks are now hour-long.; Made his entire conference deck in 4 hours with Fable the night before the talk.; At his previous company, he'd 'write down priorities and reasonably choose this quarter's focus'—now asks 'what if you just did all of it?'; Personal resolution: 'Be more productive but work less and spend more time with people I really care about.'
- Caveats: This is personal reflection and cultural observation, not empirical productivity data.; Some engineers may not experience grief or may find new forms of craft in agent orchestration.; No discussion of when hand-coding is still preferable or what skills remain valuable.; Generating value is explicitly called out as still hard—agents don't automatically solve product-market fit or strategic prioritization.; The 'prove it to the world' framing creates pressure but doesn't provide a roadmap for what proof looks like.
- Implications: Ken should expect and normalize psychological adjustment when adopting agentic workflows—it's not just tool learning but identity/craft mourning.; Frame the transition internally as 'the only way out is through' to acknowledge discomfort while committing forward.; Audit implicit tradeoffs: what have you pre-ruled-out as 'impossible' or 'would take too long'? Test those assumptions with agents.; Use agents to increase swing volume (more experiments, more prototypes, more bets) rather than just doing the same work faster.; Balance building ease with value discipline—just because you can build it doesn't mean you should; focus on value generation.; Set personal/team boundaries to avoid agent-enabled overwork—the goal is to work less, not just do more.
Notable Concepts & Terms
- Fable: Anthropic's new agentic coding model (Mythos-class) being rolled out at AI Engineer conference; represents a major capability jump requiring new operational patterns. Described as 'the map opening up' after the tutorial—enormous possibility but also confusion.
- Capability Overhang: Anthropic's term for latent model abilities that emerge when the right tools/harness are provided, not from model improvements alone. Models get smarter in spiky, non-linear ways (e.g., code execution suddenly unlocking complex reasoning).
- Unhobbling: Removing constraints (system prompt bulk, limiting examples, missing tools) that prevent models from using their full capabilities. Both the model and the operator (you) must be unhobbled to work effectively with Fable.
- Map vs. Territory: Distinction between your mental spec/prompt (the map) and the actual codebase/world constraints (the territory). Unknowns occur when Claude navigates territory not covered by your map.
- Unknowns Matrix: 2x2 framework: known knowns (what you prompt), known unknowns (acknowledged gaps), unknown knowns (implicit taste/judgment), unknown unknowns (unconsidered factors). Fable is bottlenecked by your ability to find and specify these.
- Blind Spot Pass: Technique where you ask Claude to audit an unfamiliar domain for gotchas/unknowns before execution, pointing it to git diffs, Slack, or docs for context. Used to surface known unknowns and unknown unknowns.
- Tradeoffs Are Not Real: Anthropic internal culture of rejecting reasonable prioritization and attempting everything simultaneously, forcing reality (not planning assumptions) to reveal actual constraints. Enables 'unreasonable' ambition.
- Context Not Constraints: Evolved system prompt philosophy for newer models: provide contextual framing rather than restrictive 'do not' rules or limiting examples. Models are more imaginative than constraining examples suggest.
- Claude Code: Anthropic's agentic coding product; gained 'arms' (bash tool, environment manipulation) to build its own context rather than relying on pasted code. Evolved from chat → proactive (Claude Tag) capabilities.
- Ask User Question Tool: Claude Code feature allowing the model to show multiple-choice dialogues or interview users. Evolved dramatically across model generations: barely usable in Opus 4 → 40-question interviews in 4.5 → HTML-embedded questions in Fable.
Operator Notes / Why Ken Should Care
- This is a rare insider view from Anthropic's Claude Code team on how to actually operate frontier agentic models—not just what they can do but how prompting/system design must evolve to unlock them.
- The 'unhobbling' frame is critical: Ken should audit his agent systems for over-constraining prompts (excessive examples, 'do not' rules) that may limit newer models. Consider removing rather than adding to system prompts.
- The six unknown-finding techniques (blind spot passes, interviews, brainstorms, references, implementation notes, quizzes) are immediately actionable workflows Ken can template into agent systems for discovery phases before large delegations.
- The system prompt evolution pattern (many examples → fewer examples because models surpass them) suggests Ken should version-control prompts and systematically A/B test prompt minimization with each model update.
- Fable's longer autonomous traversal means unknowns become bottleneck—Ken should invest in instrumentation (logging deviations, surfacing decision points) and multi-stage workflows (discovery → execution → review) rather than one-shot prompting.
- The 'tradeoffs are not real' culture is a strategic lens: Ken should use agents to test impossible-seeming combinations (speed + quality + breadth) and let reality show constraints, not pre-compromise based on pre-agent assumptions.
- The grief observation is operationally important: expect team adjustment when rolling out agentic coding. Frame it as 'the only way out is through' and focus on working less (not just doing more) to avoid burnout.
- Building ease ≠ value ease: Ken must maintain discipline around product strategy and value generation even as execution becomes trivial. Use agents to increase experiment velocity, not just build backlog faster.
- The talk itself was created in 4 hours with Fable (Shihipar's claim)—strong signal that content/deck generation is a high-value Fable use case Ken should test immediately.
- Watch for the 12:30 fireside chat mentioned (with Cat Woo and Simon Wilson) for additional Fable updates and operational details not covered in this talk.
Watch Map
- 00:00: Introduction, selfie tradition, Fable rollout announcement
- 01:30: Fable as 'the map opening up' metaphor; four-part structure overview
- 02:15: Part 1: Unhobbling Claude - models are grown not designed
- 03:20: Pokémon AW example: capability overhang via code execution tool
- 04:45: System prompt evolution: 80% removal, from examples to context
- 06:30: Ask_user_question tool progression across model generations
- 07:45: Markdown/HTML evolution; biology vs. physics framing
- 08:30: Part 2: Finding Your Unknowns - map vs. territory distinction
- 09:15: 2x2 unknowns matrix introduction (known/unknown knowns/unknowns)
- 10:00: Six techniques: blind spot passes, brainstorms, interviews
- 11:00: References, implementation notes, quizzes - staying in the loop
- 13:00: Part 3: Dealing with the Grief - pre-LLM coding as foreign country
- 16:00: YC startup revisit, weeks→hours, grief and gain, cannot go back
- 17:30: Part 4: Being Unreasonable - tradeoffs are not real culture
- 18:45: Pick three not two; do all of it, force reality to show constraints
- 20:00: Made deck in 4 hours; prove agents work; building vs. value generation
- 21:30: Closing: go explore, make it real, be less reasonable
Source/Metadata
- Title: Field Guide to Fable — Thariq Shihipar, Anthropic
- Transcript words: 4766
- Duration seconds: 1168
- Timestamp note: Timestamps derived from video duration (1168 seconds) and transcript structure; precise chapter markers not present in transcript but sections are clearly delineated by speaker transitions and topic shifts.
Transcript
Music Please welcome to the stage, member of technical staff at Anthropic, Tharik Shihapar. Music Music [SPEAKER_01] Hey, everyone. I'm Tharik. I work at Anthropic on Cloud Code. [SPEAKER_01] Before we get started, we have a tradition on Cloud Code where we take a selfie before a talk. So, if you don't mind, if you strike a pose with me, I'll take a quick selfie at AI Engineer. Music Okay. Incredible. Well, yeah, to kick things off, Fable is back. We're rolling it out later today. Stay tuned for exact timeline. Me and Cat Woo and Simon Wilson will be doing a fireside chat at 12:30. We might have some updates for you then. But Fable is a model I'm just so, so excited about. It's one of those Anthropic models where you're just going to remember it. Sonnet 3.5 New, Opus 4, Opus 4.5. It's a model that I just have a lot of affection and excitement for. And the best way to describe Fable to me is the map is opening up. You know, you are playing an RPG and you've been on the tutorial. And now you get to the point where the open world starts, right? And there's so much that you can do and explore. But there's also a little bit intimidating and confusing, right? Because there's so much you can do. And so what I wanted to do in this talk is give you guys a field guide to Fable, right? How do you work with this new class of models? So I've got four parts to it. I've been working on this as a series of articles and blog posts. But when we announced Fable was coming out, I was like, okay, let me do all of this at once at the talk, speedrun. So there are four parts. Unhobbling Claude, finding your unknowns, dealing with the grief, and being unreasonable. So first, unhobbling Claude. But I think something we say really often is that the models are grown, not designed, right? We don't wake up and be like, we need 99% on Sweebench, right? The models are something we grow carefully. We give it data and feedback and compute. But ultimately, it's something that we figure out and learn with the model as we use it. And so, what that also means is that what contains them is us, right? The harness we put them in and the way we prompt them is basically a function of our understanding of Claude, right? And by unhobbling it, I mean how can we understand Claude better to unleash it? And we need to understand Fable more. So I think one of my points is that we're still so early. And I think there's a lot more understanding in Fable to unlock. And I think I'll give you a quick example about how models get smarter, because it's a little bit unintuitive, right? I saw this viral tweet a couple weeks ago being, why can't LLMs say which Pokemon end in AW? There are a thousand Pokemon, right? And it turns out there are two whose names end in AW, Croconaw and Dreadnought, right? And it turns out if you ask a normal chat model, it can't answer it. Which is kind of confusing, because, you know, it definitely knows all the names of the Pokemon, right? But if you ask Claude Code, it can, right? Because what it does is fetch every Pokemon and write a script to filter for AW, right? And so this is what I mean by unhobbling Claude. We call this capability overhang, right? Claude gets smarter in spiky ways. So it doesn't just remember every Pokemon and reason through it. But if you give it the code execution tool, it can find the two Pokemon's enemies. And then it ends with AW, right? And so this is part of the challenge with Fable: figuring out this capability overhang. What is now possible? And I think this is a discovery that I'm excited to go on with you. To make this a little bit clearer, I'm going to talk about a few different examples of how models have progressed in the past. One of the big examples, obviously, is chat. The chat models had to be given context, right? Maybe you paste in your code base. And maybe naively you might have thought the way we solve coding is by the context just getting really large. And I can just paste in my entire code base. It will be a hundred million context window. But it turns out that instead, if you give it arms, give it the bash tool and ways to work with the environment, it can build and search its own context. And that's what led to Claude Code, right? And so, again, spiky, a new innovation in how we think about and work with the model. And then recently we've rolled out Claude Tag. And what unlocked Claude Tag is its ability to work proactively and multiplayer. Claude Code, you know, is something that you have to prompt for it to do work, right? And this ability for Claude to wake itself up and do work is something that we think is unlocking the new wave of agents. But there's more here. So, for example, we recently removed 80% of the system prompt for Claude Code, right? And this is one of the ways in which models and what they need changes over time. So, originally, back in Sonnet 3.5 New, the best practices for a system prompt was a small system prompt, few tools, and lots of examples, right? And then as the models get smarter, you can give them more information and more instructions and they start following them. And so it's a larger system prompt with lots of examples and many tools, right? But most recently we found this new class of models wants fewer, wants a smaller system prompt. The examples tend to constrain it because it's actually more imaginative than the examples we give it. And so we try to give it context and not just constraints. We really try and avoid being like, do not do this, which is really necessary for the previous models. And so this is a way that the system prompt is changing and probably will continue to change. Another feature I really like is the ask user question tool. This is something I worked on when I first got to Cloud Code. And so it's a larger system prompt with lots of examples and many tools, right? But most recently we found this new class of models want fewer, want a smaller system prompt. The examples tend to constrain it because it's actually more imaginative than the examples we give it. And so we try to give it context and not just constraints. We really try and avoid being like, do not do this, which is really necessary for the previous models. And so this is a way that the system prompt is changing and probably will continue to change. Another feature I really like is the ask user question tool. This is something I worked on when I first got to Cloud Code. And it's when Cloud is planning or wants to ask you a question, it can show you a multiple-choice dialogue. For Opus 4, it could barely call it. I had to really tweak the tool to make sure that it would work, right? And then sometime at Opus 4.5, I was thinking, what if I asked it to ask me 40 questions about the spec? It could start interviewing me, right? And so its ability to ask questions jumped, right? And then most recently with Opus 4.8 and Fable, I can now build a whole HTML report with the questions embedded inside of them. And it's a whole new way of interacting with Cloud, right? And so this progression of how Cloud can get information from you has also changed. Speaking of which, Markdown and HTML is something I've also talked a lot about. Initially Markdown was a good output for the model. It could show a little bit of rich information. And then with plan mode, it started to be for you. You could understand what Cloud was about to do. And now Cloud can build you these in-depth HTML reports, right? And so again, a way of the models getting smarter in a spiky way. I really like to emphasize that this is closer to biology than physics, right? It's still very empirical, very organic. We don't know all the rules, but there is some sort of science behind it, right? There is an intuition to build as well. And so I really encourage you to treat Fable like that. One of my favorite papers that Anthropic has written is on the biology of a large language model. All of our research papers are meant to be read by people with various degrees of technical expertise, but this is one of my favorites. So if you're looking to learn a little bit more, I suggest you check it out. But so, yeah, we talked about unhobbling Claude. But it turns out when you're working with Fable, you also need to unhobble yourself, right? And so one of the things that I think a lot about is that the map is not the territory, right? When I'm working on a coding problem, the plan and prompt and spec that I have in my mind is the map, right? But the territory is the actual code base, the real world, the constraints that Claude needs to navigate, right? And whenever Claude runs into something in the territory that's not in the map, I call that an unknown, right? Claude has to figure out what to do about it. It's a decision point that I haven't specified. And Fable is one of the first models where I felt that I really have to figure out my unknowns, because if not, it's going to traverse such a large area that it's going to run into a lot of them. So how do you figure out your unknowns? Fable is bottlenecked by my ability to match the map and the territory to find my unknowns. So a few ways to think about this. I like to think of it in a matrix. So for any problem, I have a bunch of known knowns. This is usually what I write in my prompt. What do I want, right? Then I have known unknowns. Things that I know I don't really know yet, but I just haven't figured out yet. Then I've got unknown knowns. What's so obvious that I just wouldn't write it down, you know? But I know it when I see it, right? And then finally, unknowns, unknowns. What haven't I considered at all? What do I not know, right? What is something that, if I knew, could change how I prompt Claude? And luckily, you can use Claude, you can use Fable to find your unknowns. So I'm going to go over a few examples of how I do that with Fable. The first is I like to do what I call a blind spot pass. So I like to say something, like, hey, I'm working on a new auth provider that I know nothing about. In this code base, can you do a blind spot pass to help me figure out my relevant unknowns, known unknowns and help me prompt better, right? And so this might have Claude go through the auth module and figure out, oh, you know, this is a hairy little dead end that comes up a lot. Maybe search as my Git diff or Slack. I might tell it where there's context, right? So that I can learn about all the gotchas. And you can use this very broadly, right? You can use it to teach you about new fields. I recently did this for color grading when doing video editing. But I think it's really powerful and Fable is incredible at it. In many ways the model knows more about almost everything than I do. I just need to get it out of it. Then I like to use brainstorms and prototypes. This helps me figure out my unknown knowns, right? Things, especially for design, for me it's know it when you see it. Right? So I might ask it to create a dashboard. And I tell it I have no visual taste. Make me an HTML page with four wildly different design decisions so I can react to them. Right? And then you tweak this as you want, but the idea is to sort of get an idea of what are the things that you can't describe in words. Right? And work with the model to help figure that out. Things like, especially for design, for me it's know it when you see it. Right? So I might ask it to create a dashboard, and I tell it I have no visual taste. Make me an HTML page with four wildly different design decisions so I can react to them. Right? And then you tweak this as you want, but the idea is to get an idea of what are the things that you can't describe in words. Right? And work with the model to help figure that out. Then interviews. So once I have an idea of what I want to do, there's probably still a lot of unknowns here, right? Where I might not have considered something, I might not have specified it, and so I'll ask Claude to interview me. Right? And I'll give it a little bit more context. In any of these questions, giving it a little bit more context about you and the work and the stage you're at, hey, prioritize questions that would change the architecture, is extremely helpful. Then references. One of the best ways to give Claude a map is to give it another map. Right? So instead of me writing out the spec, I can just say, hey, here is some code that represents what I want to be done. Right? It could be in a different system or language. But just read this code, understand it, and then use that to start your work. Right? And again, this can be in a lot of different ways. If I'm making a React component, I might have an HTML mockup that is my map, right, that I pass in as a reference. I think this is really powerful, and Fable is really incredible at it. Something else I'd really appreciated is implementation notes. So if while you're running Fable, and it runs into an unknown, ask it to log it. Right? So that you can see where the deviations happened, and then you can figure out why as well. You know, it will usually give you some context about what happened. And then finally, I like to get Fable to quiz me about what happened, just to make sure I understand what I'm doing, and I can represent this work when I'm creating a PR or merging it. This is a really great way of making sure that you're really in the loop with Fable. And I think that's one of the most important parts of Fable, is staying in the loop and making sure that you get what you want. So those are some of my tips for working with Fable. I also want to say that the first time I used a Mythos class model, used Fable, I felt both a huge sense of gain, but also a sense of loss. And I wanted to talk a little bit about that. When I think about coding before LLMs, it feels like a foreign country. I used to run a YC startup, about 30 people, and we were constantly forced into tradeoffs because of how hard code was, right? Like, we could make the app fast, or we could try prototyping a new feature, and this might take a month, or this would take two months, and so we had to choose. It was really hard. And now I went back to that code base a couple weeks ago, and I thought about some of the things that I wanted to do, and it was way easier. It was the things that would have taken me weeks, I could do in hours. And at some point it's yeah, how can you not laugh, but also how can you not cry, honestly? It's one of these things where I really loved programming and writing code by hand. I loved the feeling of seeing the code base in my mind, and rotating it. But I also remember staying up late nights trying to debug, working on things for weeks without working, right? I just remember swimming in failure. I just remember that most of the projects I've ever worked on have failed. Most startups go bankrupt. I think just overall programming and coding is extremely hard, and as much as I enjoy those highs, I cannot go back, right? And my reflection here is the only way out is through, right? There's still a lot to learn with agentic coding. There's a lot to learn with Fable. But I think if we try really hard, and if we stay in the loop, we un-hobble it, we can get there. And we can come out on the other side with so much more. And so the last bit I wanted to talk about is the so much more part. I call this being unreasonable. One of my favorite parts of Anthropic is that we believe that tradeoffs are not real. I think that very often, in my previous company, I was very used to being reasonable. So I'd write down this list of priorities, and I'd be like, well, I guess we can prioritize this against this, right? And you know, that makes sense, so we'll this will be our priority this quarter. But what if you just did all of it? What if you force reality to show you the tradeoff? This is something I've really valued as our culture in Anthropic. In my previous company, I was very used to being reasonable. So I'd write down this list of priorities, and I'd be like, well, I guess we can prioritize this against this, right? And that makes sense, so we'll, this will be our priority this quarter. But what if you just did all of it? What if you force reality to show you the tradeoff, right? This is something I've really valued as our culture in Anthropic, and that my reflection going forward is that I'm gonna be a lot less reasonable. I think one of this, the math of Claude and Fable really changes how you think about tradeoffs, and there are so many tradeoffs that you make implicitly in your head, right? Like, good, fast, cheap, now it's pick three, right? I think that the best way to do more ambitious work is to reframe and make ourselves more ambitious. Because I think the only way to prove that agents work is to do the best work of our lives faster than ever before. You know, for example, I made this deck last night in about four hours with Fable. I feel like it's a deck I really like, and I really enjoyed it. But I also did it really fast. And I think that if you're here at AI Engineer, the world is kind of looking at you to prove that AI works, right? That it's not just a fad or something, but that it can make us more productive and also save us time. And that's my resolution for this year, the goal here is to work, be more productive but work less and spend more time with people I really care about. I think it's also worth calling out that building is easier, but generating value is still hard. And I think this is something that we run into as AI engineers sometimes, where we think so much about the process of building and our setups. But the point is to generate value, right? And there, it takes a lot of swings. It takes a lot of tries to find the valuable stuff. But that really is the goal. And that's what the world is looking to us to prove that AI can really transform it. So, to end, I just wanted to say, go explore, make it real, and, yeah, be less reasonable. Thank you. Thank you. that you, uh, you know, you can't describe in words. Right? And, uh, like work with the model to help figure that out. Uh, then, then interviews. So once I have an idea of like, you know, this is what I want to do, uh, there's probably still a lot of like, uh, unknowns here, right? Where I might not have considered something, I might not have specified it, and so I'll ask Claude to interview me. Right? And I'll give it a little bit more context. In any of these questions, like giving it a little bit more context about you and the work and the stage you're at, like, hey, yeah, prioritize questions that would change the architecture, is extremely helpful. Uh, then references. One of the best ways to give Claude a map is to give it another map. Right? So instead of me writing out the spec, uh, I can just say, hey, here is some code that represents what I want to be done. Right? It could be in a different, uh, system or language. Uh, but just read this code, understand it, and then use that to start your work. Right? And, uh, again, this can be in a lot of different ways. If I'm making a React component, I might have an HTML mockup that is my map, right, that I pass in as a reference. I think this is really, really powerful, and Fable is really incredible at it. Uh, something else I'd like really appreciated is implementation notes. So if, uh, while you're running Fable, uh, and it runs into an unknown, ask it to log it. Right? So that, um, you, uh, you can see where the deviations happened, and then you can sort of figure out why as well. You know, it will usually give you some context about what happened. And then finally, I like to get a Fable to quiz me about what happened, uh, just to make sure I understand what I'm doing, and I can represent this work, you know, when I'm creating a PR or merging it. Um, this is a really great way of, like, making sure that you're, like, really in the loop with Fable. And I think that's, like, one of the most important parts of Fable, is, like, staying in the loop and making sure that you, uh, you get what you want. So, um, those are some of my tips for working with Fable. Uh, I also want to say that the first time I used a Mythos class model, uh, used Fable, I felt both a huge sense of, like, gain, but also a sense of loss. And I wanted to talk a little bit about that, you know? Um, when I think about coding before LLMs, it feels like a foreign country. You know, like, I used to run a YC startup, about 30 people, and we were just constantly forced into tradeoffs because of how hard code was, right? Like, we could make the app fast, or we could try prototyping a new feature, and this might take a month, or this would take two months, and so we had to choose. It was just really, really hard. Um, and now I went back to that code base a couple weeks ago, and I thought about some of the things that I wanted to do, and, uh, it was just way easier. It was, like, the things that would have taken me weeks, I could do in hours, you know? And, uh, at some point it's like, yeah, like, how can you not laugh, but also how can you not cry, honestly? Like, it's, like, one of these things where, um, I really, really loved programming and writing code by hand. I loved the feeling of, like, seeing the code base in my mind, and, like, rotating it. But I also remember just, you know, like, staying up late nights trying to debug, working on things for weeks without working, right? I just remember swimming in failure. I just remember that, like, the most of the projects I've ever worked on have failed. Most startups go bankrupt, you know? I think just overall programming and coding is extremely hard, and, like, as much as I enjoy those highs, I cannot go back, right? And, uh, my reflection here is, like, the only way out is through, right? There's still a lot to learn with agentic coding. There's a lot to learn with Fable. Uh, but I think if we try really hard, and if we, like, stay in the loop, we un-hobble it, uh, we can get there, you know? And we can come out on the other side, uh, with just, um, so much more. And so the last bit I wanted to talk about is the so much more part, right? I call this being unreasonable. Um, one of my favorite parts of Anthropic is that we believe that tradeoffs are not real. Um, like, I think that very often, I, like, in my previous company, I was very used to being reasonable. So I'd, like, write down this list of priorities, and I'd be like, well, I guess we can prioritize this against this, right? Um, and, uh, like, you know, that makes sense, so we'll, we'll, this will be our priority this quarter. But what if you, uh, just did all of it, you know? What if you forest reality to show you the tradeoff, right? Um, this is something I've really valued as our culture in Anthropic, and that my reflection going forward is that I'm gonna be a lot less reasonable. Um, I think one of this, like, the math of Claude and Fable really changes how you think about tradeoffs, and there are so many tradeoffs that you make implicitly in your head, right? Like, good, fast, cheap, now it's pick three, right? Um, I think that, like, the best way to, like, do more ambitious work is to, uh, like, reframe and make, make ourselves more ambitious. Because I think the only way to prove that agents work is to do the best work of our lives faster than ever before. Um, you know, for example, I made this deck last night in about four hours with Fable. I feel like it's a, it's a deck I really like, and I, I really enjoyed it. But I also, um, you know, did it really fast. Uh, and I think that if you're here, you know, at AI Engineer, the world is kind of looking at you to prove that AI works, right? That it's not just, like, a fad or something, but that it can make us more productive and also save us time. And, and that's my resolution for this year, the goal here is to work, uh, be more productive but work less and spend more time with people I really care about. Uh, I think it's also worth calling out that building is easier, but generating value is still hard. And I think this is something that we run into, you know, as AI engineers sometimes, where we think so much about the process of building and our, our setups. Um, but the, the point is to generate value, right? And, uh, there, it takes a lot of swings. It takes a lot of tries to find the valuable stuff. Uh, but that really is the goal. And that's like, you know, again, what the world is looking to us to prove that AI can really transform it. So, to, to end, I just wanted to say, like, go explore, make it real, and, uh, yeah, be less reasonable. Thank you. Thank you.