Every

Anthropic’s Fable 5: A Warp Drive for Coding

3593 summary words 16 min summary Watch video

Start with the signal

16 min read

Summary

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: Anthropic's Claude 5 (marketed as Fable) is a 'warp drive' model that compresses months/years of autonomous coding work into hours/days, scoring 91/100 on senior engineer benchmarks (vs. 62-63 for GPT-4.5/Claude 4.1), but is expensive, slow, and requires level 7-8 AI adoption to be useful—not a daily driver for most users.
  • Why it matters: First frontier model to demonstrate sustained multi-hour autonomous execution with production-grade taste/judgment, raising both the floor for non-experts and ceiling for experts, while revealing AI adoption is now a skill gap issue.
  • Best use: Reference for understanding autonomous agent workflows, pricing/adoption curves, and what 'warp drive vs daily driver' models mean for strategic planning and hiring.

Executive Summary

Dan Schipper (CEO of Every, an AI research subscription) provides a week-long real-world test of Claude 5 (Fable), Anthropic's largest 'mythos-class' model, released with cyber/bio safeguards after being deemed 'too dangerous' a month prior. The model costs $10/M input tokens and $50/M output tokens—double Opus pricing—but scored 91/100 on Every's 'senior engineer benchmark' (rewriting production codebases from first principles in one prompt), versus 63 for Claude 4.1 and 62 for GPT-4.5. Schipper's central metaphor: this is a 'warp drive' for crossing galaxies (big multi-hour tasks), not a vehicle for getting around town (quick iteration).

Every's team of ~7 testers (engineers, writers, marketers) found a bimodal response: users at AI adoption levels 7-8 (orchestrating multiple agents, delegating work 24/7) found it 'crazy' useful; users below level 6 struggled to find use cases because they lack problems big enough to justify the cost/slowness. The model excels at sustained autonomous execution (give it a task, leave for 3-4 hours), taste/attention to detail (drop caps, font weight, UX polish in one-shot outputs), multi-context research (synthesizing thousands of survey responses into falsifiable growth hypotheses), and batch GitHub issue triage. It does not substantially outperform Claude 4.8 at writing (dense, literary prose) and is overkill for daily collaboration.

Schipper built a 3D Borges Library of Babel browser game in one prompt over 3-4 hours, a Hubert Dreyfus lecture mini-site with synced audio playback highlights (no link provided—model sourced lectures autonomously), and pointed Fable at Every's survey data to surface 'you have a conversion merchandising problem' with a falsifiable next step (pricing transparency + trial offer). The model also auto-closed and wrote fixes for Every's Proof app GitHub backlog. No architectural novelty exists—it's 'just bigger and better' than prior Anthropic models, per internal sources.

The piece closes with a paradox from Schipper's 'After Automation' essay: automation creates more human work, not less. Fable raises the floor for vibe coders (one-shot games) and the ceiling for experts (solo AAA games), but using it is a skill that requires adoption maturity. Schipper expects this capability to be cheap/accessible in 6-12 months, making it a strategic preview rather than immediate mass tool. He still uses GPT-5.5 as his daily driver in Codex for fast iteration.

Key Takeaways

  • Claim: Claude 5 scored 91/100 on Every's 'senior engineer benchmark,' matching a human senior engineer in one prompt and saturating a benchmark Schipper expected to last six months. | Evidence: The benchmark tests rewriting a real vibe-coded production codebase from first principles. Prior best: Claude 4.1 (63), GPT-4.5 (62). Hexagon chart shows Fable fills all skill dimensions vs. spiky prior models. | Caveat: The benchmark is Every's internal creation and not public/peer-reviewed; no details on what 'vibe-coded production codebase' means or how 'human senior engineer' baseline was established. | Implication: If repeatable, this suggests coding agent capabilities have leapfrogged incrementally and may obsolete certain junior/mid-level implementation roles faster than expected; also implies benchmark design is now a bottleneck for model differentiation. | Timestamp: 02:40
  • Claim: The model is a 'warp drive': it compresses months/years into hours/days for big autonomous tasks, but is too slow/expensive/token-hungry for daily iteration or collaboration. | Evidence: Built a 3D Library of Babel game in 3-4 hours from one prompt; wrote fixes for Every's Proof app GitHub backlog in batch. Costs $10/M input, $50/M output (2x Opus). Schipper still uses GPT-5.5 for daily driver tasks in Codex. | Caveat: No cost breakdown provided (e.g., how many tokens the 3-4 hour game build consumed, or total cost per project). 'Warp drive' metaphor may understate that most knowledge work is iterative, not batch. | Implication: Strategic use case is delegating backlog/greenfield projects overnight or over weekends, not replacing Copilot/Cursor for live coding. Pricing will gate adoption until it drops, which Schipper expects in 6-12 months. | Timestamp: 04:15
  • Claim: The model has 'taste and attention to detail': font choices (all-caps headers, drop caps, non-default weights), UX patterns (±15s playback controls, highlight sync), and non-slop aesthetics emerged in one-shot outputs. | Evidence: Dreyfus lecture site used drop caps, better-than-default fonts, synced audio/text playback with custom player controls—all from 'can you grab his lectures and make a mini-site' prompt. Library of Babel had hexagonal galleries, 20 shelves per side, accurate to book mathematics. | Caveat: No comparison to other models on same prompts; unclear if 'taste' is prompt-following or actual aesthetic judgment. Schipper expects this will look sloppier once everyone uses it (slop saturation). | Implication: Non-technical users (vibe coders) can now ship polished-looking MVPs without design handoff; experts may use it to prototype UX faster. Also implies prompt engineering for aesthetics may become less necessary. | Timestamp: 07:30
  • Claim: The model synthesized thousands of Every survey responses into 'you have a conversion merchandising problem' with a falsifiable next step (pricing transparency + trial offer) better than weeks of internal team analysis. | Evidence: Fed survey data + analytics + site content; model output: 'Your free-to-paid conversion ratio is lower than it should be. If we ship pricing transparency and a trial offer, I think it's going to go up.' Team had not surfaced this insight despite weeks of AI-assisted analysis. | Caveat: No validation provided on whether the hypothesis is correct or if the model is pattern-matching common SaaS advice. No detail on prompt structure or how much context (token count) was used. | Implication: High-context synthesis for strategic analysis (growth, ops, research) is now a one-prompt task for advanced users. This could replace junior analyst/PM work or become a daily exec ritual. | Timestamp: 09:15
  • Claim: Only users at AI adoption levels 7-8 (orchestrating multiple agents, delegating 24/7 work) found Fable transformative; below level 6, it felt like overkill with no clear use case. | Evidence: Every's team of ~7 (engineers, writers, marketers) split ~50/50 on utility. Article reference: Every's 'eight levels of adoption' framework (level 1 = Google replacement, level 8 = multi-agent orchestration). Vibe coders/non-experts need 'big meaty problems' to justify cost/slowness. | Caveat: No definition of what distinguishes level 6 from 7, or how users self-assess. Adoption framework is Every's own and not externally validated. Schipper predicts this workflow will be 'useful for everyone soon,' but gives no timeline beyond 'eventually.' | Implication: AI adoption is now a skill/maturity gap, not just a tool access issue. Orgs should assess user AI maturity before rolling out expensive frontier models. Also suggests that mass demos (like this video) may overpromise utility for median users. | Timestamp: 13:00
  • Claim: The model did not substantially outperform Claude 4.8 at writing; outputs are dense, literary, and not ideal for copywriting. GPT-5.5 is better for writing tasks. | Evidence: Schipper and team tested writing tasks; Fable's prose is 'big blocks, pretty literary.' Schipper prefers GPT-5.5 for daily writing in Codex. | Caveat: No writing samples or benchmark scores provided. 'Dense and literary' could be prompt-tunable or a reflection of default style. | Implication: Frontier model ≠ best-in-class for all tasks. Users should maintain a multi-model workflow (Fable for autonomous tasks, GPT-5.5 for writing, etc.). Also implies OpenAI still leads on prose/copywriting. | Timestamp: 14:30
  • Claim: Anthropic placed 'strict safeguards' on cyber and bio use cases to make the model safe for public release after deeming it 'too dangerous' a month ago. | Evidence: Schipper states: 'You can't use it for anything cyber related. You can't use it for anything biological related. That's what makes Anthropic comfortable releasing it to the general public.' | Caveat: No detail on how safeguards are enforced (prompt filters, output moderation, usage policies) or what 'cyber/bio' scope means (e.g., pentesting vs. malware, drug design vs. CRISPR). No independent verification of effectiveness. | Implication: Frontier models are being released with known risks and post-hoc controls rather than inherent safety. This could set precedent for future 'conditional release' strategies. Also: adversarial users will test these boundaries immediately. | Timestamp: 01:50

Detailed Brief

Model Architecture, Pricing, and Positioning

  • Claims: Claude 5 (marketed as Fable) is Anthropic's 'mythos-class' model, the largest in their lineup (Haiku/Sonnet/Opus/Mythos).; No architectural novelty—'just bigger and better' than prior models, per internal Anthropic sources.; Costs $10/M input tokens, $50/M output tokens (2x Opus pricing).; Released with cyber/bio safeguards after being deemed 'too dangerous' a month prior.
  • Evidence: Schipper had early access ~4 days before public launch; Every team tested for ~1 week.; 91/100 on Every's senior engineer benchmark vs. 63 (Claude 4.1), 62 (GPT-4.5).; Hexagon skill chart shows Fable saturating all dimensions vs. spiky prior models.; No specifics on safeguard enforcement or scope.
  • Caveats: Benchmark is internal and not peer-reviewed; no public replication data.; No token consumption or cost breakdowns for example projects (e.g., Library of Babel game).; Safeguards are stated but not demonstrated or tested in the video.
  • Implications: Pricing will limit adoption to high-value tasks or orgs with budget; Schipper expects price drop in 6-12 months.; If architecture is not novel, gains are from scale/training—suggests compute scaling still yields major jumps.; Cyber/bio restrictions may be circumvented by users or may not cover edge cases (e.g., social engineering, bio informatics).

Core Use Case: Sustained Autonomous Execution ('Warp Drive')

  • Claims: Best for multi-hour autonomous tasks where you 'give it a task and leave' for 3-4 hours or overnight.; Built a 3D Library of Babel game in one prompt over 3-4 hours: hexagonal galleries, 20 shelves/side, accurate to Borges story mathematics.; Created a Hubert Dreyfus lecture mini-site: sourced lectures autonomously, synced audio/text playback, custom player UI (±15s, 1x/2x speed, follow mode).; Batch-processed Every's Proof app GitHub issues: closed irrelevant issues, wrote fixes for the rest, merged to production.
  • Evidence: Library demo shown in video with navigable 3D environment, bookmarks, stairs.; Dreyfus site shown with drop caps, all-caps headers, font weight variation, playback sync.; No code or issue diffs shown for Proof app, but Schipper states 'fixes we merged.'
  • Caveats: No before/after comparisons with other models on same tasks.; Unclear if 'merged to production' means tests passed or just accepted; no mention of bugs/rework.; One-shot demos may not represent iterative use or edge case handling.
  • Implications: Overnight/weekend batch work (greenfield projects, backlog triage, research synthesis) becomes viable for solo operators.; This shifts value to problem definition/validation rather than implementation.; May obsolete certain contracting/outsourcing use cases (e.g., 'build me a quick prototype').

Taste, Attention to Detail, and Judgment

  • Claims: Outputs have 'taste': non-default fonts, drop caps, all-caps headers, better-than-slop aesthetics.; Model exhibits 'judgment': not 'try-hard' like prior Claudes (e.g., 'purple accents' overdrive); will push back if task is unclear or unfeasible.; UX patterns are thoughtful: playback controls (±15s, speed toggle, follow mode) for lecture site, accurate Borges mathematics for game.
  • Evidence: Dreyfus site shown with design details; Library of Babel shown with hexagonal geometry matching story.; Schipper contrasts with prior Claude 'purple accents' behavior (no timestamp for prior model example).
  • Caveats: No A/B test with other models on same prompts to isolate 'taste' vs. prompt interpretation.; Schipper predicts aesthetic quality will degrade as model becomes widely used (slop saturation effect).; Judgment/pushback behavior not demonstrated; stated only.
  • Implications: Non-designers can ship polished MVPs; design-as-a-service may compress further.; Model may reduce need for detailed prompts if it can infer intent and push back on ambiguity.; Slop saturation risk means early adopters get aesthetic advantage; later adopters compete on differentiation.

Multi-Context Research and Strategic Synthesis

  • Claims: Model excels at synthesizing large context (survey data, analytics, site content) into strategic insights.; For Every's 10K paid subscribers + 100K free, model surfaced: 'You have a conversion merchandising problem. Your free-to-paid conversion ratio is lower than it should be. Falsifiable bet: if we ship pricing transparency and a trial offer, it will go up.'; This insight outperformed weeks of internal team analysis using AI.
  • Evidence: Schipper states team had been analyzing survey data for weeks without this punchline.; Survey data described as 'hundreds to thousands of responses,' plus analytics and site content.
  • Caveats: No validation of hypothesis correctness or follow-up results.; No prompt shown; unclear how much context engineering or prior work informed the output.; Model may be pattern-matching common SaaS conversion advice rather than novel insight.
  • Implications: Strategic research (growth, ops, market analysis) becomes a one-prompt task for execs with data access.; Junior analyst/PM roles focused on synthesis may compress; value shifts to hypothesis testing and execution.; This capability could be a daily ritual for CEOs/execs (e.g., 'analyze all survey data and give me one action').

Adoption Maturity and User Segmentation

  • Claims: Only users at AI adoption levels 7-8 (orchestrating multiple agents, delegating 24/7 work) found Fable transformative.; Below level 6, users lack 'big meaty problems' and find the model overkill.; Every's internal test: ~7 people (engineers, writers, marketers); ~50/50 split on utility.; Vibe coders can use it if they can afford it, but should be 'careful' due to cost/token consumption.; Knowledge workers below advanced AI workflows won't see value.
  • Evidence: Reference to Every's 'eight levels of adoption' article (linked in show notes).; Schipper states 'about half of us are at a level where we can see the problems this thing solves.'
  • Caveats: Adoption framework is Every's own; no external validation or calibration.; No definition of what separates level 6 from 7 or how to self-assess.; Schipper predicts utility will expand to everyone 'soon,' but gives no timeline beyond 'eventually.'
  • Implications: AI adoption is now a skill/maturity gap, not just tool access. Orgs should assess user readiness before deploying expensive frontier models.; Marketing/demos for frontier models may overpromise utility for median users; GTM should segment by adoption maturity.; Training programs should focus on problem identification and multi-agent orchestration, not just prompt engineering.

Weaknesses and Daily Driver Comparison

  • Claims: Fable does not substantially outperform Claude 4.8 at writing; prose is dense, literary, not ideal for copywriting.; GPT-5.5 is better for writing tasks and Schipper's daily driver in Codex for fast iteration.; Fable is too slow, expensive, and token-hungry for daily collaboration or quick questions.; Pro tip: set reasoning level to medium or low (vs. max/extra high) for basic questions to reduce cost/speed.
  • Evidence: Schipper states he 'personally prefers GPT-5.5' and 'still uses 5.5 inside of Codex.'; Reasoning level tip attributed to internal Anthropic usage.
  • Caveats: No writing samples or benchmark scores provided to compare Fable vs. 4.8 vs. GPT-5.5.; Reasoning level adjustment not demonstrated; unclear if it significantly affects cost/speed.; No mention of whether Fable is better for 'thinking through writing issues' (stated) vs. sentence-level output.
  • Implications: Frontier model ≠ best-in-class for all tasks. Multi-model workflows (Fable for big tasks, GPT-5.5 for writing) are optimal.; OpenAI still leads on prose/copywriting, which may matter more for content/marketing teams.; Reasoning level tuning may become a standard cost-optimization practice for frontier models.

Philosophical and Strategic Framing: 'After Automation'

  • Claims: Automation creates more human work, not less—a paradox from Schipper's 'After Automation' essay.; Fable raises the floor for non-experts (vibe coders can make one-shot games) and the ceiling for experts (solo AAA games).; Change is normal to find sad/scary, but it's an opportunity to ask 'what can I do now that this makes possible?'; Capability will be cheap/accessible in 6-12 months, so this is a strategic preview.
  • Evidence: Reference to 'After Automation' essay (linked in show notes).; Schipper's personal framing: 'If you're someone who's used to typing code into your computer, this changes that a lot.'
  • Caveats: No evidence provided for the 'automation creates more work' claim; it's stated as a thesis.; 6-12 month timeline for price drop is Schipper's expectation, not Anthropic's commitment.; AAA game claim is speculative; no examples of solo AAA games built with prior models.
  • Implications: Orgs should plan for both displacement (typing code) and new roles (orchestrating agents, validating outputs).; Strategic value is in using this as a preview to plan 6-12 months ahead (hiring, process redesign, competitive moats).; The 'more human work' thesis suggests focus should shift to problem definition, taste-making, and decision-making vs. implementation.

Notable Concepts & Terms

  • Warp Drive (metaphor): Model compresses months/years of work into hours/days for big autonomous tasks, but is not suited for quick iteration ('getting around town'). Core mental model for when to use Fable vs. daily drivers like GPT-5.5.
  • Senior Engineer Benchmark (Every's internal): Tests model's ability to rewrite a real 'vibe-coded production codebase' from first principles. Fable scored 91/100 (vs. 63 for Claude 4.1, 62 for GPT-4.5), matching a human senior engineer. Used to anchor performance claims.
  • Eight Levels of AI Adoption (Every framework): Spectrum from level 1 (Google replacement for basic questions) to level 8 (orchestrating multiple agents, delegating 24/7 work). Only levels 7-8 users found Fable transformative. Key for assessing readiness.
  • Mythos Class (Anthropic's term): Largest model tier in Anthropic's lineup (Haiku/Sonnet/Opus/Mythos). Fable is the first mythos-class release. Not architecturally novel—'just bigger and better.'
  • Vibe Coder: Non-technical user who builds software via AI prompts/agents. Schipper says vibe coders can now make one-shot games with Fable, but should be careful due to cost/token consumption.
  • Slop Saturation: Schipper's prediction that aesthetic quality of AI outputs will degrade as models become widely used, because everyone will use the same defaults/patterns. Early adopters get aesthetic advantage.
  • Reasoning Level (Fable feature): User can set model to max/extra high/medium/low reasoning for cost/speed optimization. Internal Anthropic users set lower levels for basic questions. Not intuitive; requires awareness.
  • After Automation (Schipper essay): Thesis that automation creates more human work, not less. Used to frame Fable's impact: raises floor for non-experts and ceiling for experts, shifting value to problem definition and taste.

Operator Notes / Why Ken Should Care

  • Adoption maturity is now a strategic variable: segment users by 'eight levels' before deploying expensive frontier models. Levels 7-8 (multi-agent orchestration) see value; below level 6, it's overkill.
  • Cost model for frontier agents: Fable is $10/M input, $50/M output (2x Opus). Overnight/weekend batch tasks (greenfield, backlog triage, research synthesis) are the sweet spot. Price expected to drop in 6-12 months—plan workflows accordingly.
  • Multi-model workflows are optimal: Fable for warp-drive tasks, GPT-5.5 for writing/iteration, Claude 4.8 for prose if staying in Anthropic ecosystem. Daily driver ≠ frontier model.
  • Benchmarking is a bottleneck: Every's 'senior engineer benchmark' saturated in one release cycle. If you're building agents/tools, create harder benchmarks or risk commoditization.
  • Taste/aesthetics as moat: early Fable outputs have better-than-default design (drop caps, fonts, UX). Slop saturation predicted as model spreads—differentiate on taste and brand, not just AI use.
  • Strategic synthesis as one-prompt task: feeding survey + analytics + site content into Fable surfaced 'conversion merchandising problem' with falsifiable next step. This workflow could replace junior analyst/PM roles or become daily exec ritual.
  • Safeguards as GTM strategy: Anthropic released Fable with cyber/bio restrictions after calling it 'too dangerous.' This sets precedent for 'conditional release' vs. holding back models. Watch for adversarial testing of boundaries.
  • Vibe coding compression: non-technical users can now ship polished MVPs (games, tools, sites) in one prompt. This could obsolete low-end dev contracting or create new market for 'agent-native design.'
  • 6-12 month horizon for mass access: Schipper expects Fable-level capability to be cheap/accessible within a year. Use this period to build proprietary workflows and train teams, not just wait for price drop.

Watch Map

  • 00:00: Library of Babel 3D game demo—built in one prompt over 3-4 hours.
  • 01:15: Introduction: Dan Schipper (Every CEO), model launch context, 'too dangerous' leak a month ago.
  • 01:50: Model specs: mythos-class, $10/M input + $50/M output (2x Opus), cyber/bio safeguards.
  • 02:40: Senior engineer benchmark: 91/100 (Fable) vs. 63 (Claude 4.1) vs. 62 (GPT-4.5). Hexagon chart.
  • 04:15: Warp drive metaphor: compresses months/years into hours/days for big tasks, not for daily iteration.
  • 05:30: Strengths: sustained autonomous execution, taste/detail, judgment, multi-context research.
  • 07:30: Example 1: Hubert Dreyfus lecture mini-site—sourced lectures, synced audio/text playback, custom UI.
  • 09:15: Example 2: Every survey synthesis—'conversion merchandising problem' + falsifiable bet.
  • 11:00: Example 3: Proof app GitHub issue triage—batch closed/fixed issues, merged to production.
  • 13:00: Who should use: levels 7-8 (multi-agent orchestration) see value; below 6, it's overkill. 50/50 split on Every's team.
  • 14:30: Weaknesses: not better than Claude 4.8 at writing (dense, literary), slow/expensive for daily use. GPT-5.5 preferred for iteration.
  • 15:45: Reasoning level tip: set to medium/low for basic questions to reduce cost/speed.
  • 16:30: Meaning of Fable: raises floor (vibe coders → one-shot games) and ceiling (experts → solo AAA games). 'After Automation' thesis: automation creates more human work.
  • 17:45: Closing: price drop expected in 6-12 months, go use your warp drive, read 'After Automation' and Every's full vibe check.

Source/Metadata

  • Title: Anthropic's Fable 5: A Warp Drive for Coding
  • Transcript words: 7800
  • Duration seconds: 997
  • Timestamp note: Timestamps extracted from video duration (997 seconds ≈ 16:37); aligned to transcript segment transitions where possible.
Full transcript 3872 words · 35 min read
0:00

SPEAKER_00

This is the Infinite Library of Babel from the Borges story. It contains all of the books in the universe because books are just strings. If you look, you can even go into bookmarks and I can click one of my articles after automation and it finds it in the library. It's truly infinite. And look, I could go up the stairs. I can look down. I can look up. This seems like it took a long time to make, right? Wrong. I made this entire thing in a single prompt with Claude 5, the new model from Anthropic. Let me show you. So this is a prompt from four days ago or so. I got this model a little bit ahead of time. Read Jorge Luis Borges's The Library of Babel and then plan and execute end to end a browser playable 3D game in which the player has dropped in, blah, blah, blah, loop until it's done. Just wrote that, press enter and it just went off and read the story and then it just ran and ran and ran. You can see it's looping itself and it's checking its work. After three or four hours or so, done. We've got hexagonal galleries stacked endlessly. We've got 20 shelves, five per side. Everyone would expect it's accurate to the book. It's got the mathematics, right? It says the part I'm proudest of. This is crazy. It just made it in one shot in three or four hours running on its own. Claude 5 launches today. Here's your day zero vibe check. But first, remember to never make any major life decisions within 30 days of a meditation retreat, a psychedelic experience, or your first encounter with a frontier model. Cheers.

0:05

SPEAKER_00

So before we get into it, you're probably wondering how is this video even out? My name is Dan Schipper. I'm the co-founder and CEO of Every. Every is the only subscription you need to stay at the edge of AI. You can think of us as an AI lab for the future of work. We spend all of our time testing new models, using them to do our work from programming to writing to design to business building to decision making. We use them hands on when we tell you about what works and what doesn't for real use cases. And I'm incredibly excited to be doing this because the first encounter with any new model could be crazy. But Claude, which is a frontier class model, I think is a particularly big moment. It's the most hyped model when it leaked a month and a half ago. Anthropic said it was too dangerous to even release it. And now it's out. And I have a feeling that if you're like me, you might be excited, but you're also a little scared.

0:09

SPEAKER_00

Because we've been using this model for about a week now, we get to pull back the curtain a little bit and show you what it's like to have lived with this model a little bit more. It does change things, but hopefully this can help alleviate if you're feeling a little bit of AI psychosis. I'm sure that's going to be going around on X and YouTube and the news and all that. This is a place for you to see how this thing might fit into your work and into your life in a realistic way. So let's get into it.

0:14

SPEAKER_00

Okay, so Claude is a frontier class model. For Anthropic, frontier is the model category. It's the largest model that they make. There's Haiku, Sonnet, Opus, and then Claude. As far as I can tell from talking to people internally at Anthropic, there's nothing special about it architecturally. It's basically the same thing as their other models, it's just bigger and better. In order to make it safe to release, they have put pretty strict safeguards on it. So you can't use it for anything cyber related. You can't use it for anything biological related. That's what makes Anthropic comfortable releasing it to the general public. It's pretty expensive. It's $10 per million input tokens and $50 per million output tokens, which is about twice the cost of Opus. So it's a lot, but it is genuinely the most powerful coding model I've ever used by far.

0:21

SPEAKER_00

To give you a sense, we have a senior engineer benchmark, which basically tests the model on its ability to act like a human senior engineer. We give it a vibe coded production code base, a real production code base. And we ask it, if you're going to rewrite this from first principles, how would you do it? And then we see how it does. We score it out of 100. The best model score is a 63 out of 100, which is Claude 4.1, which came out like two weeks ago. And right behind that is GPT-4.5, which is a 62 out of 100. Claude scored a 91 on this benchmark. 91 out of 100. That's the same score as a human engineer with just one prompt. That's crazy. I knew that this benchmark was going to get saturated, but I thought it would happen in like six months.

0:30

SPEAKER_00

Look at this view of it when we break it down by what it's good at versus other models. This is Claude 4.1. The orange stuff is what it does versus what it's going to do. Pretty spiky, not that great. GPT-4.5. We're starting to fill out the hexagon a little bit. This is just, yeah, it just did it.

0:35

SPEAKER_00

If I try to break down for you what it's really good at, because it's not good for everything. I think it's fantastic at sustained autonomous execution. For example, the way to work with this model is to give it a task and then leave, go do something else, let it go for three or four hours, set it up overnight. It's amazing. It just figures stuff out. And it just does good work. It has good taste. It has good attention to detail. There are all these little details that it does pretty well that I'll show you. That's really impressive with a not very well specified prompt. It has more judgment than previous models. With prior models, you'd be like, "Do this thing." And it'd be like, "Oh my God, yes, I'm going to do it. I'm going to do it. Purple accents, purple accents." It was a little try hard, to be honest. And this model, it feels like it's going to go do it and it's going to think it through and think about how to do it well. And if it doesn't think it can do it well, it might say something to you, which is really helpful. And it's also incredibly good at using a lot of context, doing a bunch of research, digging into data, giving you a bunch of things from the data that you wouldn't have known beforehand. I'm going to go through a bunch of really specific examples of exactly what this is and how it works. But if I step back and I think about it, in particular for programming,

0:39

SPEAKER_00

A little try hard, to be honest. And this model, it feels like it's going to go do it. And it's going to think it through and think about how to do it well. And if it doesn't think it can do it well, it might say something to you, which is really helpful. And it's also just incredibly good at using a lot of context, doing a bunch of research, digging into data, giving you a bunch of things from the data that you wouldn't have known beforehand. I'm going to go through a bunch of really specific examples of exactly what this is and how it works. But if I step back and I think about in particular for programming, what is this model? It's like a warp drive. You know, like in Star Wars, you want to jump across the galaxy. You punch the coordinates into the computer. The computer does some calculations to see, okay, we don't want to go through a star or whatever. Then you punch it and the stars blur into blue. And you're like, you don't get there instantly with a warp drive. You're going across the galaxy, but it takes you a couple hours or a couple days, which is pretty good because it used to take you a year or two, or maybe 20 or 100 years. But with a warp drive, you get there in a couple hours or a couple days. And that's what this model is. You can specify a destination for a big trip, and it compresses what normally would have been years or months into hours or days. But it's not really that good for getting around town. You wouldn't use a warp drive to get around town. You need more control. You need more feedback back and forth between you and the vehicle. And the same thing is true for this model. If you're using it for true collaboration or quick questions or things that need tight back and forth, I don't think it's that good for that. It's very slow. It's very expensive. It's extremely token hungry. One thing that's a pro tip that's not intuitive for this model is you can set it to lower reasoning levels. Instead of max or extra high, you can put it to medium or low for more basic questions. That's not something that was very intuitive to me, but that's how people inside of Anthropic use it. So if you want to use it a lot and you're trying for more of the everyday thing, you can do this. But I think it's a warp drive. It's really good for those big things, but you have to have big meaty things to give it. I want to show you some examples of some of the things that we saw in testing this that'll tell you some of the properties of this model and maybe give you some inspiration for where you might want to use it.

0:45

SPEAKER_00

Okay, so first example, there's this philosopher I love. His name is Hubert Dreyfus. He wrote a book called "What Computers Can't Do" in the 70s, which is a famous explanation of why AI at the time was not going to work. I really like him. He's really into Heidegger, a philosopher, and he has these lectures on Heidegger from 2007. But they're hard to follow. The audio recording is not that good. So I just told Claude, hey, can you go get his lectures and turn it into a little mini site for me so that they're easier to consume. That's literally all I did. I didn't even link it to the lectures. I just said grab the lectures. I didn't know where they were. It went and grabbed the lectures. It wrote a little summary of why does this matter. These are audio lectures. I broke it down into a table of contents. And then it has this player experience where I can press play and it'll play him, and it highlights what he's saying as he's saying it. It synced the playback to the text. And it has this player thing with minus 15 and plus 15, and you can do one x or two x, and you can follow it or not follow it. This is from one prompt. And that's what I mean when I talk about this model having really exceptional taste and attention to detail. There's a lot of different dimensions along which this is happening. Look at the font choices. This is all caps. There's a little more font weight on here. It does the drop cap here. These are not the defaults that you would expect. It's not the bland purple highlight vibe slop thing. I assume some of this stuff will look sloppier at some point because everyone's going to be using this model to do stuff. But for now, it really looks good. And the things that it builds feel much more thoughtful. These are things obviously you could have built with a model before, but now they're just available to you almost for free. And I think that's a good test of a model: does it get you to try a bunch of new stuff because now stuff that used to be hard is easy? And it opens up a whole realm of things. That's 100% true with this model.

0:50

SPEAKER_00

This is my first one. This is a taste and detail example. Another thing that it does, which I think is so cool, is it's so good at using its context and doing research. At every paid subscription, we've got 10,000 paying subscribers. If you have not subscribed, you should subscribe. Every dot to slash subscribe. You'll get a text version of this vibe check if you want more detailed insights. So check that out. Maybe 100,000 free subscribers, and we do surveys. We want to understand what do people think about us? We have a bunch of recent survey data. We fed it into this model and look at what it did. A team of us have been looking at this data for weeks now with AI, and we have not come out with anything this interesting or succinct. This is from hundreds and hundreds and hundreds of survey responses, maybe thousands of survey responses. The punchline at the top: you have a conversion merchandising problem. Your free to paid conversion ratio is lower than it should be. Okay, next, the falsifiable bet. If we ship pricing transparency and a trial offer, I think it's going to go up. That is something I would expect a really good growth person to do with a lot of time and thought and research. It requires thinking across so many different dimensions. It has to look at all the data. It has to look at all the survey responses. It has to look at our analytics data. It has to then go through the site itself and put that all together in a way where it can tell you, here's the punchline, and here's what I would do next in a way that I

0:55

SPEAKER_00

Okay, next, the falsifiable bet. If we ship pricing transparency and a trial offer, I think it's going to go up. That is something that I would expect a really good growth person to do with a lot of time and thought and research. It requires thinking across so many different dimensions. It has to look at all the data. It has to look at all the survey responses. It has to look at our analytics data. It has to then go through the site itself and put that all together in a way where it can tell you, here's the punchline. And here's what I would do next in a way that I could just scan. We've not seen a model that's able to do this. And I think this applies to lots and lots of different challenges, whether you're using it for coding or knowledge work or whatever you want to use it for. So far, I've shown you a lot of cool one-shot demos and stuff like that. The other thing I did is we have this app called Proof. It's an agent-native markdown editor. This is Proof. We have a bunch of issues in GitHub with Proof because agents can submit issues. So as you're using it with your agent, it just submits a bunch of issues, a bunch of them come in every day. I just pointed it at our issues in GitHub. And I just said, okay, I want you to take all the issues for the last couple weeks and close any that aren't relevant and write fixes for all the rest. And it just went boom, boom, boom, boom, boom, boom, boom, and actually wrote fixes that we merged. Again, other models can do this, but it's much more, okay, you go one at a time, you make sure it's doing well, you can't just be like, go do it. This is why it's a warp drive. It just speeds through the backlog of simple things like this in a way that would be impossible with other models. And now the question might be in your mind, like who should use this? And I really don't think that this model as it is right now is for everyone. Again, it's slow, expensive, it's super powerful. But because of that, it's not for everyone. We have this article on Anthropic called the eight levels of adoption, which you should definitely read. You can actually throw it into your agent, and we'll put the link in the show notes. Your agent can just go through how you use AI, and it'll break it down into eight levels. Everything from at the bottom level, you're just using it essentially like a Google replacement to just ask basic questions to the top level. It's like you're orchestrating many different agents where you're delegating work and it's working 24/7 and all that kind of stuff. That's the spectrum. And everyone falls on a different place in the spectrum. We probably had seven or so people testing it internally over the last week, everyone from programmers to writers to editors to marketers. What we found is there's actually a pretty high spread of who liked it and who didn't. And it's not like anyone hated it. But if you're not at a certain level of AI workflow, you're kind of like, I don't know what to use this for. Because you don't have a problem that's big enough where you need to speed through the galaxy. What we found is, if you're like a seven or eight on the scale, you know, you're using multiple agents and you're orchestrating them and all that kind of stuff, you've got big meaty problems, you're like, wow, this is crazy. And usually that's technical people. If you're non-technical, if you're a vibe coder, you're probably watching this being like, I have so many projects I want to do. As long as you can afford it, I'll say it again, this model is expensive and it hogs a ton of tokens. At least for now, if you're vibe coding, I would be careful with it. But I would definitely try it. And then if you're a knowledge worker and you're just using it to get your job done, unless you're a very advanced knowledge worker using it in this sort of way, where you're orchestrating multiple agents together and delegating a lot of your work, it's going to feel like overkill. And I think that's a really interesting thing, is that what we're finding is, yes, you have this AI, but using it is a skill. You need to be exposed to problems and working at a level of expertise where the problems come up in order for it to be useful.

1:00

SPEAKER_00

[SPEAKER_01] Internally, like all of us are early adopters of AI, I'd say probably about half of us are at a level where we can see the problems that this thing solves. And about half of us are still getting there. I think we will. I actually think that this kind of workflow is going to be available to everyone soon, or it's going to be useful for everyone soon. But it just depends on where you are on the adoption curve.

1:05

SPEAKER_00

And to get into a couple things the model doesn't do as well, it didn't actually do substantially better at writing than Opus 4.8. So if you're using it for writing, we found its sentences to be pretty dense, big blocks, and they're pretty literary. So for some things, it can be good for that. And I do think it's very good for thinking through writing issues. But for actually writing sentences for copywriting, it's probably not the thing you're going to want to use. If you're a Claude person, just use 4.8. If you're a GPT person, 5.5 is much better. I personally prefer 5.5. And I still use 5.5 inside of Codex as my daily driver, because most of the stuff I'm doing, I want to go back and forth pretty quick. And a model like this is going to be overkill. So this model increases my confidence for big projects, or maybe writing production code like that kind of stuff. But for my day-to-day, it's a bit overkill, even for me. Finally, let's talk a little bit about what is the meaning of Fable. We want to go past "hype, oh my God, it's going to change everything." And to some extent, that's actually right. But it's not going to change everything in a way that I think people imagine. I just wrote this piece called "After Automation," which you should read. It's about what is work like after we've automated everything. It turns out automation actually creates a lot more human work. It's a very interesting paradox. I think the same is true here. What we're going to find is this model increases the floor of capability for non-experts, but it also raises the ceiling for experts. So a vibe coder might be able to make a one-shot video game. And an expert might be able to make

1:11

SPEAKER_00

Oh my God, it's going to change everything. And to some extent, that's actually right. But it's not going to change everything in a way that I think people imagine. I just wrote this piece called "After Automation," which you should read, which is about what is work like after we've automated everything. It turns out automation actually creates a lot more human work. It's a very interesting paradox. I think the same is true here. What we're going to find is this model increases the floor of capability for non-experts, but it also raises the ceiling for experts. So a novice coder might be able to make a one-shot video game. And an expert might be able to make a true triple-A game just by themselves. And I think that is so cool. Obviously, the fact that things are changing this much, and I think it's really important for us to say, this does change things. If you're someone who's used to typing code into your computer, this changes that a lot. And I think it's normal to be sad or angry or weirded out by that. And it also changes things even if you're using AI already. This changes how you can expect to use it in the future. A lot of the skills, a lot of the things that you thought you might have to do, are starting to change a bit because this model is so much more powerful. Change can be scary, but it's also an opportunity to be like, wow, what can I do now that I might be into that this now makes possible that I can just do? I don't need to ask permission, I don't need more money, I don't need anything. And because this capability is out now, we can expect that even if it's too expensive for you to use right now, it's going to be pretty cheap soon. Let's say within the next six months to a year, everyone will be able to have this. And I think that is incredible. So if you like this video, you should really watch my video "After Automation," which talks about what happens when we automate everything and read the article. You should also read our vibe check on everything. We go in depth through every part of the testing that we did for this model, all the benchmarks from coding to writing to knowledge work. We have takes from the entire team. We have a bunch of people testing this from different perspectives and different walks of life and different ways that they like to use AI. But if you're psyched about this, the thing I recommend most is go use your new warp drive and let me know what you make.

1:16

SPEAKER_00

Here's your day zero vibe check. But first, remember to never make any major life decisions within 30 days of a meditation retreat, a psychedelic experience, or your first encounter with a frontier model. Cheers. So before we get into it, you're probably wondering how is this video even out? My name is Dan Schipper. I'm the co-founder and CEO of Every. Every is the only subscription you need to stay at the edge of AI. You can kind of think of us as like an AI lab for the future of work. We spend all of our time testing new models, using them to do our work from programming to writing to design

1:47

SPEAKER_00

to business building to decision making. We use them hands on when we tell you about what works and what doesn't for real use cases. And I'm incredibly excited to be doing this because the first encounter with any new model could be crazy. But Fable, which is a mythos class model, I think is like is a particularly big moment. It's like the most hyped model when it leaked a month and a half ago. Anthropic said it was too dangerous to even release it. And now it's out. And I have a feeling that if you're like me, you might be excited, but you're also like a little scared. Because we've been using this model for about a week now, we get to pull back the curtain a little

2:18

SPEAKER_00

bit and show you what it's like to have lived with this model a little bit more. It does change things, but hopefully this can help alleviate. If you're feeling a little bit of AI psychosis, I'm sure that that's going to be going around on X and YouTube and the news and all that kind of stuff. This is a place for you to see how this thing might fit into your work and into your life and in a realistic way. So let's get into it. Okay, so Fable is a mythos class model. Mythos is a model for Anthropic. It's the largest model that they make. There's haiku, sonnet, opus, and then mythos. As far as I can tell from talking to people internally at Anthropic,

2:50

SPEAKER_00

there's nothing special about it architecturally. It's basically the same thing as their other models, it's just bigger and better. In order to make it safe to release, they have put pretty strict safeguards on it. So you can't use it for anything cyber related. You can't use it for anything biological related. That's what makes Anthropic comfortable. We're releasing it to the general public. It's pretty expensive. It's $10 per million input tokens and $50 per million output

3:09

SPEAKER_01

tokens, which is about twice the cost of Opus. So it's a lot, but it is just genuinely the most powerful coding model I've ever used by far. To give you a sense, we have a senior engineer

3:18

SPEAKER_00

benchmark, which basically tests the model on its ability to act like a human senior engineer. We give it a vibe coded slop production code base, a real production code base. And we ask it, if you're going to rewrite this from first principles, how would you do it? And then we see how it does. We score it out of 100. The best model score is a 63 out of 100, which is Opus 4.8, which came out like two weeks ago. And right behind that is GPT 5.5, which is a 62 out of 100. Fable scored a 91 on this benchmark. 91 out of 100. That's the same score as a human engineer with just a just just one prompt. That's that's it's crazy. I like I knew that this benchmark was

3:55

SPEAKER_00

going to get saturated, but I thought it would happen in like six months. Look at this view of it when we break it down by what it's good at versus other models. This is Opus 4.7. The you know, the orange stuff is what it what it does versus what it's going to do. You know, pretty spiky, not that great. GPT 5.5. Like we're starting to fill out the hexagon a little bit. This is just like, oh, yeah, it just did it. If I try to like break down for you, okay, what is it really good at? Because it's not good for everything. I think it's fantastic at sustained autonomous execution. Like, for example, the way to work with this model is to give it a task and then leave,

4:29

SPEAKER_00

go do something else, let it go for three or four hours, set it up overnight. It's, it's amazing. Like it just figures stuff out. And it just does good work. It has good taste. It has good attention to detail. There's all these like little details that it does pretty well that I'll show you. That's that's really impressive with a not very well specified prompt. It has it has more judgment. I think previous cloud models, you'd be like, Oh, do this thing. And it'd be like, Oh, my God, yes, I'm going to do it. I'm going to do it. And then purple accents, purple accents. It was like a little try hard, to be honest. And this model, it feels like it, it's going to go do it. And

5:00

SPEAKER_00

it's going to think it through and think about how to do it well. And if it doesn't think it's, it can do it well, it'll, it might say something to you, which is, which is really helpful. And it's also just incredibly good at like using a lot of context, like doing a bunch of research, digging into data, giving you a bunch of things from the data that you wouldn't have known beforehand. I'm going to go through a bunch of really specific examples of exactly what this is and how it works. But the, if I step back and I think about in particular for programming, like, what is this model? It's like a warp drive. You know, like in Star Wars, they, you know, you

5:30

SPEAKER_00

want to, you want to like jump across the galaxy, you like, you know, you punch out, you punch the coordinates into the computer, the computer, just some calculations to see like, okay, we don't want to, we don't want to like go through a star or whatever. Then you punch it into like the stars blur into blue. And you're like, you don't get there instantly with a warp drive, you're going across the galaxy, but it takes you like, you know, a couple hours or a couple days, which is pretty good because it used to take you like a year or two, or maybe 20 or 100 years. But with a warp drive,

5:58

SPEAKER_00

you get there in a couple hours or a couple days. And that's kind of like what this model is like, you can specify a destination for a big trip. And it just like it compresses what normally would have been like years or months into like hours or days. But also like, it's not really that good for getting around town, you know, you wouldn't use a warp drive to get around town, you need like, you need more control, you need more feedback back and forth between you and the vehicle. And the same thing is true for this model. If you're using it for like true collaboration or quick questions or

6:27

SPEAKER_00

things that need tight back and forth. I don't think it's that good for that. I mean, it's very slow. It's very expensive. It's extremely token hungry. One thing that's a sort of a pro tip that's not intuitive for this model is you can set it to lower reasoning levels, like, you know, instead of max or extra high, you can put it to medium or low for more basic questions. That's not something that was very intuitive to me. But that's how people inside of Anthropic use it. So if you want to use it a lot, and you're trying for more of the like everyday thing, you can do this.

6:54

SPEAKER_00

But I just I think it's I think it's a warp drive. It's really good for those big things. But you have to you have to have big meaty things to give it. I want to show you some examples of some of the things that we saw in testing this that'll tell you some of the properties of this model and maybe give you some inspiration for where you might want to use it. Okay, so first example, there's this philosopher I love his name is Hubert Dreyfus. He wrote a book called what what computers can't do in the 70s, which is like a famous explanation of why AI at the time was not going to work really like him. He's

7:23

SPEAKER_00

really into Heidegger, a philosopher and he has these lectures on Heidegger from like 2007. But they're kind of hard to follow the audio recording is not not that good. So I just told fable, hey, can you like go get his go get his lectures and turn it into a little mini site for me so that they're easier to consume. That's literally all I did. I didn't even link it to lectures. I just said, grab the lectures. I didn't know where they were. It went and grabbed the lectures. It wrote like, look, it wrote like a little summary. Why does this matter? These are audio lectures, I broke it down into a table of contents. And then it has this like player experience where,

7:54

SPEAKER_00

you know, I can press play. And it'll play him. And, and it highlights what he's saying as he's saying it like it synced the the playback to the text. And it has this like player thing. And it's got like, you know, minus 15 and plus 15. And you know, you can do one x or two x and you can follow it or not follow it like this is from one prompt. And that's what I mean when I talk about this model having really exceptional taste and attention to detail. There's a lot of different dimensions along which this is happening. Look at the font choices. This is all caps. There's a little bit more font

8:27

SPEAKER_00

weight on here. It does the drop cap here. These are not like the defaults that you would expect. It's not like the clod purple highlight vibe slop thing. I assume some of this stuff will look sloppier at some point because everyone's going to be using this model and using to do stuff. But for now, it really looks good. And the things that it builds feel much more thoughtful. These are things obviously you could have built with a model before, but now they're just available to you almost for free. And I think that's like a good test of a model is does it get you to just like try a bunch of new stuff because now stuff that used to be hard is easy. And so it just opens up

9:00

SPEAKER_00

a whole realm of things. That's 100% true with this model. So this is my first one. This is a kind of taste and detail example. Another thing that it does, which I think is so cool is it's so good at using its context and doing research at every paid subscription. We've got like 10,000 paying subscribers. If you have not subscribed, you should subscribe every dot to slash subscribe. You'll get a text version of this vibe check if you want more detailed insights. So check that out. So whatever, maybe 100,000 free subscribers. And we do surveys. So we want to understand what do people think about

9:33

SPEAKER_00

us? We have a bunch of recent survey data. So we fed it into this model and look at what it did. Like a team of us have been looking at this data for weeks now with AI, and we have not come out with anything this interesting or succinct. This is from like hundreds and hundreds and hundreds of survey responses, maybe in the thousands of survey responses. The punchline at the top, you have a conversion merchandising problem. Your free to paid conversion ratio is lower than it should be. Okay, next, the falsifiable bet. If we ship pricing transparency and a trial offer, I think it's

10:05

SPEAKER_00

going to go up. Like that is something that I would expect a really, really good growth person to do with a lot of time and thought and research. It requires it to think across so many different dimensions. It has to go look at all the data. It has to go look at all the survey responses. It has to go look at our analytics data. It has to then go through the site itself and put that all together in a way where it can tell you, here's the punchline. And here's what I would do next in a way that I could just scan. We've not seen a model that's able to do this. And I think this applies to lots and

10:36

SPEAKER_00

lots of different challenges, whether you're using it for coding or knowledge work or whatever you want to use it for. So far, I've shown you like a lot of cool one shot at demos and stuff like that. The other thing I did is we have this we have this app called proof. It's a agent native markdown editor. This is this is proof. We have a bunch of issues in the GitHub with proof because agents can submit issues. So as you're using it with your agent, it just submits a bunch of issues, a bunch of them come in every day. I just pointed it at our issues in GitHub. And I just said like, okay, I want you to

11:01

SPEAKER_00

take all the issues for the last couple weeks and close on any that aren't relevant and write fixes for all the all the rest. And it just went boom, boom, boom, boom, boom, boom, boom, and actually wrote fixes that we merged. Again, other models can do this, but it's much more, okay, you go one at a time, you make sure it's doing well, you can't just be like, go do it. This is this is why it's it's a it's a warp drive. It just speeds through the backlog of simple things like this in a way that would be impossible with other models. And now the question might be in your mind, like who should use this? And I really

11:27

SPEAKER_00

don't think that this model as it is right now is for everyone. Again, it's slow, expensive, it's super powerful. But because of that, it's not for everyone. We have this article on every called the eight levels of adoption, which you should definitely read, you can actually throw it into your agent, and we'll put the link in the show notes, your agent can just go through how you how you use AI, and it'll break it down into eight levels. Everything from at the bottom level, you're just using it essentially like a Google replacement to just like ask basic questions to the top level. It's like, you're orchestrating many, many different agents where you're delegating work

11:55

SPEAKER_00

and and it's working 24 seven and all that kind of stuff. That's the kind of that's the spectrum. And everyone falls on a different place in the spectrum, we probably had seven or so people testing it internally over the last week, everyone from programmers to writers to editors to marketers, what we found is there's actually like a pretty high spread of who liked it and who didn't. And it's not like anyone hated it. But if you're if you're not at a certain level of AI workflow, you're kind of like, I don't know what to use this for. Because you don't you don't have a problem

12:26

SPEAKER_00

that's big enough where you need to speed through the galaxy. What we found is, if you're like a seven or eight on the scale, you know, you're you're using multiple agents, and you're orchestrating them and all that kind of stuff, you've got big meaty problems, you're like, wow, this is crazy. And usually that's technical people. If you're non technical, if you're a vibe coder, you you probably are watching this being like, holy shit, I have so many projects like I want to do as long as you can afford it. I'll say it again, this model is expensive, and it hogs a ton of tokens, at least for now, if you're vibe coding, I would be I would be careful with it. But I would

12:54

SPEAKER_00

definitely definitely definitely try it. And then if you're kind of like if you're a knowledge worker, and you're just using it to get your job done, unless you're a very advanced knowledge worker, that's using it in this sort of way, like you're you're orchestrating multiple agents together and delegating a lot of your work, it's going to feel like overkill. And and I think that's that's a really interesting thing is like, what we're finding is, yeah, yeah, you have this AI, but using it is a skill, you need to be exposed to problems and working at a level of expertise where the problems come up in

13:24

SPEAKER_01

order for it to be useful. Internally, like all of us are early adopters of AI, I'd say,

13:27

SPEAKER_00

probably about maybe half of us are at a level where we can see the problems that this thing solves. And about half of us are still getting there. I think we will, I actually think that this kind of workflow is going to be available to everyone soon, or it's going to be useful for everyone soon. But it just depends on your on your where you are on the adoption curve. And to get into a couple things the model doesn't do as well, it didn't actually do substantially better at writing than Opus 4.8. So if you're using it for writing, we found it sentences to be pretty dense, you know, like big blocks, and they're pretty literary. So for some things, it can be good for

13:59

SPEAKER_00

that. And I do think it's very good for thinking through writing issues. But for actually writing sentences for copywriting, it's probably not the thing you're going to want to use. If you're a cloud person, just use 4.8. If you're a GPT person, 5.5 is much better. I personally prefer 5.5. And I still use 5.5 inside of codex as my daily driver, because most of the stuff I'm doing, I want to go back and forth pretty quick. And a model like this is going to is going to be overkill. So this this model, like it increases my confidence with for big projects, or maybe writing production

14:29

SPEAKER_00

code like that kind of stuff. But for my day to day, it's a bit overkill, even for me, finally, like, let's talk a little bit about what is what is the meaning of fable, we want to go past home, oh, my God, it's going to change everything. And to some extent, that's actually right. But it's not going to change everything in a way that I think people imagine. I just wrote this piece called after automation, which you should read, which is about what is it what is work like after we've automated everything, it turns out automation actually creates a lot more human work. It's a very interesting paradox. I think the same is true here. What we're going to find is this model

15:00

SPEAKER_00

increases the floor of capability for non experts, but it also raises the ceiling for experts. So a vibe coder might be able to make a one shot video game. And an expert might be able to make like a true triple A game just by themselves. And I think that I think that is so cool. Obviously, the fact that things are changing this much, and I think it's really important for us to say, this does change things. If you're someone who's used to typing code into your computer, this changes that a lot. And I think it's normal to be sad or angry or weirded out by that. And it also changes things even if you're using AI already, this this changes how you can expect to use it in the

15:35

SPEAKER_00

future. A lot of the skills, a lot of the things that you thought you might have to do, are starting to change a bit because this model is so much more powerful change can be scary. But it's also an opportunity to be like, wow, what can I do now that I might be into that this now makes possible that I can just do I don't need to ask permission, I don't need more money, I don't need anything. And and because this capability is out now, we can expect that even if even if it's too expensive for you to use right now, it's going to be pretty cheap soon. Like, let's say within the next six months to a year, everyone will be able to have this. And I think that's, I think that is

16:06

SPEAKER_00

incredible. So if you like this video, you should really watch my video after automation, which talks about what happens when we automate everything and read the article, you should also read our vibe check on every we go in depth through every part of the testing that we did for this model, all the benchmarks from coding to writing to knowledge work, we have takes from the entire team, we have a bunch of people testing this from different perspectives and different walks of life and different ways that they like to use AI. But if you're psyched about this, the thing I recommend most is go use your new warp drive. And let me know what you make.

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note