FABLE 5.1 IS A BEAST
Description
Read the full Fable 5.1 vibe check: https://every.to/vibe-check/fable-5-1-vibe-check?utm_source=youtube&utm_medium=video&utm_campaign=fable_5_1&utm_content=vibe_check_article Fable 5.1 prompt library: https://every.to/claude-fable-5-prompt-library?utm_source=youtube&utm_medium=video&utm_campaign=fable_5_1&utm_content=prompt_library Follow Dan Shipper: https://x.com/danshipper Follow Kieran Klaassen: https://x.com/kieranklaassen Follow Katie Parrott: https://x.com/kplikethebird Follow Mike Taylor: https://x.com/hammer_mt Follow Every: https://x.com/every
Summary
Generated by gpt-5.6-terraAt-a-Glance
- Verdict: Skim
- Core thesis: Fable 5.1 is presented as a major usability upgrade over Fable 5: it preserves long-horizon autonomous coding capability while becoming faster, more token-efficient, more collaborative, and better at judgment-heavy knowledge work.
- Why it matters: The meaningful claim is not merely higher benchmark performance: the speakers argue that Fable 5.1 makes credible delegation possible for coding and selected knowledge-work tasks, reducing the human iteration previously required to reach polished outputs.
- Best use: Skim the demos and discussion of long-running agents, one-shot application builds, deck/dashboard generation, and effort-level routing; ignore the repeated promotion and broad model hype.
Executive Summary
Every’s team argues that Fable 5.1 fixes the central weakness of Fable 5: the prior model was exceptionally capable at large autonomous coding tasks but expensive, verbose, and difficult for nontechnical users to steer. In their view, 5.1 retains the ability to run independently for hours on complex builds while gaining a more direct, collaborative interaction style they compare to a return of the more approachable “Claude” personality.
Their strongest evidence is a series of hands-on, mostly qualitative evaluations. Kieran reports that a six-hour autonomous run rebuilt a sophisticated collaborative editor with CRDT-like collaboration, review flows, provenance, inline AI, activity tracking, and polished interaction details. The team emphasizes that the differentiator is not simply generating a working prototype, but making useful design and product judgments—such as restrained animation, appropriate UI controls, and small interaction details—without extensive human prompting.
For knowledge work, the speakers see a meaningful rise in “discernment”: identifying the relevant tension in a piece of writing, surfacing decisions worth executive attention from meeting records, extracting non-obvious patterns from NPS feedback, and building readable presentation decks. They claim Fable 5.1 is especially compelling against Opus 5 because it allegedly provides superior quality at roughly similar overall cost through lower token use, though this is based on internal usage rather than a published controlled cost study.
The video is useful primarily as an early practitioner signal about how to use a frontier model: medium/high effort for interactive work, extra-high effort for long autonomous runs where polish matters, and explicit design direction when visual storytelling or aesthetics are important. It is not a rigorous benchmark report, and the presenters themselves retain preferences for competing models in clean prose and certain writing workflows.
Key Takeaways
- Claim: Fable 5.1 is positioned as a dual-mode model: capable of both long-running autonomous execution and rapid human-in-the-loop collaboration. | Evidence: The speakers contrast Fable 5, which they describe as best when given a large task and left to run, with 5.1, which they say is easier to converse with, faster, and suitable for iterative writing, coding, and interview-style workflows. | Implication: Ken should treat 5.1-class models as candidates for a single default agent substrate across interactive operators and background autonomous jobs, rather than maintaining separate models solely because one is easier to steer. | Caveat: This is a practitioner assessment from a team with early access and a commercial relationship to publishing model reviews, not an independently reproducible comparison.
- Claim: The model’s practical advantage in coding is its ability to make product and implementation judgments without repeated human correction. | Evidence: In a reported six-hour run using a compound-engineering workflow, it rebuilt a complex collaborative editor with accept/reject review, inline AI, sharing, activity views, human-versus-AI authorship indicators, subtle animations, and collaboration functionality; the speakers say previous models could get close but required days of iteration to reach comparable polish. | Implication: For OpenClaw-style build agents, the opportunity is to delegate a full vertical slice—including UI, workflows, and refinement—not just isolated code tickets; verification and deployment hardening must remain separate gates. | Caveat: The app was demonstrated visually rather than independently tested in the transcript; claims about code quality, reliability, and underlying collaboration architecture are not verified.
- Claim: Effort level is a material routing control: extra-high is best for autonomous polish, while medium can be faster, cheaper, and sometimes better for writing. | Evidence: Kieran says extra-high produces more complete and tastefully restrained coding results without indiscriminately adding features. Katie found extra-high writing generated a long preamble and more noticeable AI texture, while medium produced sharper candidate openings such as “By every reasonable measure, I should have laid off half my company by now.” | Implication: Build effort routing into agent orchestration: reserve maximum reasoning budgets for high-value autonomous builds or audits, and test medium/default settings for operator-facing drafting, synthesis, and rapid iteration. | Caveat: The vendor PM recommends high as the default in Claude Code and medium in other products for more than 90% of tasks; the optimal setting is task-specific rather than universally “maximum effort.”
- Claim: Fable 5.1 appears better at relevance selection—what the presenters call discernment—which matters more for knowledge work than raw summarization. | Evidence: The team says a meeting feed built with the model surfaced genuinely relevant 50/50 decisions for executive input rather than merely listing technically identifiable decisions. In an NPS dashboard, it highlighted that “interesting” was a warning word while “useful,” “great,” and “smart” correlated with promoters, and rendered an adjective ladder visualization. | Implication: Use the model to prototype executive signal feeds, decision escalation, customer-voice analysis, and organizational sensing—but instrument these systems with human feedback because relevance errors create trust failure faster than ordinary summarization errors. | Caveat: The reported meeting-feed quality is anecdotal, and the transcript does not establish recall, precision, or false-negative rates across a representative corpus.
- Claim: The model can generate usable knowledge-work artifacts such as decks and narrative analyses, but its default visual storytelling is not consistently best-in-class. | Evidence: A model-generated training deck reportedly followed the organization’s brand style, kept slides legible, used a useful compounding-workflow visualization, and correctly laid out directional arrows where GPT 5.6 failed. Conversely, its default NPS dashboard was described as functional but more generic and less narrative-driven than GPT 5.6’s output. | Implication: For deck and UI agents, supply a design system, example artifacts, narrative brief, and acceptance criteria. Do not assume that strong coding judgment automatically produces strong executive storytelling. | Caveat: The team explicitly says Fable 5.1 benefits from design guidance, existing Figma/design-system references, or explicit visual direction; one-shot default dashboard aesthetics were a noted weakness.
- Claim: For writing, Fable 5.1 improves at opening with a concrete tension and making human-like editorial choices, but it still has stylistic tells and does not displace competing models for all prose. | Evidence: On an essay introduction, Fable 5.1 led with “AI answers about 95% of my email. I have never been busier,” which the speakers preferred to Fable 5’s slower setup. They still identify occasional “X, not Y” constructions, excessive verbosity at high effort, and a literary or “hippie-dippy” sentence tendency. Katie remains more often in Codex, and the hosts prefer GPT-style models for clean, straightforward sentence-level writing. | Implication: Route writing by job: use Fable 5.1 for idea development, structure, high-level tension, and exploratory revision; compare it with cleaner prose-oriented alternatives for final copy where precision and restrained style matter. | Caveat: The favorable example is a single internal writing benchmark and reflects the reviewers’ house style; stylistic fit is inherently user- and context-dependent.
- Claim: The presenters believe Fable 5.1 can replace Opus 5 for many users because it delivers better long-context understanding and approximately similar effective cost through token efficiency. | Evidence: Multiple speakers say they felt a drop in quality when returning to Opus 5 on larger projects, and claim Fable 5.1 is about 50% more token-efficient than Opus 5, making its overall price approximately comparable despite potentially higher per-token pricing. Anthropic’s Alex Albert adds that some customers are reportedly moving from Opus 5 to Fable 5.1 at medium effort for lower-cost performance. | Implication: Run a controlled routing trial on representative agent tasks measuring total cost, latency, completion quality, intervention count, and retry rate—not just list price or subjective output quality. | Caveat: No pricing table, token traces, task-normalized cost measurements, or failure-rate data are provided in the transcript, so the cost and superiority claims should be validated on Ken’s own workload.
Detailed Brief
Ambition expansion: from app generation to simulation and computer use
- Claims: The speakers argue that model capability is ahead of typical user task selection: users may still be asking for tasks sized for prior model generations.; They frame the near-term unlock as combining coding, visual generation, tool discovery, spatial reasoning, and iterative verification within a single autonomous run.
- Evidence: An Anthropic PM describes giving the model a property-lot image and asking it to create an architectural specification, render a house, and produce a real-estate-style cinematic walkthrough; the model reportedly discovered and used Blender headlessly, captured intermediate screenshots, and rendered a video over several hours.; Mike reports one-hour and three-hour prototypes of a generative-agent town based on a research paper, a message-diffusion simulation, an internal-Slack simulation intended to model who would respond to a proposed message, and invention concepts with engineering diagrams and animations.; A one-shot Mac computer-use application called Hands reportedly accepted Slack tasks, operated local tools such as Codex, ChatGPT, Claude, and a browser, then returned results to Slack; previous Fable and GPT 5.6 attempts were said to be flaky.
- Caveats: The simulation examples are presented as toys or prototypes; no evidence supports their predictive validity for employee behavior, policy effects, or real message diffusion.; Computer-use systems interacting with local accounts, browsers, and developer tools substantially increase credential, authorization, data-exfiltration, and destructive-action risk.
- Implications: The relevant shift for agent design is task decomposition at the project level: specify an outcome and allow the agent to discover libraries, generate artifacts, verify outputs, and iterate.; Do not confuse impressive generated simulations with operational forecasts; require calibration datasets and holdout evaluation before using them for organizational or investment decisions.
Practical model-selection posture
- Claims: The team does not claim Fable 5.1 wins every workflow despite its enthusiastic overall verdict.; The vendor PM says model training involves behavioral and capability trade-offs rather than a formula that reliably produces an ideal model for every user.
- Evidence: Katie says she is returning to Claude/Fable after finding prior releases less useful, but still defaults more often to Codex and tests both when one produces an unsatisfactory output.; The hosts describe GPT 5.6-style models as cleaner and more direct for sentence-level writing, while describing Fable as more variable, creative, and useful for unpacking thinking.; Anthropic’s Alex Albert says model capability overhang often takes weeks to reveal: users discover a new generation’s real limits by attempting tasks they previously considered implausible.
- Caveats: The transcript repeatedly uses brand-like labels and subjective comparisons, but gives no standardized benchmark methodology, task corpus, or reproducibility package.
- Implications: Maintain a portfolio/routing mindset rather than declaring a permanent universal winner.; Create an internal frontier-task backlog—tasks previously ruled out as too complex—and periodically rerun it as models improve.
Notable Concepts & Terms
- Reach test: Every’s practical evaluation criterion: whether reviewers naturally switch to and repeatedly reach for a model in daily work, rather than relying only on formal benchmarks.
- Discernment: The model’s ability to identify what is genuinely important or decision-relevant, such as the central tension in a draft or an escalation-worthy decision in meeting data.
- Compound engineering: A workflow that gives an agent a broad product idea and lets it plan, build, review, and refine toward a finished implementation over a long run.
- Effort level: A model-control setting that trades speed and cost against reasoning depth and polish; the video treats it as a key operational routing variable.
- Token efficiency: Effective task cost depends on total tokens consumed, not just per-token price; the speakers argue 5.1 can offset higher pricing by needing fewer tokens.
- One-shot build: Generating a functioning application or artifact from a high-level prompt with little or no iterative human steering; the speakers present this as increasingly practical for prototypes and some internal tools.
- Generative-agent simulation: A simulated population of AI agents interacting in an environment; shown as a possible interface for exploring message spread, though not validated as a forecasting system.
- Computer use: An agent operating a real desktop, browser, and applications through UI interactions; presented through the Hands Slack-to-Mac prototype and relevant to local-workflow automation.
Operator Notes / Why Ken Should Care
- Run a head-to-head evaluation of Fable 5.1 against Ken’s current primary model on 10-20 representative tasks: long-running repo work, agent-tool orchestration, executive synthesis, deck creation, and production copy.
- Measure intervention count, retries, wall-clock time, token cost, and post-run defects; do not accept the claimed Opus-5 replacement economics without workload-specific data.
- Implement explicit effort routing in the control plane: medium/default for interactive drafts and routine transforms; high/extra-high only for bounded autonomous builds, complex debugging, or high-value artifact generation.
- Add a mandatory verification layer to autonomous coding runs: tests, dependency and secret scanning, permission boundaries, change review, and deployment isolation.
- Prototype an executive decision-feed agent over meeting and operating data, but collect feedback labels on relevance before trusting it for prioritization or escalation.
- Require design-system context, artifact examples, and narrative instructions for UI/dashboard/deck generation; use default output only for rough prototypes.
- Treat local computer-use agents as high-risk: use scoped credentials, allowlisted applications/actions, audit logs, human approval for consequential steps, and a hardened sandbox rather than unrestricted workstation control.
- Create a quarterly 'previously impossible' task suite to detect capability overhang and identify workflows that have crossed the threshold from assisted to delegated execution.
Source/Metadata
- Title: FABLE 5.1 IS A BEAST
- Transcript words: 14411
- Duration seconds: 4477
- Timestamp note: No timestamps or chapters were provided. The transcript contains substantial duplicated closing and discussion passages.
Transcript
Thank you. Thank you. I would say I think it's probably better on my computer. Can you just plug the camera in? Yeah, just plug the camera into my computer and the mic. Thank you. Thank you. Thank you. Thank you. Hello, everybody. We are here. We are figuring out some of our AV issues. So give us a sec. We've got Kieran Klassen here. We've got Katie Parrott with her every hat. We love the every hats here. If you're here, you should know. Fable 5.1 just dropped. We've been testing it for about the last week or so. And it is a sick model. We've got your day zero vibe check, as you've come to expect from us. We're going to tell you everything that we did with it. We put it through its paces. We did it on writing. We did it on coding. We did it on knowledge work. We have a full vibe check coming up on the every website, every.to. You should check it out, every.to. It's the only subscription you need to stay at the edge of AI. I want to get started. I want to go through our vibe check. I need two seconds to get myself set up so I can show you a couple things. So please bear with me. While I do that, Kieran, do you want to give your overall take on this model? Yeah. This is a great model. And for who? Because Fable 5 was a great model as well. Fable 5 was a great model for people that like to just give a bunch of work to it, and then it will just do it. It blew the ceiling of what it could do in depth, and for coding especially. This model is also very good to collaborate with. So it has that feel that Katie also called this sonnet feel. It's just fun to work with. It's nice to work with in a setting of in-the-loop work, which is great because it means I can use Fable for every kind of work. Because I do those two kinds of work. It's either I'm talking to the agent, iterating, or I let it run and rip forever. And it's really good at both. And we'll go into the details. But really, what the take is for me is this is Fable for everyone. This is not some elitist model that only the people on the edge will understand. And it's actually something everyone should use and try. So if you were a skeptic, or if you didn't really feel Fable 5, any of those folks, please try this one out. I agree. I think there's something interesting here. When we say Fable for everyone, there's a couple of different ways that that comes through. The first is, like Kieran said, it's friendlier and faster and more token-efficient. So the earlier Fable model was a beast, but it was so expensive. It was so verbose or spoke to you in technical gibberish that you really couldn't get all of the benefits of Fable unless you're highly technical and you want to send it off to go do a big, big project. And this feels like the first model I've used where we're actually pushing towards true delegated work for knowledge work, or for people who are maybe vibe-coding something but aren't really totally technical engineers. And that is really, really, really exciting. I think that there is this question for me about the anthropic strategy, which is we're going to make a genius in a data center, and then we're going to use that to make all of our other models. And I felt like one of the downsides of the genius in the data center is geniuses are hard to talk to. And they left some of their cloud identity behind, like the three fives or the four fives that we loved. And I feel like this is one where they're back. They figured out both having something that's an incredible coder, runs for days at a time, and is super friendly. Katie, Mike, anything you want to add? Katie, Mike, just echoing that, as someone who did find Fable 5 a little too intimidating, as one of those people who was not 100% sold on Fable, I just echo Kieran. This is a model that can really do it all. I ran it through my iterative writing process, the back-and-forth with it, and I understood what it was saying. It asked good questions in my interview workflow. And it just feels like a much more accessible model that can still do some great, incredible stuff. I'm actually seeing if I can get it to one-shot a project for me while we're on this call. So I've been on the hunt for a CMS for managing my Bible study. So I want a view that allows me to manage a Bible study hub. So we'll see. We'll see what Fable 5.1 can do for me while we're on this call. Sick, Mike. And I got to run. I'll be right back. I need to switch my browser. Mike, take it away. Yeah. So one of the things that all my friends said to me when I said, "Oh, you have to check out Fable 5. It's the best model," they were like, "No, it's not really the best model. I'm not getting good results from it because it's really hard to work with." And we've all hired that person who's super smart but really hard to talk to. And it felt a little bit like that. And I remember talking to one of my friends. He said the real sign of intelligence is if you're smart and also you find a way to relate to people, the Richard Feynman approach of you can bring it down to the level of a fifth grader. And I feel like that's actually what's happened here with Fable 5.1. It talks more like Richard Feynman than savant genius. Daniel, love this. I saw someone somewhere. They had this really insightful thought about Claude models where they said some philosopher, I can't remember who, talks about models of thinking and how when you think at high enough levels of abstraction, you become impossible to understand because you're just not relating to the real world anymore. And that felt like Fable 5, and Fable 5.1 seems to have fixed that. That definitely sounds right. And it's probably why I love, when I'm stressed, to read extremely esoteric philosophy because it disconnects me from the real world and makes me feel calm. Okay, so I want to give, we've been testing this for a while, I want to give people, we're going to go into specific tests that we ran and show people what we built and all that kind of stuff. But the first thing I want to start with is how does this change how I work. And we have actually a really interesting, we're using, let's just use me, for example, as a laboratory, which you can see when I got access to Fable 5.1, what happened to how I use AI. And what we found, what I found, is I was using, and there's another site that I need to go find, but I was using chat to BT a lot more even after I had access to Fable 5.1. But what you can see here is I'm using way more tokens on Fable 5.1. So I'm spending most of my day still in chat to BT. So if you're a chat to BT fan, I think that's probably what will continue to happen. But if you're a chat to BT fan, I think you'll find that you can delegate these big tasks to it and have it running in parallel successfully while you're doing chat to BT. And then, because of some of the tests we found, it's actually more token-efficient. It's about, even if it's a better price, even if it's a higher price than Fable, it's about the same overall cost to use it. And so I think this is one of those models where, if you're using Opus 5, you should just switch to Fable 5.1. It'll be approximately the same price, but much, much better quality and really fast and easy to talk to. Kieran. Yeah, in price, and it works way better. I've been switching back and forth a little bit between Opus 5, and especially we had early access and then we didn't have it for a little bit. And I went to Opus, and you just feel the dumbness come in. And this is especially across bigger projects with multiple things going on at the same time. Just that trust you can have in someone actually understanding everything is in Fable, and it's not in Opus 5. And Opus 5 is great at design and doing lots of work as well. And so I think this is one of those models where, if you're using Opus 5, you should just switch to Fable 5.1. It'll be approximately the same price, but much, much better quality and really fast and easy to talk to. Kieran. Yeah, yes, in price, and it works way better. I've been switching back and forth a little bit between Opus 5, and especially we had early access, and then we didn't have it for a little bit. And I went to Opus, and you just feel the dumbness come in. And this is especially across bigger projects with multiple things going on at the same time. Just that trust you can have in someone actually understanding everything is in Fable, and it's not in Opus 5. And Opus 5 is great at design and doing lots of work as well. But why not go for a model that gives you more, that costs the same? It's just more efficient with tokens. So yeah, if you're using Opus 5, go try Fable 5.1, because it feels very similar to work with, including it's just bigger and better. And yeah, I feel we haven't been in the saddle here in a while, guys. This is fun. It's fun to be back. It's fun. Is everyone hydrating? I got my water here. I also got coffee. Sorry. You got tea. We're cutting our diet coking. The fridge cigarette is, I think, an appropriate drink for Fable 5.1. Who's got some stuff they want to show? I know, Kieran, you've got some stuff that you built. Maybe you could show that. Mike, I would love, while Kieran's starting to queue that up, to go through some of the cases. If you want to show us, for example, the outputs for slide decks or the outputs for that blog post it wrote was really good, especially for Sofas 5. Yeah, I'm just drastically trying to hide the secret model names. Any way that we can do that would be great. I can show some writing stuff while both of them get organized, because nobody listened when I told them to check this beforehand. Yes, Kieran. Public shaming. Do you need some time, Kieran? I got them here. I can do one. I have them lined up. Do one, and then we'll pass it off to Katie, and Mike will seriously try to hide his code model names. And then we'll kick it on the forum. Kieran, go for it. Cool. Yeah, I shared my screen, if you can share it. So my favorite benchmark for Fable is rewrite Proof from scratch. Proof is this amazing collaborative editor that Dan created where it's a markdown editor where you can see provenance a little bit. And it's agent native, and it has all these things, and it's very complex. And I just said, hey, can you look at this idea? And I have a prompt that just kind of describes what it is. And I just gave this to Compant Engineering LFG, which is a workflow that just brings an idea to life until the end. And this is what it did. And it's really cool. It ran, I think, for six hours. And basically, it's a whole editor with accept, reject. It even has inline AI, which found a Google API key, and it integrated Gemini to do work for me. So you can ask AI about things, maybe less long. Let's see what it does here. Features. Well, I want to accept that. Boom. Boom. So it feels really polished. And then you can ask, why, what is new in 5.1? So my take on Fable 5, which was the best one before, and Fable 5.1, is things I noticed. It goes a little bit further in details. I can focus on the copy. There is a toggle here where I just hide this. The sidebar toggle didn't exist in the other version. In the other version, the activity didn't scroll with me. Also, what I noticed is the design looks a little bit better. So they improved design sensibility. Also, they tuned animations. I think before, animations would be a little bit over the top, appearing and all that. But here, they're with taste. So they have a purpose. Like here, for example, it kind of pushes the things closer to what it was. And I kind of like the clarity there. Also, if I refresh, I think there's very subtle animations here. So they tuned some in that as well. Obviously, I looked at the code as well. Looks very good. But Fable 5 was amazing as well. This is extra high. I think if you run these long tasks in extra high, you just get amazing results. High is very good. Medium is very good. But if you just want the best, go for extra high. Extra high just goes the furthest. And the beautiful thing is it knows when not to do things as well. So it won't add a lot of things. It will just make it better. So extra high means better results. It doesn't mean, like Opus 5, way more tokens and way more things that I don't need. The judgment is really tuned well in going for adding something versus not, which is really, really cool to see. Any questions on this, or let's go to the next one. I think let's go to the next one. Actually, before we do that, I think it's important to underscore how hard it is to make an app like this. Yes. Can you explain how you struggled with your version at some point and how you, yeah. At this point, this is ancient history. But at this point, I built the first version of Proof maybe, probably at this point six months ago. And we launched it, and it went down immediately, day one. And I was trying to fix it, and it was totally impossible to fix it because I completely vibe coded it. And at that point, the models were not strong enough to, A, vibe code it from scratch in a way that would work, and then, B, fix it if it was broken. Now they absolutely can. So comments work, track changes, live text updates, all that kind of stuff is kind of crazy. Actually, there's a lot of stuff in here, Kieran, that we should just put into Proof. I know, I know. That's the cool part. I have my own spinoff of Proof, which was Fable 5. One showed up that one, and I open sourced it. And I'm like, this is better than that version. So yeah, and there are really cool things because at another intro, I just do this, and boom. And that works. There's just so many things where it looks so smooth, which is very hard to do with software, software to look smooth and actually work. And there's CRDT in here. There's multi, you can see who is where. You can also click on review. So the whole idea is that you need to review it, and you need to mark it as reviewed, and you can even endorse it. So you can see levels of how much you endorse the text, and you can click here. So you go through what you have to do. You can see how many things are written by AI versus humans. You can see what AI is at the inline AI. You can share this with a link to a CLI. You can invite your own AI to collaborate on this. There's no web MCP in here, but I should add that to the font. But it's just insane what kind of thing you can build in one shot. And that's really the power move of Fable, doing it with taste and really nailing this. And that's the thing I want to underscore, is there are a lot of models that can build something like this end to end now. But one thing that Fable still, to me, is the best at is all these little details that it just does, that you look at and you're like, I might not have thought to tell you to do that, but I'm really glad you did. That's exactly what I wanted. And that's very hard to do. Yeah. And it costs a lot of time. So if you do this with other models, like Opus 5, you'll get very far. But to push it this far will cost you days of iteration. And just not having to be in the loop as a human to give that judgment and direction is very valuable. So I think that's really the power of this demo. So yeah, I can do other demos, or maybe someone else should do now, and then we'll come back to it. Yeah, let's do something. So I want to take a second and just say to everyone, if you're here, welcome. We are Every. But I'm really glad you did. That's exactly what I wanted. And that's very hard to do. Yeah. And it costs a lot of time. So if you do this with other models, an Opus 5, you'll get very far. But to push it this far will cost you days of iteration. And just not having to be in the loop as a human to give that judgment and direction is very valuable. So I think that's really the power of this demo. So, yeah, I can do other demos, or maybe someone else should do now, and then we'll come back to it. Yeah, let's do something. So I want to take a second and just say to everyone, if you're here, welcome. We are Every. We are the only subscription you need to stay at the edge of AI. When new models drop, we get them early, and we have day-zero vibe checks. We just published our vibe check from about a week of testing on Fable 5.1. Anthropic is so back again, baby. Let's go. So if you go to Every.to, Every.to slash vibe check slash Fable 5.1 vibe check. Also, you can just go to Every.to and click on it. We have this long-form vibe check with all of our hands-on testing. You can read it with ChatGBT. You can read it with Claude. It's pretty beautiful. We've got our reach test. So this is, for the people who review this model, we say, do we reach for it? And why? This model got two golds, which is very rare. We take you through coding. We already talked about proof here. We take you through writing. We take you through knowledge work. If you care about this model, if you care about learning what it's good for, what it's not good for, and where it should fit into your workflow, you need to go to Every. Subscribe and read this. It's free if you put in your email, Every.to. And welcome to our live stream. We're going to go through it in detail. Katie Parrott, give us some writing thoughts. Yeah. So we had Claude go through our writing benchmark, which, to recap, is a set of tasks that we have defined that reflect the kinds of things that we might ask a model to do. So that includes writing an intro to an essay from scratch, filling in a missing paragraph where your brain just farts and you don't know what you want to say, but you want it to take a pass on it. And then repurposing content from an essay into a LinkedIn post, into an email, into an X post. So I can share my screen and show you. Wait, wait, wait. I got to put it up. Okay. Is it? I don't see it. Did you press share? Share window. Let me. Okay. There we go. So here is a viewer that I built for myself just to be able to review all of the model outputs, just to make it legible to myself. So here you can see on the left, this is the intro to an actual Every piece that we published. It stands after automation mega post, which you should definitely seek out if you haven't read that. And here, model X is Fable 5.1. Hide your model labels, friends. And then Fable 5 is here on the right. And what I really like, I've noticed about Fable 5.1, is that it gets to the core tension that we're trying to establish much more quickly. So you see in the published version, there's a paradox at the heart of AI. We're immediately establishing a tension there, and this idea that we've automated as much as we can, and yet there's more work to do than ever. If you look at Fable 5, it goes off on a tangent about how crazy it is that Dan can do his email. It's just taking too long to get to the “but this created more work” part. Whereas Fable 5.1 just comes out with what I think, frankly, is a banger: “And AI answers about 95% of my email. I have never been busier.” So that's what I found with this model, is that it just makes really good, really human-like choices. And you may disagree with those choices, but they're not, it's not feeling as synthetic. Now I want to caveat that by saying this always happens with new models. And then over time, we discover what the new model's tells are. So I'm prepared to eat my words on this, and if it turns out that there's some new load-bearing not X but Y that we need to watch out for. But so far, I've found it to be pretty humanistic, still a little writerly or literary, as we've said. It loves a more hippie-dippy sentence structure. I don't know if that makes sense to me. I don't know if it makes sense to anyone. Makes sense to me. Yeah, totally. That's a scientific word. Yeah. Hippie-dippy bench, right? What are the vibes? And then the other thing I'll just say quickly. So that was Fable versus Fable 5 versus Fable 5.1. We can also look at the model comparing effort levels. And Kieran mentioned that when you kick it into X high, it really improves for coding. And that's when it can do its distance running. I actually found, with this model and past models, similarly to—oh, I need to get rid of the example. Well, anyway, I found that at X high, it gets a little in its own way. So here it has this really long preamble about all the stuff that we do. And that's just not as fast and interesting as what we got here with our banger: “I've never been busier.” So I do also like this medium line: “By every reasonable measure, I should have laid off half my company by now.” I think that's another really good way of launching the piece. So medium is worth checking out for Fable, I would definitely say, because you need to get those speed gains. I did find it gets a little bit more AI smell on it as the effort level changes. So that is writing weather report. I love that. I want to just summarize that a little bit. There's something interesting. It's really fun to actually watch it redo my writing and be like, oh, I like this or not. So the opening I did, “There's a paradox at the heart of AI.” What's interesting about that opening that I'm seeing now that I didn't see before is I'm telling, not showing. And what I like about this take from Fable 5.1, “And AI answers about 95% of my email. I've never been busier,” is it's showing. I think it could be slightly stronger if it added “and yet I have never been busier” to underscore the contrast, because I might not immediately make that connection in my head, and I want to connect one sentence to the next. But I do like this a lot. And I do think it's showing a move that AI has historically been hard to get the AI to do, which is figure out what's interesting and put it at the top. And the fact that this model is stronger at that, which I would put in the category of what I've been calling discernment, discerning whether or not it's relevant to the situation or my interest or anything like that, the fact that it's stronger at that is helpful for writing. And it's also helpful for a lot of other things. One of the things I did with this model that worked for the first time—I’ve got to see if I can bring it up in a demo. But what this model did is it created a feed of all the meetings that have happened at Every. And it created little summaries that were like little tweets that were pulling out for me, okay, here's a meeting where there was a 50-50 decision that it would be helpful for you to weigh in on. And normally the models would find all these decisions, and I'd be like, these are dumb. I can see why you think this is relevant, but it's actually not relevant. The fact that it's stronger at that is helpful for writing. And it's also helpful for a lot of other things. One of the things I did with this model that worked for the first time, I got to see if I can bring it up in a demo. But what it did, what this model did, is it created a feed of all the meetings that have happened at every. And it created little summaries that were little tweets that were pulling out for me. Okay. Here's a meeting where there was a 50-50 decision that it would be helpful for you to weigh in on. And normally the models would find all these decisions. And I'd be like, these are dumb. They're like, I can see why you think this is relevant, but it's actually not relevant. Your discernment is not good enough. And this is the first model that pulls out real things from meetings where I'm like, holy shit. I had this whole new lens into my company, and getting to a place where models have good enough discernment to do that, I think opens up all these different things. It's writing, but it's also, how do you manage your company? If you can surface good information to yourself that's relevant to you. It used to be that feeds were the domain of social media companies. And you had to have millions and millions and millions and millions of data points to make feeds. Now you can make your own feed in natural language, and it's getting pretty good. And I think it'll be way better in a year. And I think Fable 5.1 is the first one where I'm like, oh yeah, I would check this. So I want to keep the show rolling on. We've got Mike Taylor, our head of evals. Mike, do you want to show us, is there anything you want to show us? Yeah, yeah, I'll show you some stuff. Let me share my screen. Here we go. No. All right. Hopefully you can see that. One second, please hold. Yes, we can. Okay. Before you start, I just got to say, Mike built this thing to help us run our internal evals on this model. And what it does is it lets you see every single model and every single result from that model. So it's really easy to compare. And we've done this for a long time, professionally. And this just is such a pleasure to use, and it's just so cool. So I'm just really glad, Mike, that you're demoing this. Yeah. No, no. It's good to dust it off and get it out there. But yeah, these are a series of tasks that, at the start of the year, really mattered a lot to me when I was testing AI. And you can run all the 12 tasks that matter to me and get the inputs and outputs. And it's running each one in isolation. So you know that it's a fair test, essentially. So, for example, one of the things that I do a lot was creating dashboards. And you can see this is the dashboard it created. Now I want to share this one first because I would say that Fable 5.1 wasn't the gold for me because I think the default design isn't that great. I know, Kieran, you liked it for design. But to me, this is a pretty boring dashboard. I also have this other, this is something that I ask it to do, is create a Typeform competitor where you can talk instead of type. And it does the same standard vibe-coded stuff. The caveat is that this is the default design. I didn't give it any specific design chops. I didn't say, you should follow this design. So I think it's probably much better if you give it some more guidance on design, but that was the only downside I could find with this model. Mike, do you want to go, can you go to the write-up one for a second? Yeah. This is one of my favorite ones with Fable. So what's really interesting about this run, actually, do you want to describe the task first? Yeah. So the task, this is actually something I did in real life. So I had AI write this post. This was Opus 4.8 originally. And I was like, I created a banger tweet, and I was like, this needs to be a post. So I gave it a huge transcript of me just blabbing using monologue of all the things that I was thinking about when I wrote that tweet. And then I gave it a sample of my writing and said, this is how I write. And then the test is basically, does it do a good job of writing this up for me? Yeah. Yeah. And so this is a fully vibe-written post. And I just want to read some of it. So raise the ceiling, not the floor. So the nice thing about this is it's interesting. It pulled out the right interesting thing. The problem is it's X, not Y, and it's not specific enough to the post, but it's still decent. Seven things I learned about getting teams to actually use AI from my first few months running tech consulting and every, like perfect straight-ahead subhead. That's the kind of thing that we would hit publish on immediately. I think it did a good job here in the opening paragraph. Sam Par asked me a question on X, blah, blah, blah. Now that we've set up what this article is about, situating it in the right place. And then each section here, if you keep going, when a company decides to do AI, the first instinct is to evaluate tools. Someone assembles a spreadsheet of vendors of Claude or Codex or Gemini under the hood. And the team spends six weeks. This is the kind of thing where if I saw this, I'd be like, yeah, this is a reasonable way to open a paragraph. I want to show you, by comparison, go look at Claude Opus 5. Yeah. Okay. Nobody wants to use your AI tools. This is the same task. That's a crazy headline. It's very similar. It's very similar to what Opus 5 did in our benchmarks. It was like, we haven't fired anybody. It's like, why are you leading with that? Exactly. What are you doing, dude? And I think this is a really nice one because it's the exact kind of thing that people hated about Opus 5 that is clearly just fixed in Fable 5.1. And you can see that in other parts. If you scroll down a little bit on the, yeah, on Opus or back? No, no, no, no. Go back to the Opus 5 one. I want to keep, there's some really crazy shit in here. Wait, wait, wait, wait. Okay. Go down to number two. I don't know. There's one of these numbers where it's just completely going off the rails and talking about hell. And like, there's a special place in hell reserved for XYZ thing. And it was just. Yeah. You know. So Opus 5, as everyone knows, we gave pretty poor reviews too. And it's really nice to see that they figured out how to package that. And this is why I think Fable 5.1 is an Opus 5 killer. It's just, you get better quality, it's faster, and it's around the same price. Yeah. And where it didn't get a hundred percent, it was marginal. It was a little bit too verbose, but most of the models fail on that. I don't know. There's one of these numbers where it's just completely going off the rails and talking about hell. And there's a special place in hell reserved for XYZ thing. And it was just. Yeah. So Opus 5, as everyone knows, Antonia 5, we gave pretty poor reviews too. And it's really nice to see that they figured out how to package that. And this is why I think Fable 5.1 is an Opus 5 killer. It's just, you get better quality. It's faster. And it's around the same price. Yeah. And where it didn't get a hundred percent, it was marginal. It was a little bit too verbose, but most of the models fail on that. It had a couple of AI tells. We didn't automate the work where we used the new capacity to make the word an order of magnitude better. Right. It's the X not Y antithesis thing. So it still has a little bit of it, but this feels much more human. It's not obviously a Claude issue. Do you have the PowerPoint? I want to show the PowerPoint. Because there's a couple. Yeah. Yeah. That's a PowerPoint. That's a crowd pleaser. Yeah. So set the stage for us. Yeah. So this task is basically, here's a link to the repo for compound engineering that Kieran made, no pardon, but create a PowerPoint to teach this in a short training session. And this is something I really had to do at one point. And I used compound engineering a bunch, but I didn't know all of the ins and outs of the repo. So it was a task that's especially close to my heart. Because I was almost teaching myself at the same time. But one of the things I really like about this is we have a particularly specific brand style, and it finally really follows it. So it adds all these extra, the swirls and the background gradient and stuff, which previous models had failed at. And it doesn't try to shove too many different ideas in one slide. So you can see it's got good spacing. It's not hard to read. So I would say it's pretty close to something I would actually teach. And then the other thing I just wanted to point out is this is a slide that I immediately was like, oh, we need to steal this. Because I hadn't shown it this way before. I just showed it as bullet points. And I was like, oh, this visualizes it much better, with compounding and without compounding. So it even had a couple of flashes where I was like, oh, okay. Yeah, this is better than what I had. There are a couple of little things in here too that I think are really interesting. Did you show, do you know the one with the arrows pointing in a circle? I think it's toward the top. Just trying to find it. We're almost there. There it is. So, okay. Getting a model to automatically lay out, okay, we've got plan, we've got work, we've got review, we've got compound, with these arrows pointing the right direction. If you go to slide seven on GBT 5.6. Yeah, I want to show you what that looks like, or maybe it's not slide seven because it's not deterministic, but one of the slides is that. Yeah, this one. Oh yeah. Yeah, look at that. Yeah. It's trying to do the same thing and it just can't quite get it. Pointing arrows is apparently very, very difficult. Yeah. AGI is a pointing arrows problem. Forget Pelican on a bicycle. Okay. Making arrows in a PowerPoint is very, very difficult. And so this is one of those things where you look at it and it's a small detail, but if you really want a model to make a deck for you one shot, you need these little details to be right. Otherwise you're like, I just have to do it myself. Yeah. And Fable Club won't want it. This is why we've been talking about it as the first time you can truly delegate knowledge work. It's really much better on a lot of those dimensions. Yeah, for sure. Yeah. Yeah. Yeah. But if I say it's good at design, I mean it's good at doing design work and following what it needs to do so that it doesn't make mistakes. So it's super helpful in that way. And Claude, I've been using Claude design a lot with 5.1, and it's really good. You can also sync from Claude's code there now very easily. But yeah, is it the best looking design? Not always if you don't have the skills loaded or the direction, but at least if you say Figma to sync to your website or Claude design to your website, stuff like that, it really just works well, which is great. What I mean by the default design as well, this is the Fable 5.1 dashboard. And it's not wrong, and it has a lot of useful stuff in there, but it's a little bit less beautiful and a little bit less narrative-driven than, say, what GPT 5.6 sold it. So it came up with a better title, and it's picked out more interesting stuff. So I would say there's still a little bit of, that's one that we should talk about. Yeah. Let's talk about that. Yeah. So look at this. This is GPT 5.6. This is one thing. I think that Fable 5.1 can do this, but this is a task where I'm like, oh, it's interesting that it didn't choose to do it on this one. 5.6 was like a strong core clear promise, right? To set this up, we took our NPS survey from several years ago and threw it in the model. And this is a test we run all the time. And GPT 5.6 decided to tell a story, like a strong core clear promise. Customers see every as a rare mix of thoughtful AI media and useful products. Advocacy rises sharply when the mix translates into a simple outcome, practical help at the frontier without the noise. That's actually a very useful summary. That's not about the numbers. And if you go back to 5.1, it chose not to do that. It chose to just say straight ahead, every MPS dashboard, survey question, blah, blah, blah. Could Fable, if you asked it to do the task the way GPT 5.6 did, could it do it? Probably. But it's always interesting to see, on any given prompt, what the model decides to do on its own without being asked to. And what's also interesting about this is we've been praising its desire or its ability to get all the little details right without you asking. And this is one place where, for whatever reason, it's not as good. Yeah. And I think I have strong confidence that if I gave it a little bit of direction, it would do a good job, maybe even a better job, but I like to see what the defaults are. Yeah. And if you go down on this one, there were some things that it pulled out. blah, blah. Could Fable, if you asked it to do the task in the way that, to ask it to do the task like GPT 5.6, could it do it? Probably. But it's always interesting to see, on any given prompt, what the model decides to do on its own without being asked to. And what's also interesting about this is we've been praising its desire or its ability to get all the little details right without you asking. And this is one place where, for whatever reason, it's not as good. Yeah. And I think I have strong confidence that if I gave it a little bit of direction, it would do a good job, maybe even a better job, but I like to see what the defaults are. Yeah. And if you go down on this one, there were some things that it pulled out. This is, again, we're going back to what is its level of discernment? Can it pull things out that are interesting? And these are examples of people who can't describe every "don't recommend it." That's an interesting one. Oh, this is a good one. Scroll up slightly. Interesting is a warning word; useful, great, and smart predict promoters. So people who are just like, "Oh yeah, that's interesting," are not really likely to be our actual promoters. And the fact that it just pulled that out without us saying anything about what it needed to do is a really interesting insight that might change how we build the business. Yeah. A hundred percent. This is actually useful. And the other thing I wanted to point out that I really liked is it had this adjective ladder, where I guess this is where it gets that insight from, and I hadn't seen that design before for an MPS survey. And I thought that was pretty cool. It did some kind of keyword analysis, essentially, and then showed it in a visual way that makes a lot of sense to me. Yeah. That is really cool. All right. So what I want to do now, we've got a lot of people here. Now we've got 2287-ish and climbing rapidly. So if you're here, welcome. We are the only subscription you need to stay at the edge of AI. See, it says that right at the bottom, not lying right at the bottom. Every.to. But Fable, what's really cool about this version of Fable 5.1 is, what I want you to know, what the verdict is so far, you should watch our stream right now. But what you can also do is go to Every's website, which is right here, every.to. And we have a vibe check, a written vibe check, on Fable 5.1. The verdict: Anthropic is so back again. But Fable, what's really cool about this version of Fable 5.1 is, on a coding front, I would say slightly better than its predecessor Fable, which was definitely the best coding model in the world. But it is also now fast and friendly, like it speaks English, and that's a big, big, big achievement. So we are fans of this model. It runs for days at a time. Ooh, this is something I can demo for you in a second. It runs for days at a time doing things. And it happens to be about 50% more token efficient than Opus 5. It's also faster than Opus 5. So if you're using Opus 5, this should be your new daily driver. You should go read this vibe check. We're going to go through more of it on Every, but you should go check it out. It's totally free. You just have to put in your email. I've got everything from what Anthropic is saying to our reach test, where we talk about, I think the best way to tell if a model is any good is did we switch to it and are we reaching for it every day? If so, why, and who are we? This is our reach test for the model. We've got coding, which we've been going through. We've got writing. We've got, what else have we got? We've got knowledge work, which we just went through in detail, but it's really a faster and friendlier Fable, and I think that's a big achievement. What I want to talk about now on the coding front, I actually had it build me this thing called Hands, and Hands is kind of sick. I think actually Katie's also using Hands, which is pretty cool. And if, Katie, you have a take on this, let me know. But Hands is a computer use agent, a Mac app that runs on my computer, and I can send it tasks. So yesterday I asked it to create and share a history podcast at the history of Every. I did this from Slack, and the Slack message goes from my computer to Hands. It uses my computer, so it just types stuff into Codex. It's a little hard to demo a lot, but it types stuff into Codex or Chat GP or Claude or my browser. It gets the result, and then it sends it back to Slack. So I can basically use my computer fully from Slack. Computer use is one of those things where I'm like, oh, it's actually really, I think it's improving really fast, and I want to get good at it. And I'm interested in it, and it's nice to be able to use my computer from Slack. And this was a one-shot build from 5.1, and regular Fable and 5.6 both couldn't do it. They got close, but it was very flaky. And this is the first one that actually seems to work, which is pretty cool. And I'm using it all the time. Katie, it looked like you used Hands. Did you? I mean, I invoked Hands in our Slack, but I may not have actually had it set up, and the agent might have just told me that it was using it, and it lied. Oh, really? Yeah. You didn't download it? I didn't know it was something I had to download. I thought it was a skill I could just find. Oh, no. Holy hallucination, Batman. Okay. Well, anyway, you should try Hands. It's, I think, a pretty cool Mac app. We will probably release it as an experiment at some point, but for now it's just something we like to play with internally. For the people who have joined, Kieran, anything new you want to talk about or demo? I can share just a few quick things. So from the LFG bench, which is the compound engineering skill that runs everything from one prompt, I have a bunch of results here that I can go through and maybe say what is cool about it. So this is for long-running tasks mostly, but on the vibe coding side. Yeah. Yeah. So I really love this model for vibe coding as well, but the examples here are the results. So this is design. The thing I test here is, can it just do a design? This is kind of saturated, but what I look for here is, can I find anything that's an AI tell or a new AI tell or anything that's annoying or I wouldn't publish? I think we're at a point where it can one-shot landing pages easily. The other one is the 3D spatial. This is the cozy island, and you can see this looks pretty good. Other models look pretty good, but here, a new addition is reflections. So you can see reflections of the sun that I never saw before, which is really cool. You can see birds that stop flying once in a while, so they don't flap the whole time, which is also a new thing that I've never seen before. You can see it's very detailed. There's no mistakes here, I think. So yeah, I think this one is pretty. Yeah. There is always stuff you can add, which is kind of cool. And you can see here the butterflies are flying. And the fireflies. And the fireflies. Yeah. The other one is the 3D spatial. This is the cozy island, and you can see this looks pretty good. Other models look pretty good, but here the new addition is reflections. So you can see reflections of the sun that I never saw before, which is really cool. You can see birds that have, they stop flying once in a while. So they don't flap the whole time, which is also a new thing that I never, never saw before. You can see it's very detailed. There's no mistakes here, I think. So yeah, I think this one is pretty. Yeah. There is always stuff you can add, which is cool. And you can see here the butterflies are flying. Next one. And the fireflies. And the fireflies. Yeah. So I've seen those before, but it's just cool, the details. This is the breathwork garden things. So this is isometric. And yeah, you can just go to a desert thing, do a breathwork session. And then if you finish it, it will turn green and it will grow a flower. And if you do longer, you can do more flowers. Things here are cool. It's this butterfly here. I've never seen that it flies from flower to flower. So it goes from this flower, goes to there, and flies around. So there's a whole pathfinding mechanism for a butterfly flying to the flowers there, which is something new. And this is one of those details. I've never prompted it to do that, but it delights me. It's like, oh, that's smart. I should have done that. And Dan said with the proof thing and a writing thing, it's just fun to have a model where you see details that you're like, oh, that's delight. And it's very unique. Normally we are annoyed by all the mistakes it makes. It does not work. And we're now pushing ourselves in that we're being delighted by the results of the model, which is a very special thing. And I think it makes it a lot of fun also to do these things. You get more delight with more effort, I think. So extra high brings way more delight than high or medium, although medium is very good at vibe coding or just doing things in the loop. But yeah, that's the things from the LFG run. Amazing. I love the look of the breathwork garden. We have, and on that note, on the note of little details and how important they are, we have a very special guest coming on, Alex Albert from Anthropic. Hello. Hello. Alex is a research PM on this model. So you helped make this model what it is. Thank you for all your work. Tell us about the model. Tell us about your thoughts on it and the journey to build it. Yeah. Well, thanks for having me on, Dan and Karen, Mike, Katie. Great to see you guys again. This is a fun one for me. I've been very, very involved in this release, and it has been a process. It has been a big team, of course, that brought this model out into the world. And we're all very excited about it. I think this is best described as the friendly fable, as you put it, Dan, in that same sort of sense. This is the cherry on top to what fable five was back in the middle of the summer. I think there were just a bunch of little edges that we smoothed out with this release. That makes it something that's, one, on the model front, really, really easy to engage with and to just use as your daily driver. And then two, for broader adoption of this model, there's a ton of stuff that we really focused on here to make sure that this actually lands really well with all the folks out there that do want to roll this out to their enterprise or within their organization or just use it on their own. So a ton to love here. I would love to just chat about all this. It's great. That's awesome. Well, tell us a little bit about, I feel like there is this arc that happened with fable where it dropped and we're like, holy shit, this model's crazy. And then from our perspective, the son of five, the opus five launches, they felt like they inherited a little bit of fable's personality, which is a little bit hard to talk to. And a little bit, it just didn't have that Claude thing anymore. And this feels like it's back. And this feels like it's back. So can you tell us a little bit more about that arc and if there's anything you can share with what was going on behind the scenes and what you're trying to do? I would love to hear it. Yeah. I think model training is just not a perfect science yet, right? So whatever you do always carries some sort of trade-off, and we're constantly in this mixture of balancing one trade-off versus the other. And that's not to say that capability necessarily trades off entirely versus how the model feels or behavioral things, but there's no formula, right, that you can just plug in the right inputs and get this perfect solution that fits everybody's needs. I actually think opus five and sonnet five, for all their reported flaws, are still really great models, and they're daily drivers for many millions of people out there that use them all the time. And they're very, very performant as well. Opus five, if you just run that thing in the background, that will cook, and it will be really, really strong. I think there were a ton of learnings that we took from those models, and especially through feedback from just talking to a bunch of users and customers and folks like yourself, that really helped shape fable 5.1 into something that is great, kind of taking what was best from fable and what was best from opus and putting everything together. Which is really special to me. And I think this is actually now the start of a very continuing upward trajectory here of, okay, now we're cooking. This is going to get really fun, I think, with this model and future models. Amazing. Kieran, any questions? Yeah. So I love everything in this model, except one thing. I'm curious what your thoughts are on that, which is I still have to select an effort level, which is kind of annoying because it's so good at long-running short. I don't have to think about go for this for a short versus long-running, but I kind of have to think about the effort level. And I know at some point in the past you had this dynamic routing system set up, and I'm curious. That's kind of the only thing that's still annoying in the model to me, is that I have to decide whether I want extra high or medium, even though they're very well tuned. I still have to decide. Yeah. Oh, I appreciate you acknowledging they are well tuned. We've definitely put a ton of time into selecting the right defaults. So for most users out there, high, which is the default in Claude Code, and medium, which is on our other platforms or our other services, is what you should just stick with and use for 90-plus percent of your tasks. It just works. It's the best balance of intelligence while also still being efficient and fast. And it's what I just default to using all the time as well. I do think that the medium and the extra high, and even dropping down to low, do give you that spectrum of choice, which, Karen, you're a power user. You're wanting to tune the model to the perfect exact level for each task. And I think for you, there is probably a little bit of that, okay, do I have to switch down or do I move up? If you move down to medium or low, that really is the cost savings while preserving intelligence and speed option. So you're going to get fable five, better than fable five performance, at a cheaper cost, at a faster time. I mean, we're seeing customers even switch down from opus five to a fable medium, fable 5.1 medium, because of how token efficient it is at those low-cost points. All right, those low-effort points, which translates into lower cost. And then dialing up to extra high as well. you're a power user. You, you're wanting to tune the model, right? To the perfect exact level for each task. And I think, for you, there is probably a little bit of that, okay, do I have to switch down or do I move up? If you move down to medium low, that really is the cost savings while preserving intelligence and speed option. So you're going to get fable five better than fable five performance at a cheaper cost at a faster time. We're seeing customers even switch down opus five to a fable medium fable 5.1 medium because of how token efficient it is at those low cost points. All right. Those low effort points, which translates into lower cost. And then dialing up to extra high as well. It just gets you that you're going to be running this more autonomously. Maybe you want it to really put in all the sweat on the details and get something that's super polished. I think in terms of the future here, there's more that we're going to be exploring in terms of how we can make this even easier. I think from a user standpoint, and when I think about using the models myself, I want less knobs and dials that I have to think about every time I send a prompt. And I just want it to work for me. So we're very aligned on that approach, and it's something that we're going to keep pushing on. Great. I wanted to ask something, if you don't mind. Yeah. So Alex, one of the things I realized when I ran the benchmarks is my tasks are not ambitious enough to really show how good this model is. And so I had to actually spend most of the time really rethinking, what can I do? Where can I push it? And I actually don't even really feel like I got it to break on anything. How are you guys internally thinking about this? You see people on the team who've maybe had this for longer. Are they starting to reach for more ambitious tasks? Oh yeah. Yes. I think this is a funny thing, the capability overhang that exists with every single model generation. It usually takes a few weeks, a month even, for the model to fully get fleshed out in terms of what's possible. And the most striking example of this to me was Opus four or five back when it was released in November of last year. Then it wasn't really fully through winter break time towards January that people were like, oh wow. Okay. This is a really special model in autonomous background tasks. This finally clicks for a lot of things I've been wanting to do. I think every model has a little bit of that time period as well. So over the next week, two weeks, we'll start to see examples popping up. Internally, folks have just been starting to throw this model at everything. Everything, without fail. And I think this is the best way to get a sense of where the limits are. It's just really unbound your ambition here. Be like, why can't it do this task? Have I tried it before? Maybe I'm just artificially constraining it. So that's one thing. I think the second thing is each model has kind of a moment of some sort of task that now gets unlocked, and that usually kind of goes viral on Twitter or is really, really common and a bunch of people now start to try it. And with Opus five, one I remember is 3d games. A lot of people are suddenly making three JS games and really pushing the model hard in that sense. One thing we've been noticing internally, and I just actually posted something about this 20 minutes ago, is renderings, 3d models, spatial reasoning. In this example I had, I gave the model an image of a property lot. And I said, here's a lot. Can you design a full architectural spec for this and a house, render it, and then create a cinematography walkthrough, something that a real estate agent would put on their listing. Do you want to show it to us? Yeah, I don't know how I can pull up my thing here. Maybe if somebody can just pull up my Twitter, it's on there. Yeah. My new favorite benchmark test is to have every model try to build the set of the musical Hadestown. And I found that it's very good at showing the model. And I got a pretty good one from Fable.1. Yeah. That's a good one. We love it. Wait, wait, wait, we're going share screen. Okay. Share screen. Here we go. Can you see my stuff? You can't, not yet, but you will. Okay. There it is. I think this is what you're talking about. Yeah. Yeah. This is what I'm talking about. So I mean, this just blew my mind. I barely specified a prompt here for this, and it just figured out, oh, there's a blender library that I can use headlessly to run and render. And then it started taking screenshots, verify all the things as it was working. And eventually I come back to this a few hours later and I see this video that it pulled and ran on my CPU to render it efficiently for me. And I think this was something I didn't even know it was possible to do. I just didn't, I've never used blender before personally. I didn't know this was something you could do. Yeah. I agree. That's starting to happen, that specification. Blender. Blender is having a moment right now for some reason. Blender is definitely, I mean, yeah. And Linux. Linux and Blender. I just downloaded it today. I was like, I gotta, I guess I gotta get on the blender train. So, yeah, I mean that was just a library that the model found and used. But I think not even anchoring on that specifically, it's just that sort of task where it's like, wait, what if I did want to redesign my backyard? Can I just ask the model to do that? I think Mike has some demos like that that I want you to see. Because I think it'll please you, having worked on this model, what it is capable of. Mike, do you want to show us some stuff? Yeah, sure. Let me, so I was, I also kind of reached this. I was like, what can I do that I feel is too ambitious? So I really enjoyed the generative agents paper a few years ago and I've been thinking about it a lot. And I liked this idea of having a little AI town where they're just running around and talking to each other. And so I was like, here's the paper. Can you watch about this? And I came back an hour later and this was working. Thank you, concerned ape, for the aesthetics. Yeah. Yeah. And this is specifically from the paper. I think it's a Pokemon style aesthetic, which I also love. But this is pretty cool. And then I was like, you could use this to figure out how messages diffuse through a population. I was like, let's take this even more crazy. So it was like dropping new things. This is from the paper as well, but it was like, if we tell someone there's going to be a Valentine's Day party, how does that spread throughout the community? And I just think this is going to be a toy version of what you could do. But this is going to be pretty nuts for economists and policymakers. They could really AB test different messages and see what works. Yeah. That's amazing. It's like combining this analysis with this visual generation and all these different skills kind of mixed into one. I had it running in a loop to generate invention ideas of things that don't exist, but should exist for me. And it started coming up with some crazy ones, like a walking stick that automatically This is from the paper as well, but it was, it was, if we tell someone there's going to be a Valentine's Day party, how does that spread throughout the community? And I just think this is going to be, this is obviously a toy version of what you could do. But this is going to be pretty nuts for economists and policymakers. They could really AB test different messages and see what works. Yeah. That's amazing. It's combining this analysis with this visual generation and all these different skills mixed into one. I had it running in a loop to generate invention ideas of things that don't exist, but should exist for me. And it started coming up with some crazy ones, a walking stick that automatically picks itself up, a whiteboard that'll self-clean. And I can share some of these later on Twitter. Please do. But that's amazing. These were, these were fully spec'd out engineering diagrams, with arts lists and everything. And then animations of how it actually would work. Just insane stuff. I don't know engineering like that. I could not guide the model through that sort of thing, but it just inherently knew how to take this loosely defined spec into a working product. Yeah, that's brilliant. There was one other thing, because I thought, oh, it just did this without really thinking about it, so it could be even more ambitious. And so what I did is I got it to simulate our Slack. So it basically pulled everyone from our Slack, pulled their Slack history, generates, you can basically compose a message, paste it in here. And then you have a God mode for that message, like who responds? What do people think about it? What is Dan going to say? And so obviously, this again, you need to improve the accuracy of this. And this is probably, I think there's a similarly AI, they do this as a product, but I was blown away that it could. I think it took three hours to do this. And again, pretty amazing. And actually taught us a little bit about how information spreads within every, if you post in the everyone channel, it spreads much faster than if you spread it in wins or the consulting team. So I feel like there's a lot we're not even thinking about here. Yeah. So I have one question, because this is amazing, but I think the strongest thing of this model is that it's actually good at the smaller task. Have you seen internally people switching from Opus maybe to Fable more for the smaller in-the-loop tasks as well? Because yeah, Fable 5 could do pretty insane things already, and it's even more insane now, but day to day, I'm not going to do this all the time. It's really cool. But have you seen a shift in how and when people reach for Fable internally as well? Yeah. Yeah. I think that is, of course, the other end of this spectrum, is you have this super intelligent, capable model that can do these very crazy things. But at the same time, a lot of people have work that needs to be done, very practical work. And that's the focus on just making this model work out of the box for folks. So how does it really integrate well into our products like cloud code, across our services, into co-work, or anything else? Internally, this is a model that everyone is reaching for for these tasks. I find that ability for it to be token efficient and to be really targeted in edits that it's making in code, or whether it's writing a blog and revising something, those are tasks I use it for all the time. And I do feel that, as you guys were mentioning earlier on the stream, it just has a better, it's a hard thing to pin down, but it just feels better in my hands in that way. And that's something that you can't really capture in a benchmark or an eval. It really just comes out through a few days of usage. So that's what I'm super excited about for folks to start to pick up on as they're trying this model out. Amazing. Alex, I know you got a busy day. I appreciate you spending this time with us and sharing about the model, and thanks for your work on it. Of course. Thank you guys for using it. Appreciate it. Of course. Bye. Thanks. All right. You heard it from the man himself. A really, I think, a really interesting mix of coding ability, creative reasoning, taste, I don't know what you want to call it, all in one package. It's a really good model. Who's got a demo? We've just, we just got a bunch from Mike. Katie, anything else you want to share? Anything we haven't talked about on the writing front? Not on the writing front, but my live building attempt is still in the planning phase. So do you, it's being thorough. Can you show us the Hadestown rendering that it made? Yeah, sure. Let me pull it up. While you're doing that, basically, that's the thing now. Everyone's making 3D worlds. Kieran, I think, actually was the first person to do this a long time ago when it was not really even able to make a 3D world. And now it's like everyone's one-shotting 3D worlds. And so I had to think about what kind of 3D world I want to see. So I've been having the models do 3D reconstructions of the Battle of Waterloo, which I would totally show you, except my computer will die if I do that. I need to get a new computer. Basically, I'm constantly on video calls and people are like, oh, is your internet bad? And I'm like, no, it's a rogue agent that's trying to build something, and it is crashing my CPU. So you could go Linux next. I could go Linux. You're selling me hard on it. I think I honestly, I'm just waiting for the CPU RAM shortage to end so I can get a better MacBook, but we'll see. Oh, we have a good question. So be mask maker ask Katie, do you still prefer codex for writing and or knowledge work? And I have a good answer to this too. I'm curious what you think, Katie. I am so on the fence right now between being back in Claude and being inside of and staying with codex. Codex feels more familiar because it has been several weeks since I was in Claude code reliably because I wasn't finding so much utility in the previous models. And so I made the joke in our channel that going back into the Claude desktop app, I was like, Gandalf, I have no memory of this place. But I'm getting more familiar and I am having really good results for my style of writing, my voice, with Claude, with Fable 5.1. But Soul is still really strong. And so I'm hesitant. I'm back in. I don't want to choose territory. Because if I don't like it in one, I'll try it in the other. And sometimes I get a surprise that way. If we looked at your usage today, which app are you in more today? I'm still more in codex today. As I put in my reach test, I'm dipping my toes back in. I've been burned by a couple of models, but they're earning my trust back gradually. It's a process. I love it. Yeah. I'm also similar. I'm still in chat GBT most of the time. And so I'm hesitant. I'm back in, I don't want to choose territory. Because if I don't, if I don't like it in one, I'll try it in the other. And sometimes, I just, sometimes I get a surprise that way. If we looked at your usage today, which app are you in more today? I'm still more in codex today. As I put in my reach test, I'm dipping my toes back in. I've been burned by a couple of models, but they're earning my trust back gradually. It's a process. I love it. Yeah. I'm also, I'm similar. I'm still in chat GBT most of the time. But I've been such a huge chat GBT stand for the last six months that it's actually, I think, impressive that I'm now starting to be like, oh yeah, I'm flipping back into Claude every once in a while for the bigger tasks. And yeah, I agree. The thing I like about 5.6 and chat GBT models generally is I think their writing is just cleaner. It's clean, straight ahead. It just does the thing. And I think Claude is, and we've talked about this for a while, it has a little more variance, a little more creativity, a little more literariness. And I think that chunkiness of it means that there are sometimes where I'm like, a Claude model would be good for this. I do like a Claude model for helping me unpack my thinking. But for the actual writing of sentences, I find that chat GBT just does the clean thing well. And that's, for my writing style, the important thing. Yeah. Yeah. Inside compound writing, I've found that the model does a better job of picking up my style than Fable 5.1 does. I think it's also a more flexible writer. Although I personally do like Claude for my writing style. If I'm writing a section for a context window where it tends to be more repertorial, I find that I'm still reaching for soul. Cool. So we are reaching the end of our time together on this live stream. However, if you are interested in Fable 5.1, if you want to know more about it, if you want to hear all the things we covered and more, you should check out every, every.to. We have our full long form written vibe check on every. We also have a Fable 5.1 prompt library. So you can get the most out of Fable. And we also have this really cool graphic. Look at this. This is so cool. We have, we've been testing this model for the last week or so. You can read this with your agent. We have all these different sections. We'll tell you what Anthropik is saying. We'll tell you about what we're reaching for, the reach test, who we are, what we use it for, what we don't use it for, how it's changed our usage or not. We go through everything from coding, all the tests that we did, how it changed our usage in terms of coding, all the things we built. We went through writing, which we've been through on this stream, but where it fits in, how it performs versus Opus 5 and 5.6. We go through knowledge work. So everything from how it builds NPS dashboards to how it builds decks. And we also went through how it works inside of an agent. So if you're using it inside of an open claw or anything similar, you'll be able to learn how that works in Fable 5.1. And we have our verdict as well. You should check this out. If you're interested in AI, there's a ton more stuff on every. It's the only subscription you need to stay at the edge of AI. Look how pretty that is. What a good homepage. And there's more stuff coming. So pay attention. We've got a lot coming up in the next several weeks. Open AI dev day is in three weeks. So if you want to stay on the edge of AI, you should subscribe to every. We also, and this is actually, thank you, be mass maker, we have a camp on Friday, this Friday. We do live streams for subscribers only. We call them camps, and we're going to talk through our usage of Fable 5.1 in more detail. You'll get a chance to ask questions, meet the team, and hang out. So if you're here for the first time, welcome. We're every, we love to see you. If you're a returning person, we love you too. We are psyched to get to share these moments with you, and we hope you have an incredible Tuesday checking out Fable 5.1. And remember, stay hydrated, everybody. Cheers. All right. And it is like crashing my, my CPU. Um, so you could go, uh, Linux next. I could go, I could go Linux. I, you're, you're, you're selling me hard on it. Um, uh, I think I, I honestly, I'm just waiting for the, the CPU RAM shortage to end so I can get a better Mac book, but, um, we'll see. Um, oh, we have a good question. Um, so be mask maker ask Katie, do you still prefer codex for writing or, and or knowledge work? And I have a good answer to this too. I'm curious what you think, Katie. Um, I am so on the fence right now, um, between, between being back in Claude and being inside of, and staying with codex. Like I'm more like codex feels more familiar because it has been several weeks since I was in Claude code reliably because, you know, I wasn't finding so much utility in the previous models. Um, but, and so like, I made the joke in our channel that going back into the Claude desktop app, I was like, Gandalf, I have no memory of this place. Um, but I'm getting more familiar and I am, am having really good results for, um, from my style of writing my voice with Claude, um, with, with fable 5.1. Um, but, um, soul is still like soul is still really strong. And so I'm kind of, I'm kind of hesitant, like I'm back in, like, I don't want to choose territory. Um, cause like if I don't, if I, if I don't like it in one, I'll try it in the other. And sometimes, and I just, sometimes I get a surprise that way. Like if we looked at your usage today, which, which app are you in more today? I'm still more in codex today. Um, as I'm, as I put in my reach test, you know, I'm dipping my toes back in. I like, I, I've been burned by a couple of models, but they're earning my trust back gradually. It's a process. I love it. Yeah. I, I'm also, I'm similar. I'm still in chat GBT most of the time. Um, but I've been such a huge chat GBT stand for the last like six months that it's actually, I think impressive that I'm now starting to be like, oh yeah, I'm flipping back into Claude every once in a while for the bigger tasks. And yeah, I agree. The thing I like about, um, 5.6 and chat GBT models generally is I think their writing is just cleaner. It's like clean straight ahead. It just sort of like does the thing. Um, and I think Claude is, and we've talked about this. This for a while, like it's a little, has a little more, a little more variance, a little more creativity, a little more literariness. And I think that that chunkiness of it means that there's sometimes where I'm like a Claude model would be good for this. I do like a Claude model for helping me unpack my thinking. Um, but for the actual writing of sentences, I find that chat GBT just like does the clean thing well. And that's a, that's for my writing style, the important thing. Yeah. Yeah. Inside compound writing, um, I've found that, um, the model does a better job of picking up my style, um, than, than Fable 5.1 does. Like that's so like, I think it's also a more flexible writer. Although I personally do like Claude for my writing style. Um, if I'm writing a section for a context window where it tends to be more repertorial, um, I find that, um, that I'm still, still reaching for, for soul. Cool. So we are reaching the end of our time together on this live stream. However, if you are interested in Fable 5.1, if you want to know more about it, if you want to hear all the things we covered and more, you should check out every, every.to. We have our full long form written vibe check on every. We also have a Fable 5.1 prompt library. So you can get the most out of Fable. And we also have this really cool graphic. Look at this. This is so cool. Um, we have, uh, we've been testing this model for the last week or so. You can read this with your agent. You can, we have all these different sections we've got. We'll tell you what Anthropik is saying. We'll tell you about, uh, what, what we're reaching for the reach test, um, who we are, what we use it for, what we don't use it for, how it's changed our usage or not. We go through everything from coding, um, all the tests that we did, how it changed our usage in terms of coding, all the things we built. We went through writing, which we've, we've been through, uh, on this stream. Um, but, uh, where it fits in, it's, uh, uh, how it performs versus Opus 5 and 5.6. Um, we go through knowledge work. So everything from how it builds NPS dashboards to how it builds decks. And we also went through how it works inside of an agent. So if you're using it inside of an open claw or anything similar, you'll be able to learn how that, uh, how that works in Fable 5.1. And we have our verdict as well. You should check this out. If you're interested in AI, there's, uh, there's a ton more stuff on every, it's the only subscription you need to say at the edge of AI. Look how pretty that is. What a good homepage. Um, uh, and there's more stuff coming. So pay attention. We've got a lot. We've got a lot, uh, coming up in the next several weeks. Open AI dev day is in three weeks. So if you want to stay on the edge of AI, you should subscribe to every, we also, and this is actually, thank you, be mass maker. Uh, we have a camp on Friday, uh, on Friday, this Friday, we do, we do live streams for subscribers only, uh, we call them camps and we're going to talk through our usage of Fable 5.1 in more detail. You'll get a chance to ask questions, meet the team, um, and hang out. So if you're here for the first time, welcome. We're every, we love to see you. If you're a returning person, we love you too. We are, we are psyched to get to share these moments with you and we hope you have an incredible Tuesday checking out Fable 5.1 and remember stay hydrated, everybody. Cheers. All right.