We Tested Anthropic's Fable 5.1 for a Week
Description
Dan Shipper, CEO of Every, spent a week testing Anthropic's latest model: Fable 5.1. This is your day one vibe check. Read the full Fable 5.1 vibe check article https://every.to/vibe-check/fable-5-1-vibe-check?utm_source=youtube&utm_medium=video&utm_campaign=fable_5_1&utm_content=vibe_check_article Fable 5.1 prompt library https://every.to/claude-fable-5-prompt-library?utm_source=youtube&utm_medium=video&utm_campaign=fable_5_1&utm_content=prompt_library Follow Dan Shipper: https://x.com/danshipper Follow Every: https://x.com/every Timestamps: 00:00 Intro 01:42 Coding 06:26 Knowledge work 12:45 Writing 19:52 Closing thoughts
Summary
Generated by gpt-5.6-terraAt-a-Glance
- Verdict: Watch fully
- Core thesis: Based on Every’s week of testing, Anthropic’s Fable 5.1 is a materially more usable, faster, and more token-efficient long-horizon agent model that can delegate substantial coding and knowledge-work projects end to end, while remaining a useful complement rather than a universal replacement for GPT 5.6.
- Why it matters: The video offers a concrete operating pattern for AI work: use an interactive model for rapid daily iteration, and send bounded but substantial projects to Fable 5.1 to execute autonomously over long runs.
- Best use: Watch for the demonstrated agent-app build, the cost/latency benchmark, and the practical model-routing framework for coding, decks, analysis, writing, and executive-information triage.
Executive Summary
Dan Shipper argues that Fable 5.1 is a rare improvement in both model capability and usability. In Every’s internal agent benchmark, he says it averaged 766 tokens per request and 22 seconds of latency, versus roughly 2,000 tokens and 37 seconds for Opus 5 on the same internal-work tasks. His core conclusion is that the model brings Fable-class autonomous delegation into a cost and speed range that ordinary knowledge workers can plausibly use.
The strongest evidence is an end-to-end build of “Hands,” a Mac-based computer-use agent accessible remotely from Slack. Shipper says Fable 5.1 built the application from a few prompts after running for a day on an “ultra code” setting with 40 sub-agents. He treats this as evidence that the model can own large, multi-component engineering projects rather than merely assist interactively.
For knowledge work, the model generated an NPS dashboard with useful qualitative and quantitative insights, and turned a written document into a visually coherent slide deck with unusually competent layout details. Shipper’s broader claim is that knowledge workers can now delegate first-pass deliverables—analysis, presentations, and selected monitoring/triage tasks—in the way developers have increasingly delegated larger coding blocks.
He does not claim Fable 5.1 replaces GPT 5.6 everywhere. GPT 5.6 produced a stronger top-level narrative in the NPS-dashboard comparison, and Shipper still prefers it for much of his own interactive daily writing because Fable’s prose can remain literary and somewhat chunky. His actual workflow is two-speed: ChatGPT/Codex for conversational iteration, Fable 5.1 for large tasks that can be dispatched and allowed to run.
Key Takeaways
- Claim: Fable 5.1’s main practical advantage is making long-horizon autonomous work cheaper and faster enough to be regularly usable. | Evidence: On Every’s internal agent benchmark, Shipper reports average consumption of about 766 tokens per run and 22 seconds latency, compared with almost 2,000 tokens and 37 seconds for Opus 5 on the same tasks. | Implication: Ken should evaluate it where responsiveness and token burn have made agentic workflows uneconomic, particularly internal operations that need repeated multi-step execution. | Caveat: The figures come from Every’s internal benchmark and task mix; they are directional rather than an independently validated general performance comparison.
- Claim: Fable 5.1 can execute complex software projects end to end rather than requiring continuous interactive coding supervision. | Evidence: Shipper says it produced “Hands,” a Mac computer-use agent that exposes remote computer control to other agents through Slack, including screenshots, task status, and logs, from a few prompts and a one-day ultra-code run using 40 sub-agents. | Implication: Use it for explicitly scoped, high-value builds with acceptance criteria, budget caps, and a review stage—not as a default substitute for interactive pair-programming. | Caveat: The demonstrated project reportedly consumed approximately 3–5 million tokens, and ultra-code mode is intentionally non-conversational: the model largely disappears to work and returns later.
- Claim: The model appears especially strong at converting source material into decision-ready knowledge-work artifacts. | Evidence: On customer-survey analysis, it generated a static HTML NPS dashboard and surfaced an actionable language signal: “interesting” correlated with weaker customer sentiment, while “useful,” “great,” and “smart” predicted promoters. It also converted a document on compound engineering into a deck with one primary idea per slide and correctly rendered visual loop arrows that competing output reportedly mishandled. | Implication: Fable is a strong candidate for delegated first drafts of analysis and presentations, but executive-facing outputs should still receive a narrative and factual review before distribution. | Caveat: In the NPS comparison, Shipper preferred GPT 5.6’s more narrative top-level framing, so Fable’s outputs are not uniformly superior across deliverable types.
- Claim: A differentiating capability is discernment: identifying what is strategically interesting and explaining why it matters, rather than merely summarizing or agreeing with any analogy. | Evidence: In a meeting-monitoring example, it flagged a marketing launch-targeting discussion as a strategic tie-breaker requiring Shipper’s involvement. He contrasts this with the “McDonald’s eval,” where models flatter an arbitrary comparison instead of judging whether the connection is genuinely meaningful. | Implication: For agent-control-plane and executive-assistant workflows, test Fable on escalation selection: whether it can reliably distinguish items that need operator judgment from routine updates. | Caveat: This is based on qualitative examples and the presenter’s judgment, not a disclosed benchmark for strategic relevance or false-positive rates.
- Claim: Fable 5.1 substantially restores Anthropic-style writing quality relative to Opus 5, but it is not an automatic replacement for GPT 5.6. | Evidence: Given a brief prompt and a transcript from Every’s head of evals, it produced a structured post on AI adoption that Shipper judged clear and non-generic. He says its measured reading ease was highest and grade level lowest among Fable 5.1, Opus 5, and GPT 5.6; he also shows the model diagnosing a weak paragraph’s lost payoff and mismatched Help Center-like voice. | Implication: Route writing by task: use Fable for structural editorial critique, alternate creative perspective, and one-shot drafts; retain the preferred interactive writing model for voice-sensitive production work. | Caveat: Shipper still uses GPT 5.6 for most day-to-day writing because he prefers its more minimal, direct prose; Fable retains some “literary” and “chunky” tendencies.
- Claim: The emerging productive workflow is a two-gear model stack, not single-model standardization. | Evidence: Shipper’s usage after receiving Fable 5.1 remained mostly interactive in ChatGPT/Codex by prompt count, while Claude/Fable token use rose sharply because he assigned it large builds such as Hands and personal-feed experiments and let it run independently. | Implication: Ken should design routing around interaction mode and horizon: short feedback loops go to the conversational model; expensive autonomous runs go to the delegation model with observability, checkpoints, and spending controls.
Detailed Brief
Operational design pattern: dispatch versus collaborate
- Claims: Fable 5.1 changes the interaction model more than merely improving output quality: its high-effort mode is designed for dispatching a substantial task, not conversing continuously while it works.; The presenter sees this as a potential expansion of delegation beyond developers into ordinary knowledge work, especially now that first-pass decks and analyses are credible enough to review rather than rebuild.
- Evidence: Shipper describes ultra-code behavior as the model going away to “cook” and returning later, rather than engaging in a live back-and-forth.; His own post-release usage pattern shows more Fable token consumption without a corresponding increase in Fable interaction frequency.
- Caveats: Autonomy shifts failure detection later in the workflow; a result that looks complete still needs validation for requirements, correctness, security, and presentation.; The transcript contains no systematic reliability rate, security analysis, or comparison of failure recovery across models.
- Implications: Treat autonomous execution as a production workflow with explicit task contracts, budget limits, artifact review, and stop conditions.; A useful evaluation should measure not only final quality but also how often the agent escalates appropriately, stays within spend, and leaves a comprehensible audit trail.
Limits of the review and comparison claims
- Claims: The video is a hands-on review built from Every’s internal tests and examples, not a neutral or comprehensive benchmark study.; The presenter repeatedly favors Fable 5.1 over recent Anthropic models but preserves GPT 5.6 as a preferred tool in several interactive and narrative tasks.
- Evidence: Every’s stated test set includes coding, data analysis, presentations, and writing, with an internal agent benchmark used for token and latency figures.; The source is associated with a paid Every subscription and ends by directing viewers to Every’s written review and prompt library.
- Caveats: Several product labels and model comparisons are presented without methodology, raw benchmark data, task distribution, pricing details, or reproducible prompts in the transcript.; The transcript includes duplicated sections, so its practical value is concentrated in the initial app-build, benchmark, knowledge-work, and workflow-routing portions.
- Implications: Use the review as a source of test hypotheses and workflow patterns, not as a procurement conclusion.; Run a private bake-off on Ken’s own agent tasks before moving workloads or relying on the presenter’s broad “no reason to use Opus” framing.
Notable Concepts & Terms
- Fable 5.1: The reviewed model, characterized as combining Fable-level long-horizon delegation with lower latency and token use.
- Hands: Every’s demonstrated computer-use application: a Mac-resident agent interface that lets other agents remotely operate the computer and exposes screenshots, status, and logs.
- Ultra code: A high-effort operating mode used for the Hands build; it emphasizes prolonged autonomous execution over conversational interaction.
- Two-gear workflow: Shipper’s proposed operating model: interactive GPT/Codex use for day-to-day work alongside Fable for large delegated jobs.
- Discernment: The model’s claimed ability to identify whether a signal, analogy, or issue is genuinely meaningful rather than merely plausible-sounding.
- McDonald’s eval: An informal test of sycophancy and relevance: whether a model agrees that an arbitrary comparison is apt instead of assessing the connection critically.
- NPS dashboard: A test artifact combining Net Promoter Score analysis with qualitative survey interpretation; used to assess analytic insight and communication quality.
Operator Notes / Why Ken Should Care
- Run a controlled bake-off on 3–5 real long-horizon tasks: a repo-level feature, an agent workflow, a source-to-deck task, and an executive-triage task. Track completion quality, human repair time, latency, token cost, and unsafe actions.
- Create a routing policy that explicitly separates interactive work from dispatched work; do not judge models solely by chat quality when the intended value is autonomous execution.
- For any computer-use implementation, require least-privilege access, task-level approval gates for consequential actions, immutable logs, screenshots or state captures, and a kill switch before enabling remote agent control.
- Set per-run token and elapsed-time budgets for ultra-code-style jobs, with intermediate checkpoints for projects that can otherwise consume millions of tokens.
- Test relevance calibration before deploying meeting or communications triage: measure false escalations, missed decisions, evidence citation quality, and whether the model can state uncertainty.
Source/Metadata
- Title: We Tested Anthropic's Fable 5.1 for a Week
- Transcript words: 6203
- Duration seconds: 1245
- Timestamp note: No timestamps or chapters were present in the supplied transcript; the latter portion contains repeated material.
Transcript
It's model release day! We've got an all-new Fable, Fable 5.1, and it is Fable for everyone. We've been testing it for about a week. This is a sick model. It is a better coding model than the original Fable, so if you love the original Fable, you're going to love this model even more. And it speaks English. It actually talks to you like a normal person. It says things that you can understand, and that's not all. It is significantly faster. It's twice as fast as old Fable, and it uses half the tokens as Opus 5 for similar tasks. So it's priced like a Fable, but you'll be able to use it for all your Opus tasks, and there's really no reason for you to use Opus anymore. Fable all day, baby. So let's get into the vibe check. Before we get into this video, who am I, and how do I possibly have a full review of a model that just dropped? My name is Dan Shipper, and I'm the co-founder and CEO of Every. Every is the only subscription you need to stay at the edge of AI. We get all of the models before they come out, and we test them with a team of writers, designers, engineers, and more to tell you what they're good for. So on the day a model drops, we have long-form reviews to tell you what to use them for and what not to use them for from our real experience doing real work. We've been testing this model on all the kinds of things that you do every day. We've done it on coding, we've done it on knowledge work, like decoration and data analysis, and we've done it on writing. And I'm going to show you everything that we found in our usage and how that has pulled through into how I use it day to day. What are the actual numbers for how it has changed my usage over the week since I got access to it? Let's get started. One of the coolest things I had to build is this thing called hands. It is a computer use agent that sits on my computer and allows other agents to use my computer remotely. So that's the arrow for hands. That's hands going and using my computer. And you can see it has this little window on my computer that shows me what hands is doing. And what this does is it lets me remotely use my computer from any agent I want. And this is a hard thing to build. It's hard to make a Mac app that is a computer use agent that is accessible from the web that can do any task you ask it to do. And I tried to build hands in 5.6 in regular Fable. And it is good, but it didn't really work. And this actually just works. It just uses my computer, and I can have it use my computer from Slack. And that is fucking wild. Look at that. Look at that. That is fucking wild. And look, it's going to start typing too. And you can see it's working in ChatGPT. You can see it's typing, it's typing. And I have no idea how this works. This was built end to end by Fable 5.1 from a couple of prompts. That is a desktop agent built fully by this model. So this is just a fully fledged app. It has a screenshot of the current state. It says what it's doing. It has a log. Look, look, these are all the tasks I've been doing with this thing. This would ordinarily take months to be good. And it just worked after a day. I gave a prompt to Fable 5.1 and it just churned on ultra code with 40 sub-agents for a day. And then this came out, and it just works. I don't know how else to say it. It was a completely different experience of programming to be able to give a simple prompt to a model and have a completely usable, complex application come out a day later. That said, there are some things to be aware of. It's more token efficient than Opus 5, but if you give it a big task like this on ultra code, it's going to run for a day and it's going to use, I mean, this is like a three- to 5-million-token run. So it can get expensive. And especially on ultra code, it's not really that conversational. It is conversational on lower effort settings. But if you put it on ultra code, it's basically just going to go off and cook and come back later. So that's just a different usage pattern than you might be used to. It's better for big projects, especially not as interactive coding work, but fucking sick. And the other technical people at Every, like Kieran Klassen, who's the GM of Quora and makes compound engineering, love this model. I think for Kieran, he's been a Claude stan for a long time. And even he didn't like Opus 5, didn't like Sonnet 5. And this model, he's like, this is my favorite again. So it's a big jump in both power and usability, which usually don't go hand in hand, and speed. It's just fast enough to use for day-to-day tasks. To give you a little bit of a sense for its level of efficiency and how fast it is, we ran it on our internal agent benchmark. This measured it on tasks that our internal agent does, like work inside of a company. And we compared it to Opus 5. So in terms of the number of tokens per run, on average, this model consumed about 766 tokens for each run. So each request to the model is about 766 tokens. By comparison, Opus 5 was almost 2,000 tokens for the same tasks. Similarly, the average latency for this model was about 22 seconds. So that's how long it took for a basic request to get a response, 22 seconds. Opus was about 37 seconds. So it's about twice as fast as Opus, and it uses about half the tokens. That's really crazy, and it's a very rare jump. There was this moment in December 2025 where Opus 4.5 and GPT 5.3-ish crossed this barrier where suddenly all these vibe coding use cases opened up, and all these people discovered that you could code in this new way with these models. I think there's a similar thing happening with Fable 5.1 because it opens up these long horizon tasks in an easy-to-use form factor that's cheap enough for regular people to try. I think it has the chance to open up real delegation of coding and knowledge work tasks to more average knowledge workers. And that's a big deal. So now let's get into knowledge work. We have this really cool eval viewer so I can show you exactly what we ran it on, what the prompts were, and what the outputs were. The first knowledge work task we always do with this model is we run it on data analysis. So we feed it a survey that we gave, a real survey that we gave to our customers a couple of years ago, and we ask it to analyze the survey. So the prompt was, I want you to build a static HTML NPS dashboard, and NPS is Net Promoter Score, and pick out interesting insights after doing the analysis. So this is what we got out of the model. So it's very clean design. It makes it really easy to see the score distribution and all that kind of stuff, which is, I would say, pretty normal. It's not obviously AIE right now, so it doesn't feel like slap. One thing that it did that I think is a little bit worse than, for example, GPT 5.6 is, and we can actually go look at GPT 5.6 for comparison, GPT 5.6. This is GPT 5.6's version. It pulled out a story at the top: a strong core, a clear promise, customer see every is a rare mix of thoughtful. This is something that Fable 5.1 didn't really do. It was much more straight ahead. So I really like GPT 5.6 on this task. But one thing that I want to point out from 5.1 that's really interesting is it did an incredibly good job at pulling out real insights from both the quantitative and qualitative data and then presenting it to us. So like what people love, the writing is the core of the love, the apps are the bonus. Keeps me current without overwhelming me is the strongest value proposition. Interesting is a warning word, useful, great, and smart predict promoter. So this is really cool, right? So it went and did the analysis to understand that when one of our customers calls us customer see every is a rare mix of thoughtful. This is something that Fable 5.1 didn't really do. It was much more straight ahead. So I really, I actually really like GPT 5.6 on this task, but one thing that I want to point out from 5.1 that's really interesting is it did an incredibly good job at pulling out real insights from both the quantitative and qualitative data and then presenting it to us. So what people love, the writing is the core of the love, the apps are the bonus. Keeps me current without overwhelming me is the strongest value proposition. Interesting is a warning word, useful, great, and smart predict promoter. So this is really cool, right? So it went and did the analysis to understand that when one of our customers calls us interesting, they actually don't really like us that much. But if someone calls us useful, great, or smart, they do love us. And then it was able to package it, not only find the insight, but package it into a little sentence that could express it to me in a way that I didn't feel was slop. This is something that old Fable wouldn't do. Opus five absolutely would not do. And generally, models struggle with. So there's a leap here in how it is able to pick out interesting things and then present them to you. Another thing that we always do is we have these models create a PowerPoint. So here's the PowerPoint it created. You notice some small visual things like this. Every logo is not quite right, but in general, the prompt we give it is just take compounding engineering and turn it into a deck. It created this deck end to end. This is a very well put together deck on a couple of different levels. For one, it does a very good job of pulling out specific ideas and making sure each slide has one idea on it. It does a really good job of these subtle visual design things, like the yellow for the compound engineering and the end of that paragraph. And it even does things like, you'll see it placed these arrows to show the compound engineering loop. It created, it does these bubbles with these arrows. Most models can't get these arrows right. There are these subtle little things that it does really, really well. So it is the kind of model where if you give it a well-written doc and ask it to turn it into a deck, it'll do a really good first pass end to end without very much interference, which is something that most models can't do. So for example, this is what 5.6's deck looks like. And you can tell it's just a little bit worse. It's a little bit more bare. You can see the arrows here. The arrows are not done right. So the thing I want you to take away is developers for at least the last six months, depending on who you are, have been getting used to delegating big chunks of work to the model. That hasn't really been possible for knowledge workers. You can do it, but the details aren't really right. It looks a little bit like slop. This is a big step in the direction of actually, I can have it make a deck for me end to end. Actually, I can have it do a bunch of data analysis and get something back, and I'm like, this is actually pretty good. And to be clear, the original Fable did some of this, but not at the speed and not at this cost. So it makes it available to more and more people that previously would have looked at Fable and been like, this is not for me. For example, this is a meeting where I can be a tiebreaker. Our head of marketing is talking about a new launch and who we're targeting for that launch and whether we might want to target other customers as well. And it's flagging me as, Dan, there's a question about strategy here. You might want to be involved in this. You might want to know about this. And this is a superpower for me that Fable 5.1 can do, is pick out interesting things that you're going to care about from what's happening in your company, what's happening in the world, and tell you about it in a way that clicks for you. This is new. Normally it would show you stuff and you'd be like, ah, this doesn't really click for me. It's not actually interesting. This is the first model that I've ever found that can pull out something interesting and then present it to you in a way that makes it click. I think the thing that Fable can do, well, one, it just built this app totally end to end. But the thing that it can do is if Fable thinks something is interesting, it's more likely to actually be interesting. There's this eval that's loosely called the McDonald's eval, where if you're talking to your model and you're like, oh, that's exactly like McDonald's. And it's like, oh yeah, you're totally right. But you're talking about something completely random. That's an example of the kind of thing that models struggle with. I would call it, to some degree, people pleasing. To another, it's discernment. When you make a connection, is that connection actually interesting? Is it a real connection, or is it just that you can connect anything to anything, which these models can do because they know everything? And I think Fable's level of discernment is much higher than previous models I've tried. And that takes me to the next one, which is writing. And obviously writing is, that's what I love. So Anthropic has really struggled with writing recently. They used to be the top writing model, but with Sonar 5 and Opus 5, pretty much everyone was just like, this model doesn't speak English correctly. It doesn't give me good output anymore. And everyone internally at Every had switched to ChatGPT. With Fable 5.1, it's not only better than Opus at writing, it's actually starting to compete with ChatGPT. It writes quickly. It has pretty clear prose. It still has a lot of the Claude literariness. But if you look at the reading level that it writes at and the reading ease scores, it has the highest reading ease and the lowest grade level of Opus 5 and GBT 5.6 Sol. So it is just producing stuff that's easier to read. So as an example of the kind of one-shot writing this thing does, we asked it to create a blog post based on a simple prompt and a long transcript of Mike Taylor, who's our head of evals, just talking about some of the work that he does. And we asked it to write some of the things that he's learned about getting teams to use AI. And this is the post that it created just from a very, very simple prompt. You can see the prompt is really small. 3. Raise the ceiling, not the floor. 7 things I've learned about getting teams to actually use AI for my first few months running tech consulting at Every. Sam Parr asked me a question on an X last week that I've been asked in one form or another in almost every client meeting since I joined Every. How's everyone getting team adoption for Claude? It's the right question. The gap between what the terminally online people in a company can do with AI and what everyone else does with it has never been wider. This is pretty good prose. It doesn't feel like slop. Buy the model. Buy the model direct, not third-party tools. When a company decides to do AI, the first instinct is to evaluate tools. Someone assembles a spreadsheet of vendors, the Clutter Codex or Gemini under the hood. It doesn't have any of the X, not Ys. It has the raise the ceiling, not the floor, but still, it doesn't have that slop feeling. It's well-structured. It's thoughtful. It's actually really impressive. Let's look back at Opus, Opus 5. Same prompt, same exercise, same everything. Nobody wants to use your AI tools. Opus 5 was such a dick. Nobody wants to use your AI tools. Some of the things I've learned about AI adoption since joining Every Consulting. Sam Parr asked a question last week that I think is the real question in AI right now. That opener, it's a similar opener. It's a little bit more of a minimal sentence construction, which isn't necessarily bad, but you can just feel, let's see. The instinct is to go after laggards. We bought you the tools. Now everyone needs to use AI. This does not work. On pain of death, a good number of people are emotionally unwilling to use AI. What is this thing even talking about? Opus 5 is just so dark, and they fixed it with Fable 5.1. It's really impressive. Let's look back at Opus, Opus 5. Same prompt, same exercise, same everything. Nobody wants to use your AI tools. Opus 5 was such a dick. Nobody wants to use your AI tools. Some of the things I've learned about AI adoption since joining Every Consulting. Sam Parr asked a question last week that I think is the real question in AI right now. That opener, it's a similar opener. It's a little bit more of a minimal sentence construction, which isn't necessarily bad, but you can just feel, let's see. The instinct is to go after laggards. We bought you the tools. Now everyone needs to use AI. This does not work. On pain of death, a good number of people are emotionally unwilling to use AI. What is this thing even talking about? Opus 5 is just so dark, and they fixed it with Fable 5.1. It's really impressive. I find, for me personally, I write every day, and I'm still using 5.6 in the ChatGPT for Work app for most of my day-to-day writing. I still find with Fable 5.1, it's still a little bit literary and chunky, but different writers have different styles. Katie Parrott, who's a staff writer, is now back. She was using ChatGPT all the time, and now she's back to Claude for a lot of her writing because she's a big Claude fan for many years until the last few models. I find that, for me, I really like more minimal, straight-ahead prose, and I can collaborate with 5.6 a little bit easier to create that. But I am now flipping back to 5.1, to Fable, when the first prompt or two to ChatGPT doesn't work. So I find it to be a valuable part of my arsenal, where I wasn't using Claude at all for writing anymore. And that's new. And I certainly was never using Fable for writing because Fable is this warp drive thing. So one of the things this model knows how to do that other models don't really know is it knows how sentences connect together. I was working on this piece, I'm working on this history of Codex, and I was working on a piece of it where I asked a question: why was Codex Cloud different? And the details of this don't matter. Why was Codex Cloud different? And then there was some text after that. I had a following paragraph, and I just knew it wasn't really quite working. And I threw it into this model, and it was like, here's what's dragging it down, roughly in order of how much it matters. The biggest problem is that the paragraph goes flat right where it should pay off. You set up a real puzzle, why was Codex Cloud different? And then answer it in the voice of a Help Center article, the right software dependencies, blah, blah, blah. Everywhere else in the piece, you reach for an image. Here, you explain dependency hell literally. So it's actually zeroing in on the right part of the problem and then telling me how to fix it. And the takes that it does, generally the sentences are connecting. It's leading from one sentence to the next. It doesn't write as many of those sentences where you're like, oh yeah, that sounds right. And then you think about it, and you're like, that doesn't mean anything. I have no idea what it means. And that's new for a model of this level of power. So the takeaway here is, if you really liked Cloud models until Opus 5 or Sonata 5 for writing, you're going to really like this model. It's going to be back to the stuff that you love. And if you've been using GPT 5.6 or similar models for your writing, you should give this one a try. It's a good one-two punch to give you a little bit more variation and perspective in your writing that'll help liven it up. So where does this leave us? I want to show you my actual usage. So I got this model on August 24th, and you can see this is in green, this is Codex or ChatGPT for Work, and in orange, it's Cloud code. You can see my level of use has not really changed. I'm mostly using Codex. But then if I look at the number of tokens, my use of Cloud jumped dramatically. Why is that? That's interesting, right? What Fable 5.1 has done for me is given me a model that I set off on a big, chunky task, and I just let it cook. And I'm in ChatGPT for my more basic day-to-day stuff, and then I flip back to Fable every once in a while. And so I'm getting all this work done, but I'm not really checking it because it just does the work end to end. And this is that new style of coding and knowledge work that I'm talking about. You want both, you want both gears, but what Fable 5.1 has really given me is this extra tool that works at the cost and speed of Opus 5, but it gives you the delegation power of Fable that used to be pretty out of reach unless you're willing to spend a ton of money. So you can see that in the graph, right? I'm still using a lot of ChatGPT tokens, but I'm using way more Cloud tokens, and I'm in ChatGPT more. I'm prompting ChatGPT more. It's more of an interactive model. I think that's really interesting. Okay. What did I actually use the tokens on? It's hands. It's these personal feed experiments. It's these big things that I'm building end to end, and the rest of my work I'm mostly still doing in ChatGPT. So final thoughts. If you like Cloud models and you've been disappointed by the recent releases, you're going to love this. You need to check it out. And if you're a big GPT 5.6 fan like I am, it's worth really adding to your arsenal, especially for your big projects and especially for trying to understand, okay, where's knowledge work going? How is knowledge work going to work when I'm delegating lots of things, like making a slide deck end to end, that used to be impossible is starting to be possible now. And so if you want to get a glimpse of this new way of working in a way that isn't going to totally break your bank account, check out Fable 5.1. In addition to this video, we have a long-form written vibe check with all of our benchmarks and all of our testing live now on every.to. We also have a prompt library that you can download to help you get the most out of Fable 5.1. So if you want to go deeper after this video, check it out, every.to. Check it out. not for me. For example, this is a meeting where I can be a tiebreaker. Our head of marketing is talking about a new launch and who we're targeting for that launch and whether we might want to target other customers as well. And it's flagging me as like, Dan, there's, there's a, there's a question about strategy here. You might want to be involved in this. You might want to know about this. And this is like a superpower for me that Fable 5.1 can do is pick out interesting things that you're going to care about from what's happening in your company, what's happening in the world and tell you about it in a way that kind of clicks for you. This is new. Normally it would show you stuff and you'd be like, ah, like this doesn't really click for me. It's like, it's not actually interesting. This is the first model that I've ever found that can pull out something interesting and then present it to you in a way that makes it click. I think the thing that Fable can do, well, one, it just like built this app totally end to end. But the thing that it can do is if when Fable thinks something is interesting, it's more likely to actually be interesting. You know, there's this, um, there's this eval that's like loosely called the McDonald's eval, where if you're talking to your model and you're like, oh, that's exactly like McDonald's. And it's like, oh yeah, you're totally right. But you're talking about something completely random. That's a, that's an example of the kind of thing that models struggle with is I would call it, I mean, to one, to some degree, it's like people pleasing to another, it's basically discernment. Like when you make a connection, is that connection actually interesting? Is it a real connection or is it just that you can connect anything to anything, which these models can do because they know everything. And I think Fable's level of discernment is much higher than previous models I've tried. And that takes me to the next one, which is writing. And obviously like writing is that's, that's what I love. So Anthropic has really struggled with writing recently. They used to be the top writing model, but with Sonar 5 and Opus 5, pretty much everyone was just like, this model doesn't speak English correctly. It doesn't give me good output anymore. And everyone internally at every had switched to ChatGPT. With Fable 5.1, it's not only better than Opus at writing, it's actually starting to compete with ChatGPT. It writes quickly. It has pretty clear prose. It still has a lot of the Claude literariness. But if you look at the reading level that it writes at and the reading ease scores, it has the highest reading ease and the lowest grade level of Opus 5 and GBT 5.6 Sol. So it is just producing stuff that's easier to read. So as an example of the kind of one shot writing this thing does, we asked it to create a blog post based on a simple prompt and a long transcript of Mike Taylor, who's our head of evals, just talking about some of the work that he does. And we asked it to write some of the things that he's learned about getting teams to use AI. And this is the post that it created just from a very, very simple prompt. You can see the prompt is really small. 3. Raise the ceiling, not the floor. 7 things I've learned about getting teams to actually use AI for my first few months running tech consulting at Every. Sam Parr asked me a question on an X last week that I've been asked in one form or another in almost every client meeting since I joined Every. How's everyone getting team adoption for Claude? It's the right question. The gap between what the terminally online people in a company can do with AI and what everyone else does with it has never been wider. This is like pretty good prose. It doesn't feel like slop. Buy the model. Buy the model direct, not third party tools. When a company decides to do AI, the first instinct is to evaluate tools. Someone assembles a spreadsheet of vendors, the Clutter Codex or Gemini under the hood. It doesn't have any of the X, not Ys. It has the raise the ceiling, not the floor, but still, it doesn't have that slop feeling. It's well-structured. It's thoughtful. It's actually really impressive. Let's look back at Opus, Opus 5. Same prompt, same exercise, same everything. Nobody wants to use your AI tools. Opus 5 was such a dick. Nobody wants to use your AI tools. Some of the things I've learned about AI adoption since joining Every Consulting. Sam Parr asked a question last week that I think is the real question in AI right now. That opener, it's a similar opener. It's a little bit more of a minimal sentence construction, which isn't necessarily bad, but you can just feel, let's see. The instinct is to go after laggards. We bought you the tools. Now everyone needs to use AI. This does not work. On pain of death, a good number of people are emotionally unwilling to use AI. What is this thing even talking about? Opus 5 is just so dark and they fixed it with Fable 5.1. It's really impressive. I find for me personally, I write every day and I'm still using 5.6 in the ChatGPT for Work app for most of my day-to-day writing. I still find with Fable 5.1, it's still a little bit literary and chunky, but different writers have different styles. Katie Parrott, who's a staff writer, is now back. She was using ChatGPT all the time and now she's back to Claude for a lot of her writing because she's a big Claude fan for many years until the last few models. I find that for me, I really like more minimal straight-ahead prose and I can collaborate with 5.6 a little bit easier to create that. But I am now flipping back to 5.1 to Fable when the first prompt or two to ChatGPT doesn't work. So I find it to be a valuable part of my arsenal where I wasn't using Claude at all for writing anymore. And that's new. And I certainly was never using Fable for writing because Fable is like this warp drive thing. So one of the things this model knows how to do that other models don't really know, is it knows how sentences connect together. Like I was working on this piece, I'm working on this history of Codex, and I was working on a piece of it where I asked a question, why was Codex Cloud different? And the details of this don't matter. Why was Codex Cloud different? And then there was some text after that. I had a falling paragraph and I just knew it wasn't really quite working. And I threw it into this model and it was like, here's what's dragging it down, roughly in order of how much it matters. The biggest problem is that the paragraph goes flat, right where it should pay off. You set up a real puzzle, why was Codex Cloud different? And then answer it in the voice of a Help Center article, the right software dependencies, blah, blah, blah. Everywhere else in the piece you reach for an image, here you explain dependency hell literally. So it's actually zeroing in on the right part of the problem and then telling me how to fix it. And the takes that it does, generally the sentence are connecting. It's leading from one sentence to the next. It doesn't write as many of those sentences where you're like, oh yeah, that sounds right. And then you think about it and you're like, that doesn't mean anything. I have no idea what it means. And that's new for a model of this level of power. So the takeaway here is if you really liked Cloud models until Opus 5 or Sonata 5 for writing, you're going to really like this model. It's going to be back to the stuff that you love. And if you've been using GPT 5.6 or similar models for your writing, you should give this one a try. It's a good one-two punch to give you a little bit more variation and perspective in your writing that'll help liven it up. So where does this leave us? I want to show you like my actual usage. So I got this model on August 24th and you can see, you know, this is in green, this is Codex or ChatGPT for work and in orange, it's Cloud code. You can see it. My level of use has not really changed. I'm mostly using Codex. But then if I look at the number of tokens, my use of Cloud jumped dramatically. Why is that? That's interesting, right? What Fable 5.1 has done for me is given me a model that I set off on a big, chunky task and I just let it cook. And I'm in ChatGPT for my more basic day-to-day stuff. And then I flip back to Fable every once in a while. And so I'm getting like all this work done, but I'm not really checking it because it just does the work end to end. And this is that new style of coding and knowledge work that I'm talking about. You want both, you want both gears, but what Fable 5.1 has really given me is this extra tool that works at the cost and speed of Opus 5, but it gives you the delegation power of Fable that used to be pretty out of reach unless you're willing to spend a ton of money. So you can sort of see that in the graph, right? Like I'm still using a lot of ChatGPT tokens, but I'm using way more Cloud tokens and I'm in ChatGPT more. I'm prompting ChatGPT more. It's more of an interactive model. I think that's really interesting. Okay. What did I actually use the tokens on? It's like, it's hands. It's this personal feed experiments. It's these big things that I'm building end to end and the rest of my work I'm mostly still doing in ChatGPT. So final thoughts. If you like Cloud models and you've been disappointed by the recent releases, you're going to love this. You need to check it out. And if you're a big GPT 5.6 fan like I am, it's worth really adding to your arsenal, especially for your big projects and especially for trying to understand, okay, where's knowledge work going? How is knowledge work going to work when I'm delegating lots of things like making a slide deck end to end that used to be impossible is starting to be possible now. And so if you want to get a glimpse of this new way of working in a way that isn't going to totally break your bank account, check out Fable 5.1. In addition to this video, we have a long form written vibe check with all of our benchmarks and all of our testing live now on every.to. We also have a prompt library that you can download to help you get the most out of Fable 5.1. So if you want to go deeper after this video, check it out, every.to. check it out.