Open Reader

VIBE CHECK: GPT-6 ASTRA

completed 1:17:32 Sep 04, 2026 Watch on YouTube

Current Status

completed

Video ID

JTvE7v_rMIw

RAG / Chat

Enabled
VIBE CHECK: GPT-6 ASTRA
Description

Read Our Vibe Check: https://every.to/vibe-check/gpt-6-astra-vibe-check?utm_source=youtube&utm_medium=social&utm_campaign=gpt_6_astra&utm_content=launchlive

Summary

Generated by gpt-5.6-terra

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: The team judges GPT-6 Astra as a major upgrade for computer use, one-shot knowledge-work scaffolding, visually rich coding, and writing, but less trustworthy than Fable 5.1 on high-stakes, tightly scoped work because it routinely adds unnecessary features, copy, and UI.
  • Why it matters: Astra appears to shift computer use from a novelty into a practical delegation layer for repetitive browser, desktop, media, and administrative workflows, while exposing a central operating problem: model capability is now high enough that taste, constraint-following, and evaluation harnesses determine real value.
  • Best use: Use the video as an implementation-oriented model-selection review: adopt Astra for supervised computer-use and expansive first drafts, retain Fable for critical execution where simplicity and restraint matter, and build task-specific blind tests before standardizing either.

Executive Summary

Every's team presents GPT-6 Astra as an unusually capable frontier model that is strongest when it can act across software rather than merely generate text or code. Their standout evidence is delegated computer use: Astra reportedly edited a 20-minute video in Premiere Pro, generated and formatted screenshots for a publication, handled browser-based vendor and Gmail tasks, and can operate in the background through the ChatGPT for Work/Codex harness. The practical framing is not that every task should be automated, but that recurring, multi-step tasks are now candidates for delegation.

The central tradeoff is judgment. Astra can one-shot impressive artifacts—a historically styled Battle of Waterloo simulation, 3D set and Blender scenes, a Bible-study dashboard, landing pages, dashboards, and an app for scanning handwritten notebooks—but it repeatedly defaults to surplus. Reviewers describe unnecessary navigation bars, labels, buttons, marketing copy, formatting treatments, and workflows that require users to remove rather than refine output. Fable 5.1 is described as the more restrained and trusted colleague for the largest or most exacting tasks.

Writing results are notably stronger than prior models, although mixed by task. Astra produced the skeleton of its own 4,000-word review overnight from internal context and followed an editorial structure better than older models; a 50-comparison blind writing test slightly favored Astra over Fable 5.1. Yet another benchmark found Astra sanding down distinctive language, omitting supplied statistics, and linking to source material rather than embedding it. The implication is that its prose should be evaluated through the organization's own editorial standards, not by first-impression fluency.

The video’s deeper operational argument is that generic leaderboard comparisons are insufficient. Mike Taylor's "checks" system converts personal standards—correctness, readability, distinctive design, factual grounding, and task completion—into reusable automated evaluations followed by human review. The team also flags two material constraints on agent deployment: Astra can refuse benign administrative tasks due to safeguards, and its economics under real token pricing remain untested.

Key Takeaways

  • Claim: Astra is the team's preferred model for computer use and may make recurring desktop and browser tasks meaningfully delegable. | Evidence: The team says Astra edited their approximately 20-minute Fable 5.1 review video in Premiere Pro after a human loaded raw footage; Katie says it gathered, formatted, captioned, and placed screenshots according to a Notion style guide; other examples include Gmail management, Facebook Marketplace posting/selling, customer-support interactions, vendor procurement forms, and moving a prior Claude conversation into Codex. | Implication: Ken should prioritize Astra trials around repeated, auditable browser/desktop workflows with clear completion criteria, rather than use it indiscriminately for every individual action. | Caveat: Speakers distinguish recurring or form-heavy work from one-off tasks, which can still be faster to do manually; they also report refusals on sensitive-looking administrative workflows.
  • Claim: For complex one-shot coding and design, Astra is highly generative and visually capable, but it is less constrained and dependable than Fable 5.1. | Evidence: Astra produced a one-prompt Battle of Waterloo simulation, a detailed Bible-study dashboard from a file of notes, a 3D Hadestown set that reviewers felt reflected visual references, a Blender vineyard scene with generated textures, and a handwritten-notebook transcription app. In the notebook-app comparison, however, Astra added multiple capture controls and steps while Fable proposed a simpler hold-page/press-space workflow; Kieran also found functional errors in Astra's recreation of a Markdown editor that Fable versions did not exhibit. | Implication: Use Astra to generate ambitious prototypes and visual explorations, but apply acceptance tests and a simplification pass before promoting output into production workflows. | Caveat: The reviewers regard Astra as particularly compelling for landing pages, 3D environments, design exploration, and idea generation; the comparative weakness is most consequential on large production apps and tasks needing crisp workflow choices.
  • Claim: Astra’s characteristic failure mode is overproduction: it makes attractive artifacts but persistently adds UI, copy, and formatting that were not requested. | Evidence: Across app, dashboard, "cozy island," breathwork, and landing-page tests, reviewers cite gratuitous navigation bars, logos, status labels, excessive buttons, repeated warm-paper styling, dense instructional copy, and headline conventions such as a two-line treatment with an italicized second line. Kieran calls the output more like "fast food"—immediately impressive but quickly tiring—than Fable's more focused output. | Implication: Ken should encode negative constraints in prompts and evals—no navigation unless required, minimal controls, no decorative labels, preserve supplied layout—and make “deletion burden” a first-class quality metric. | Caveat: The same tendency is a feature when the goal is broad ideation or a polished marketing-style landing page; one reviewer called Astra's landing page the best-looking generated page they had seen.
  • Claim: Astra can create unusually useful first drafts from broad internal context, reducing the work of collection, organization, and production formatting rather than only sentence drafting. | Evidence: After learning at 3 a.m. that the model would launch, the team had Astra inspect relevant Slack channels, benchmarks, and links and produce the base draft of a 4,000-word review by about 7 a.m. Katie says the draft included the right structure, contextual use of evidence, blank screenshot placement, and later image formatting/caption work. | Implication: Deploy Astra as a research-to-draft pipeline over curated internal sources, with a mandatory evidence-validation and editorial-review stage rather than treating a polished draft as a verified conclusion. | Caveat: Katie explicitly says she would normally first compile and verify a finding summary, identify contradictions and disagreements, and validate the sequencing; human editorial judgment remained necessary.
  • Claim: Astra's writing quality is competitive with Fable 5.1 but task-dependent, so a general claim that it is the best writing model is not supported by this review. | Evidence: In Mike Taylor's blind test of 50 A/B paragraph comparisons from a real 1,200-word article, Astra narrowly beat Fable 5.1. But in a separate consulting article task, it was judged to have flattened distinctive language, dropped supplied statistics, and linked to a tweet rather than displaying it. It performed strongly on a client training-plan task, receiving 100% on Mike's checks for depth, interview grounding, and absence of hallucination. | Implication: Benchmark Astra against Ken's actual artifacts—especially factual retention, source presentation, and voice—using blind comparisons and rich context documents rather than choosing it from generic writing demos. | Caveat: The team expects recognizable stylistic patterns to emerge with wider use, and Allie Miller says Fable 5.1 currently performs better after being supplied explicit theory-of-mind guidance via SOPs, skills, or context documents.
  • Claim: The model-selection process should be personalized and automated through explicit checks rather than driven by novelty, demos, or aggregate benchmarks. | Evidence: Mike's checks runner executes standard tasks across writing, dashboards, PowerPoint, and app generation, then evaluates both objective requirements such as page load and correct NPS calculation and subjective standards such as readable typography and distinctive design. He proposes giving each employee a personal benchmark so new model releases can be assessed without manual retesting from scratch. | Implication: Ken should maintain versioned, role-specific eval suites for agents and models, combining deterministic tests, artifact inspection, and human taste criteria; this is more durable than adopting a single vendor-wide default. | Caveat: The presenters acknowledge that aesthetic checks are subjective and intended to answer whether a model is useful to a particular operator, not establish a universal best model.
  • Claim: Astra's safeguards and unproven production economics are material adoption constraints despite its capability gains. | Evidence: One speaker reports Astra refused roughly seven times to complete a school medical form for a child with a mild peanut allergy, requiring attempted prompt injection before succeeding; the group also anticipates that safety changes following an unspecified cyber-related incident may reduce usability. Dan adds that he has not yet tested whether Astra remains a daily driver once token costs and limits apply. | Implication: Before operationalizing Astra, measure refusal behavior, recovery paths, cost per completed workflow, and human-intervention rate on the exact classes of tasks Ken intends to delegate. | Caveat: The transcript does not provide pricing, refusal rates, policy boundaries, or controlled reliability data, so the scope of either limitation is unknown.

Detailed Brief

Comparative model positioning and the reach test

  • Claims: Dan and Katie classify Astra as a green, daily-use model, while Kieran and Mike classify it as yellow because Fable 5.1 remains their default for trusted execution.; Katie's preference is partly interactional: Astra feels approachable and conversational despite being a large model, whereas she treats Fable as a scarce, expensive resource that requires more deliberate prompting.; The team repeatedly frames Fable as a model that infers the intended workflow and adds a small number of useful improvements, while Astra tends to broaden the scope of the artifact.
  • Evidence: Dan says Astra is his daily driver across coding, writing, and knowledge work, but he still reaches for Fable on the biggest tasks.; Kieran says he would use Astra for landing pages, 3D games, or multi-shot ideation, but not as a daily driver at equal price to Fable.; Mike ran Astra and other models roughly 50/50 while checking whether he had missed a major capability, ultimately identifying computer use as the area most likely to change his view.
  • Caveats: The review is largely qualitative and reflects a small expert team's preferences and private testing, not a published controlled benchmark.; The transcript contains promotional material for Every and repeats sections, so its strongest value is the concrete examples and evaluation methodology rather than an independent market-wide ranking.
  • Implications: A portfolio approach is preferable: select models by task class and failure tolerance rather than seeking a universal daily driver.; The gap between a spectacular demo and a trusted operator is increasingly defined by scope discipline, reliability, and cost.

Design, spatial awareness, and visual-generation observations

  • Claims: Astra represents a substantial advance in spatial and visual reasoning relative to older models, particularly in 3D scenes, interactive environments, and presentation design.; Visual sophistication does not equate to product taste: Astra's generated interfaces often resemble polished marketing sites rather than the simplest functional product.; The team identifies Grok as an under-discussed competitor that performed strongly in some dashboard, writing, and presentation tests, including generation of custom SVGs.
  • Evidence: Astra created a 3D Hadestown stage whose layout and details appeared to reflect visual references, versus Fable's version, which the reviewer felt resembled a textual interpretation.; In a PowerPoint spatial-layout benchmark, Astra was described as nearly correct while older Soul output was visibly janky; Fable 5.1 still missed aspects of the visual arrangement.; Grok was said to produce a more distinctive dashboard narrative and attractive custom SVGs in a presentation task.
  • Caveats: The reviewer's aesthetic assessment is subjective, and the transcript does not provide the raw benchmark prompts, artifacts, or scoring data needed to independently verify comparisons.; Some apparent scene detail used shortcuts: in the Blender vineyard, the background was an image projected onto a plane rather than fully modeled geometry.
  • Implications: Visual-agent evaluation should separately test spatial correctness, interaction design, originality, and task appropriateness; a single 'looks good' score masks these distinct qualities.; Competitive routing should remain open: the team’s own discussion suggests model strengths are converging and can vary sharply by artifact type.

Notable Concepts & Terms

  • Reach test: The team's practical model-quality test: whether an operator actually reaches for a model in daily work, and for which tasks instead of alternatives.
  • Computer use: An agent operating browser and desktop applications via a cursor or equivalent interface; Astra is presented as the model's most consequential capability area.
  • ChatGPT for Work / Codex harness: The execution environment the speakers favor for computer use because it can run tasks in the background with its own cursor across browser and desktop apps.
  • Mike's checks: A reusable evaluation runner that generates standard model outputs and grades them against objective requirements plus encoded personal taste criteria.
  • Theory of mind context: SOPs, skills, or context documents that teach a model how to account for audience needs, narrative structure, and visual communication; presented as a way to improve writing performance.
  • Deletion burden: The practical cost of a model doing extra work that the user must remove—Astra's central usability weakness in this review.
  • Vibe slop: Formulaic generated design or copy marked by generic landing-page patterns, visual sameness, excessive decoration, and recognizable model fingerprints.
  • Warm-paper aesthetic: A recurring beige/green, editorial-style visual treatment the team sees Astra applying too broadly, regardless of the requested product concept.

Operator Notes / Why Ken Should Care

  • Create a controlled Astra pilot for 3-5 recurring computer-use workflows, with sandboxed accounts where possible, explicit stop conditions, logs, and human approval before consequential submissions.
  • Add an agent-eval dimension for unnecessary output: count superfluous controls, labels, workflow steps, copy blocks, and requested revisions needed to reach a minimal usable artifact.
  • Build a model-routing matrix: Astra for supervised UI/browser execution, research synthesis, visual prototyping, and broad first drafts; retain the most trusted alternative for high-stakes code changes and tightly specified workflows.
  • Run blind A/B tests on Ken's own investment memos, operating documents, decks, and agent outputs; score factual retention, source traceability, voice, decision usefulness, and edit time.
  • Supply reusable audience/theory-of-mind context packs to both Astra and competing models, then compare outputs with and without the pack before attributing quality differences to the base model.
  • Track production metrics before committing: token cost per completed task, refusal rate by workflow type, task-completion rate, recovery success, and human minutes saved.
  • Do not rely on prompt-injection workarounds for safety refusals; establish an alternate manual or model-routing path for forms, financial, health, tax, and other potentially restricted tasks.

Source/Metadata

  • Title: VIBE CHECK: GPT-6 ASTRA
  • Transcript words: 21661
  • Duration seconds: 4646
  • Timestamp note: No timestamps or chapter markers were present in the supplied transcript; substantial portions of the transcript are duplicated.

Transcript

13986 words en Processed in 445.6s

And we are live. Welcome, everybody. We're back again in a familiar place. We are reviewing by checking a new model. This time, today, it is from OpenAI. It is GBT-6. GBT-6. Astra. In honor of Astra, I'm wearing a NASA T-shirt. Thrifted, may I say. And the big question is, how good is it? And also, is it AGI? We've got Greg Brockman going on Axios saying he thinks it's pretty close. We have been testing it for the last, well, we actually can't say how long we've been testing it, but we've been testing it. We've been testing it extensively, extensively across coding, writing, across knowledge work. And we've got your Day Zero vibe check. I'm joined by the Every Team. Let's go around the corner and introduce ourselves. Kieran, you want to go first? Yes, I am Kieran. I run many agents and have a plugin called Compout Engineering Plugin. Amazing. Now we've got Katie Parrott. I'm Katie Parrott, staff writer. And as of this week, I also have my own plugin called Compound Writing. So two can play at that game. Let's go. And now we've got Mike. I'm Mike Taylor, or Michael if I'm in trouble. And I don't have a Compound plugin yet. Yet. But I head up evals at Every. Awesome. And I am Dan Shipper. I'm the co-founder and CEO of Every. Every is the only subscription you need to stay at the Edge of AI. Hi. We have a lot, a lot, a lot for you today. If you are here and you're interested in what we've made, what we've been able to do with Astra, you should know that everything is on Every. We have a full vibe check. It's like 4,000 words. It's at every.to/slash vibe check slash GBT6 Astra vibe check. I'm highlighting the URL here. You should check it out. It's got everything you need and more, written by our very own Katie Parrott. The headline is, vibe check, GBT6 Astra is a big upgrade with some bad habits. Interesting. Interesting. One thing that I think we should start with, which should tell you something about how good this model is, is we got word that this model was going to drop at 3 a.m. Eastern time, and it is now 3:30 p.m. Eastern. So it's about 12 hours later. And this entire vibe check was written front to back in that time. We also just published a video on our YouTube. It's like a three- or four-minute video. We're doing an insane amount of work to review this model, and we used the model to do the review. I want to show you, this is our very own Katie Parrott. We used Astra to one-shot its own review. This was the base of the review that the vibe check we just put out. There's some really good writing in here. So I think that should give you a sense for how powerful this is, that we actually one-shot our own review of this model in a way that served as the base and let us get to a really great 4,000-word review in less than 12 hours. It was really like eight hours, something like that. Let's go through the high level here. It's a very good model. For me, it's my favorite writing model. I use it every day. It's still my daily driver. It is incredible at computer use. We actually dropped a Fable 5.1 video on YouTube. It was like a 20-minute vibe check of Fable 5.1 two days ago. And Astra actually edited that. Our head of video, Randy, just set it off and let it churn in Premiere for like five hours, and it did a lot of the editing for that video. So there's something really interesting here going on with computer use. It is an insane model for that. It's also good at 3D games, 3D designs, and just design in general. Actually, I'll go into this more, but I had it build a reconstruction of the Battle of Waterloo that I can scrub through. It's historically accurate. It's got all the geography. It's got all the little uniforms. This is one shot, folks. So that's just a little taste of what's possible with Astra. And it is a new model class above Soul. So it's more expensive. It's the same price as Fable. So we're really comparing it to Fable apples to apples. I think our feeling from testing this for a little while is that this is a really good model, has some incredible things that it does, and the headline of this is with some bad habits. There are a couple of things that, at the top end, I don't think it quite reaches Fable. We have some different opinions on the team, but I think overall my sense is it doesn't quite reach Fable 5.1 on the top end. It does a couple of things, like it'll overpack landing pages with too much text and too much stuff. It's a good writer, but once you put it into software, it starts adding little labels and all this extraneous stuff in there. And I think Fable has this thing that makes it very special, which is when you give it a prompt, it intuitively understands what you want and then extends it in a way that's surprising but keeps it pretty simple and focused. And Astra does this to some extent, but it's not as good. And it sometimes goes off the rails a little bit and just adds a bunch of bells and whistles that you don't really want. So, incredibly smart model, big jump over Soul, very good at writing, very good at coding, very good at knowledge work stuff. And we're going to go through all the tests we put it through. Not quite at the top end. So for your biggest tasks, for me, I think Fable still edges it. So now I want to go around the horn. I want to talk through the reach test. So the reach test is our way of talking about, I think it's the best early judge of model quality. Are you using it every day? Are you reaching for it? And for what situations? So for me, this is now my daily driver. I'm doing pretty much all my day-to-day tasks in it. So everything from coding to writing to knowledge work. But I'm really reaching for Fable for the biggest things. So it gets a green. It's a jump for my daily stuff. I think my one caveat on this is I have to see what it's like when we're paying for the tokens, how quickly I run out, and whether I'm going back to 5.6 or not. So that's my one little caveat on this. But right now it is my daily driver. I think it's a very good model. Let's go around the horn a little bit. Kieran, you're a yellow. Can you give us your take, give us your overall take, and your reach test? Yeah. So I compare this with Fable 5.1, because that's a fair comparison, which is my daily driver. And I had a lot of fun with Astra. It's really fun. You can do insane things, and it will go over the top with things in a good way, which is really fun. But also, I work, and really what I want is I want it to do work for me. Ideally, it makes my life easier and just does the work I don't want to do or elevates my thinking. And I think Fable, I really trust. Like I said in my Fable 5.1 review, I kind of feel Fable is a trustworthy colleague that I can trust. And it's really the only model ever where I felt that, Fable 5 and Fable 5.1. And I was hoping for Astra to give me that same feeling where, oh yeah, I trust this model, but I don't. It's adding so many things here and there that sometimes are fun. But most of the time you have to say, oh, actually, just delete all these things. And that loop of work where you have to remove work is, I hoped we were past this. And I was thinking that's only with the to do or elevate my thinking. And I think fable, I really trust, as I said in my fable 5.1 review, that I feel fable is a trustworthy colleague that I can trust. And it's really the only model ever where I felt that, fable 5 and fable 5.1. And I was hoping for Astra to give me that same feeling where, oh yeah, I trust this model, but I don't. It's adding so many things here and there that sometimes are fun. But most of the time you have to say, oh, actually just delete all these things. And that loop of work where you have to remove work is, I hoped we were past this. And I was thinking that's only with the cheaper models. So it's just something that really annoys me. If I have a model and I give it a task, it needs to do the task. I don't want it to do 20 more things. It's fine if it does two things better than what I was thinking with fable. But I have to say, design for landing pages is insane. If you are a designer, this is going to be a very fun and cool model. And for normal coding, it's also really good, but it's not as good as fable 5.1. And are you reaching for it at any point in your day to day? And if so, for what? No, I would never reach for it. If it's the same price as fable, there's no reason for me to take a model that's less good and less trustworthy than fable. And maybe if I do landing pages, I would reach for it or build 3D games. Or if I just want to have fun or see what this model would do. I would use it as a multi shot and see, hey, maybe there is something that is interesting in here, and then take it from there. Like ID generation, but not as a daily driver. All right, there you have it. Not as a daily driver. That's why it's here in the yellow. Let me go back to our vibe check page. I'm going to throw it up there on the stage again. Now we've got Mike Taylor, also yellow. Mike, give us your overall thoughts and tell us about your reach test. What are you reaching for for this model, if anything, and why? Yeah, it's a strange industry we operate in because if you'd asked me last week, I would have said, yeah, this is green. This is amazing, my new daily driver. And because we just got fable 5.1, which solved a lot of the issues that I had with previous fable, now it's kind of like, yeah, yeah, this is amazing. We're living in really great times, and it's probably the most fun having it work, but I used it as much as possible and I didn't see it make any real major mistakes, but I also didn't really see it do anything special compared to fable. Hmm. And so is it part of your daily routine? I know you've switched over the last month or so from pretty much all cloud to, I think, a lot of codecs. Is that still the case, or are you switching back? Where are you? I'm right now running both 50-50, just because I feel like maybe I'm missing something. Everyone's saying Astra is AGI, but maybe it's hype. But no, I think maybe what I might be missing, and this is where I want to explore more, is on the computer use side. It is much, much faster. And I just don't really feel like I've trusted AI enough on computer use. And so maybe now I can be more ambitious with that. So that's where I think my opinion might change over the next few weeks. I think that's totally true. For people who just joined, I was saying this edited end to end a 20-minute video that has 25,000 views. Our fable 5.1 vibe check was edited by Astra, weirdly. And it did it in Premiere Pro. That's a complicated piece of software that I can't even use. So there's some really interesting things happening in computer use that should, I think, change how we think about using computers at all. Okay. Now we've got the woman of the hour, Katie Parrott. Katie wrote this entire thing again. This is a 4,000-word vibe check. Wait, I'm not sharing my screen. This is a 4,000-word vibe check that she wrote end to end with Astra in just a couple hours because we've been testing this model for a little bit, but we got word that it would be launched at 3 a.m. Eastern. And here we are. I mean, Katie, this would not be possible without Astra, or we've never done it like this before. You are a green. Can you tell us about your experience with this model, your overall take, and your reach test? Yeah, I love this model. I am in its debt, as you mentioned, in the fact that we had a very tight turnaround on this review. Honestly, I had rolled over and checked Slack, which you should not do in the middle of the night. Kids, don't do it. And seeing that it was go time, I got out of bed, walked to my office, said, hey, Astra, can you look at these channels and run and write the vibe check based on these examples and this context? And I came back to a finished draft at seven in the morning. So I went back to sleep, and there was, you know, it's not one and done. There's a ton of work, and I wrote in the live check, it takes a village to produce a live check. And so it was definitely a team effort between me and Jack and Kate and Kieran and Mike and Dan and everybody to get this thing written. But we were able to just get a really good skeleton with one prompt. And I don't think I've ever seen a model do that before. Generally, I want to draft, and I wrote this in my compound writing guide for best practices for writing with AI, you want to go section by section because a decision that you make up top can have ripple effects down the draft. But I set this up and it followed the instructions, and it found the right context and used it appropriately. So I am super green on this model for this week. And then just overall, because what really stands out to me is, for me, the reason I reach for this potentially more than fable 5.1 is that it's a big model with small model energy in a really good way. It feels very accessible and like I can go back and forth with it. And I forget that I'm using the big scary model and not the normal friendly model. Whereas with fable, I'm always aware this is fable, and I better make sure that I'm using my tokens responsibly with it because it's fable. And so this model just doesn't have that aura, which I think is a good thing from an accessibility standpoint, because it feels like something that you can approach and you don't have to be scared of. I totally agree. So what I want to do now, we're going to get into the vibe check. So again, let's just open it up. We've got the vibe check. Katie wrote this 4,000-word vibe check. It's up on every dot to every is the only subscription you need to stay at the edge of AI. And we have been testing this model. We have a ton of stuff to talk about, everything from coding to writing to design to knowledge work. We've done a lot of different tests on this model that we're going to go through. What I want to start with is coding. So, I think this is a good thing from an accessibility standpoint, because it feels like something that you can approach, and you don't have to be scared of. I totally agree. So what I wanted, what I want to do now, we're going to get into the vibe check. So again, we've got, let's just open it up. We've got the vibe check. Katie wrote this 4000 word vibe check. It's up on every, every dot to every is the only subscription you need to stay at the edge of AI. And we have been testing this model. We have a ton of stuff to talk about, everything from coding to writing to design to knowledge work, to we've done a lot of different tests on this model that we're going to go through. What I want to start with is coding. So the thing to note about coding is it does stuff like this. This is a one-shot app that reenacts the Battle of Waterloo, which is crazy. And I just said, hey, this is there, this is a, the French are in retreat right now. I just said, hey, I want you to research Waterloo and create a historically accurate rendition of the last charge and then retreat. And it just did this. And it's crazy. Anyone with programming experience want to say a little bit more about what's so crazy that it does this. Kieran. Kieran. Kieran. Kieran. I mean, I've never done this, but there are just many things that can go wrong here. And it also needs to look very good, so it has to tickle your brain in that way. There are just lots of things that can go wrong, especially with performance as well. I'm no game designer in any way, but it's a very complex thing to do this. ohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohohoh can help you see that a little bit. So this is one that we put together from Jack Chang, who's one of our editors who helped put together this vibe check and do a bunch of testing on it. You should check out Jack. He's amazing. And Jack is an editor. He's also a writer. He's actually written several children's books, which you should absolutely read. Actually, the way I found Jack is I read one of his books called See You in the Cosmos. It was one of my favorite books. And I was like, I need to meet this guy. And that's how we started chatting and hanging out. And now we work together. And it's amazing. Anyway, incredible writer. And he has all of these handwritten journals that he wants to digitize. And so he thought that would be a good thing to try, Fable 5.1 versus GBC6 digitizing his journals. So he did a little prompt that was something like, let's see, I think we have the prompt here. Actually, I don't know if we have the prompt on here. But the prompt was essentially, read the exact prompt. Here we go. You're building a complete working Mac OS app in one shot. No one will answer questions during this task. If anything is ambiguous, blah, blah, blah. What I want you to do is use the Mac's built-in webcam as its input, support a whole notebook session, the user works through many consecutive pages and no pages lost, produce an accurate transcript, and then ask the user where to save. So basically, he wants something where he can just hold up his notebook and get a transcription just by turning the pages, right? So let's see how Fable and Asher get from this task. So we have the welcome screen. So this is Fable's version of it. And you can see it's a little basic, but it just does the job. It's like, this is the notebook scanner, one, two, three. I think there's probably a little bit too much copy here, but that's the sort of bare take on it. This is Astra. So we've been talking a little bit about Astra's design sense. I do think Astra has a really nice design sense. In my view, it goes for warm paper a little too often. And this is a thing with GPT models since, I don't know, probably GPT five, five or something, five, four, five, five, six. So it's a little bit like you can tell it's a little bit GPT-ish because of the warm paper, but it's got the warm paper. It's got green accents. It's a little bit more designed. You can see there are these one, two, three, that kind of stuff. And this copy is nice, but there's also a lot more complexity, like from paper to markdown and then hold each page to your Mac's camera, view the writing, keep the whole notebook. But that's already also said over here. And then these little labels where it's like photos you choose to transcribe or sent to open router. And it's just like, there's a lot in these. There are like three buttons. That's just a lot. And that's some of the gripes that I have with Astra coming through here. It's just a little bit extra, as Kieran says. So the next thing that I think is interesting and an important part of this is, how do I show you? Okay, here we go. The next thing that is interesting and an important part of this is, what does the actual scanning experience look like in this app? This is the Fable version. For Fable, it's still pretty basic, but it's basically just like hold up a page and then press the space bar, and then we'll do the next page, hold up the page, press the space bar, do the next page, and then finish. That's actually a really nice workflow that Jack didn't have to specify. It just works. And it's an example of Fable getting the sort of intuitive, ooh, I know what you want, and I'm going to do a version of it that maybe you didn't explicitly say, but works the way that you think it should. And it's still pretty simple. And Astra does this, which again, there's a lot of similarity, but look at all these: prepare photo, rotate 90 degrees, apply crop, clear selection, this page is blank, type it myself, transcribe page, use other readers. So in order to do this, he has to hold up the page, click to take a picture, then press transcribe, then go to the next one. So it's a much more complicated process, and the whole UI is much more complicated. And I think that that's basically a good example of the difference between these two models for big one-shot tasks and their level of taste and their ability to make decisions without a lot of direction. Kieran, Mike, Katie, anyone have reactions to that from your experience? Yeah, very similar. Also, there's a navigation bar. All my apps have a top navigation bar now, even though it's a 3D island or something else, which is kind of interesting. It's really steering into marketing pages. If you share my screen, I have one app. This is my OG coding example, which is rewrites proof from scratch, which is Dan's amazing markdown editor that I extracted a prompt from and said, okay, you just build this. Dan struggled for many months to get it to this pristine quality. And we just one-shot these things now to see if that actually works. But difference between these two models for big one-shot tasks and their level of taste and their ability to make decisions without a lot of direction. Kieran, Mike, Katie, anyone have reactions to that from your experience? Yeah, very similar. Also, there's a navigation bar. All my apps have a top navigation bar now, even though it's a 3D Island or something else, which is interesting. It's really steering into marketing pages. If you share my screen, I have one app. This is my OG coding example, which is rewrites proof from scratch, which is Dan's amazing markdown editor that I extracted a prompt from and say, okay, you just build this. Dan struggled for many months to get it to this pristine quality. And we just one-shot these things now to see if that actually works. But there are errors, as you can see. Accepting this doesn't work, which Fable 5 and Fable 5.1 did not have any errors. To me, that is the top end where I'm talking about. It looks a little bit busy, but it looks cool. But for an app, I would say it's not as good as Fable here. I want the simplicity because the simplicity gives me a direction where you go. And it's a little bit, yeah, and it doesn't work. So this is the hardest feel-my- benchmark. The rest, it did pretty good. But for coding in this benchmark, it just is not as good as Fable 5 or Fable 5.1 because both did better. All right, there you have it, folks. I want to add, so we've got Allie Miller here. Allie, welcome. How you doing? We can't mute it. You're muted. Hanging outside the courthouse. How's it going, guys? Did they get you for vibe coding too hard? Exactly. Yeah. It happens to the best of us. A journey of our peers also has decided that Astra is awesome. So, let's talk about it. Awesome. We're psyched to have you. We've just been going through our Astra reviews. Whenever we talk about new models, we talk about the reach test. Are we reaching for the model? And if so, for what? And over what? How has our behavior changed? Curious what that's been like for you. I know you've also been testing this model. I think one of the, so I test things from a business lens primarily. And I think the thing that I felt like I had been reaching for with Claude that I'm no longer reaching for is just tone and vibe. I am not yelling at this model. I hear you, Kieran, on some of the errors, and I was seeing the same thing, but I'm yelling at Claude multiple times a day just to say, what the flip are you saying? And I think business professionals would appreciate the personality of Astra more than Fable. The standout thing to me, I saw you were walking through the game design stuff. I love all the 3D design. I don't think that it has a massive space yet in the business space. Computer use is out of this world. If you have not, if you are waiting for some moment to do computer use, browser use, whatever, GBT six Astra, this is the wake-up moment. And I just tell people, imagine that someone has recorded your screen for a week. What would they see that you're doing constantly? And can you ask AI to do that? With previous models, I asked it to take over my computer and unsubscribe me from every single Nextdoor email ever. And I wouldn't have asked previous models because I tried and it failed to go through my whole Gmail to manage that. And I would with Astra. Hmm. What are you using it for on the computer use front? I know we edited the, we actually edited the Fable 5.1 video that we released with Astra, but yeah, what are your top computer use cases? Because I think we're all starting to feel like there's this new way of using computers because these models can use them faster than you can, and often better. So how are you using it in your day to day? Yeah. And I would say that it's kind of like a measure twice, cut a thousand times. If it is truly a one-off task, it's still faster for me to do it. But if it's something like fill out this whole thing for me, then it's faster. I'm going to give you an insane example, which is that I was remote. I was on my phone. As you know, the codex experience remote is a million times better than Claude. And so Claude, all my chats were disconnected. So I actually used codex computer use to control my computer, to go into Claude, to find a chat that I had already had inside of Claude and ported it over to codex so that I could continue it there. So the ability to also jump between the, I know that's such a weird use case, but I trusted it. And then a lot of stuff around Gmail management, which again is just a normal businessy task. And then we've been trying stuff on very specific vendor procurement sort of things. So a logistics checkout that just has a bunch of different forms, and it handled it very, very well. I love it. I love it. All right. We're going to keep going around the horn. Allie, stick around. We loved your take as we go through. We just did coding, and we're going to get into writing in a second. Then we're going to go through knowledge work. I feel like you'll have a lot of things to say. I know you're also doing jury duty. So whenever you need to jump, just feel free to let me know, but we're excited. We're psyched to have you here on our vibe check today. Before we move on, anything else on coding, either from Mike or Katie, that you want to add? I saw the same things as Kieran. There were a bunch of additional marketing flourishes that it added. I was trying to build a simple app, and I had to keep telling it, no, simpler, simpler, simpler. And so I feel like it's almost trying a little bit too hard, but that's true of all of the AI agents, I guess. Yes. Yes. And Katie, have you been vibe coding with this at all? I know you've been busy writing, but yes, I'm curious. Well, there's always time for vibe coding with a new frontier model. I actually did. I've had several different models take a pass at building my sort of Bible study homepage. So basically, I just want a nice, friendly dashboard for viewing and interacting with all the notes and my bibliography at my small seminary of study Bibles, et cetera. And what Astra built me, Astra was the first one to just do the thing in a way that I'm like, oh yeah, I'm excited to use this. So if I share my screen, I can just show you. Yeah. Let me just add you. Whenever you share, I'll add it to the stage. Okay. All right. You're on stage. Here we go. Yep. So this is scriptorium. And so this has got, this is one of the latest notes that I made about the word righteous or righteousness. It's got a notebook with all of these. Tell us about righteousness. Tell us about righteousness. I don't redo that. May I just say that this is a warm paper background? Just saying, just saying. Anyway, yes. I never said I have very much. It is very ChatGPT. I never said that. I never said anything else. Otherwise, I love warm paper. I'm a writer. Warm paper is very much for me. So shout out to the OpenAI team for designing an app that would appeal to me. But return to this note. I can return to this note. And then it says what I noticed, word families. What did I notice? Well, I don't remember. Righteousness carries so much accumulated theological weight that an English translation can begin interpreting a passage righteous or righteousness. It's got a notebook with all of these. Tell us about righteousness. Tell us about righteousness. I don't redo that. May I just say that this is a warm paper background? Just, just saying, just saying anyway. Yes. I never said I have very much. It is very chat GPT. I never said that. I never said anything else. Otherwise, I love warm paper. I'm a writer. Warm paper is very much for me. So shout out to the open AI team for designing an app that would appeal to me. But return to this note. I can return to this note. And then it says what I noticed, word families. What did I notice? Well, I don't remember. Righteousness carries so much accumulated theological weight that an English translation can begin interpreting a passage simply by using or avoiding the word. So I just, then I was like, here's a bunch of different places that use righteous, righteousness, but different translations will say right paths versus right direction or there's just a lot of different translation stuff that you can get into. And I try to capture all of that here. And then it tracks research that I need to follow up on. And this is my journal, my actual reading log, which is not very complete. So there's some, you see around the edges, there's some stuff here that isn't completely ready to go, but it's got so many different dimensions, and the layout, and I know what all the buttons do and why I would want to use them. And I haven't been able to develop, to build something this complicated with this little specification, because I was like, here's a file full of my Bible study stuff, build me an app, build me something that would be useful for me. And this is what I got. So it's not even, yeah, it's like I didn't have some extensive PRD or anything. And this is just what Astra thought I might like. So I just want to pull us back. Do you remember being in upstate New York, summer of 2025 around the GPT-5 launch, and you were trying to vibe code your own CMS for editorial. Mm-hmm. Talk to us about the difference between then and now. I mean, then I was up until midnight then or 2 AM the night before, frantically trying to make it work. And then I brought it online, and then it died. And then I brought it online again, multiple times, only to have to Theranos it at the very last minute and make it look like it was working on the demo day when it actually wasn't. And this, I barely even, I just gave it a prompt and sent it on its merry way. And then once the edit, once the draft was out of my hands this morning, in edit land, which is none, or CMS land, which is none of my business, I went back and looked, and I was like, oh, this is really cool. And so yeah, the difference between having to learn what a branch was and it being really important because something could go very, very wrong, and now it's just ridiculous. That's amazing. So we've got something that Allie wants to say. While Allie's talking, Katie, can you queue up your Hadestown demo as well? Because I feel like that's worth showing. Allie, whatever that is. I'm very excited. Katie, your last screen share, I just want to call out something. And this is, I've experienced this with previous GBT models and continue to experience it with Astro, which is annoying, which is that it's allergic to chunky paragraphs and few paragraphs. Every single thing gets its own line, a new indentation, a new bullet, a new italicized whatever. Kieran, your demo showed the same thing. This one's italics. This one's underlined. This one's small. This one's big. This one's bold. Bro, pick one to two options and just go back and forth. I think a copywriter with a design eye, if they were really trying to perfect something, would be annoyed by this default behavior that all the GPT models do, every single one of them. And so I find myself constantly providing almost like wireframes for copy design. Katie, any response? I mean, I was talking with Douglas about how we need compound marketing because copywriting is hard, and designing layouts is hard. And it makes sense to me that it's one of the things that is still an unsolved problem just because it involves theory of mind for the human, which is something, and understanding what kind of information, it's dealing in that weird ephemeral space between a meat brain and a compute and a zero and one brain, and making those two things line up so they work together is something that the models are still kind of figuring out in that. And you see that in things, in the more squishy things like design and copywriting and regular writing, I think. I agree. All right. Here it is. Hadestown. What is Hadestown? Tell us. Katie, Hadestown is the Tony award-winning musical. The other Tony award-winning musical whose title is one word that starts with H and ends with N that is not Hamilton. And it just came out on streaming. The pro shot of the London cast recording was just released, which is mostly the original Broadway cast. And so I've just been really fixated on Hadestown lately. And I particularly love the stage design and think it's really compelling. So I've had both 6, 6 Astra and 5.1 try to build the set just as my version of testing some of these cool 3D design concepts that you hear so much about. And so this is, this one is Astra's, and you can see it's got the balcony up here and the stairs down. We've got our band, our spot for the band on either side, and the cafe tables, and then the rotating stage. So this is pretty good. But the thing that you'll notice, the color of the background and the balcony is really contained to this one place. If you compare that to this, this is 5.1 fable 5.1 version, you'll see the balcony runs all the way around and the stair, the spiral staircase is on the wrong side. So the way that I'd articulate the difference between these two versions is like the, and I'm going to go back to the Astra one, the Astra one looks like somebody's actually seen, it looks like the model has seen the stage itself. It actually went and found visual references, where the fable one just looks like the model did some reading about what the set looks like. So I think that's how I would articulate the difference I perceive between their two approaches to this task. I love it. I do think that this is, this is a nice one where it shows there's such, there's so much cool detail in this one versus fables. And you can also see a world above, a world below. This is the sort of, it turns everything into a landing page and writes too much copy and too many buttons. And there's a net bar. Why is there a navigation bar here? What does that add? Because you can look at different parts of the, you can look at the, from different perspectives. No, I mean the bar on top, like the Hadestown logo, Broadway, Act one, like about the model, like why? That doesn't add anything. Doesn't matter what you ask, it will add a logo with the title and a navigation bar because that's what you need. All right, we're going to keep going. We are going to keep going. Let's, I just, I just found out that I, in our YouTube video talking about this model, I called it GPT 5.6 instead of GPT 6, which is just perfect. We shot this video in, less than an hour, early in the morning, and there, I have a can of Celsius here, and that's for a reason. So we're just going to have to, we're just going to have to own it unless we can cut it out. I don't know. Randy says no. You got to switch to a line. But I think we're, I think we're going to have to just own it, Randy. And just, it's funny. So, like, All right, we're going to keep going. We are going to keep going. I just found out that in our YouTube video talking about this model, I called it GPT 5.6 instead of GPT 6, which is just perfect. We shot this video in less than an hour, early in the morning, and I have a can of Celsius here, and that's for a reason. So we're just going to have to own it unless we can cut it out. I don't know. Randy says no. You got to switch to a line. But I think we're going to have to just own it, Randy. It's funny, so let's just do some funny phones of ourselves in the comments there, if we can. Cool. All right. Let's keep going. We are moving on. We're going to talk about writing now. My favorite topic. So what you should know about this model is it one-shotted its own vibe check. You should know Every is the only subscription you need to stay at the edge of AI. If you want to get more, we've got an entire vibe check written by Katie. It's 4,000 words, written start to finish at 7 AM this morning, because we only found out at 3 AM that the model was dropping. Look, it's so beautiful. It makes a little six. There's some really cool stuff happening here. But it's got all of our testing, all of our reviews from our whole team. We've got screenshots. You can compare it to Fable. You can compare it to your model of choice. And now we got to the writing. And the really interesting thing is, I saw at 3 AM that the model was dropping. I got some Slack messages and some texts, and I was like, oh fuck. And then I spent an hour worrying about how we were going to even get a vibe check out, because we only had a few hours. And I was making all these plans to be like, okay, we'll do it, we'll do a tweet. I definitely can get a tweet out, and I'll promise a real vibe check on Monday or something like that, because the model is actually not even live. It's only live for enterprises today. It's not even live for regular people until, I think, tomorrow, or they're figuring it out. And then I woke up at 7 AM after worrying, and Katie had a draft, or I don't know if it was seven, but it was pretty early. Katie had a full draft in the channel. And I just opened it and started reading it without looking particularly closely, and I was like, oh, this is a pretty good draft. Katie did a good job. I'm shocked she got this much written this well, with screenshots and all this stuff. This is just Astra. Astra just one-shotted this. It's pretty crazy. I'm going to go through some of the writing in here and talk about what I think is good. But Katie, anything you want to say first before I start going into detail? Just everyone get ready for what it's like to be on an editorial call with Dan. I mean, I'm going to be complimentary. There's some good stuff here. There's some good stuff here. Okay. So vibe check GPT 6 Astra is a show horse. So, a little bit critical of itself. And actually, one thing that's really interesting here that I like is it picked up a good line from Kieran. Kieran was talking about it as a show horse, because it's a little bit extra. And so it did a good job of finding something interesting and putting it at the top. This is not the headline. This is not the headline we picked, but it's plausible. Open AI's new model has striking visual taste and a knack for consulting work. Its ambitions still outrun its judgment. That's actually really good. I think it does a nice job of picking out more things that are true. I think striking visual taste is not exactly how I would have said it, but definitely knack for consulting work. Its ambitions outrun its judgment is interesting. I think it's a little bit too hard on itself, even in this subhead, but good. Okay. Astra GPT 6 Astra is very good at getting carried away. Open AI's new model, out today, builds beautiful things. I asked it for a 3D rendering of the Hadestown set and got a little theater with weathered green walls, hanging lamps, and wooden stage rings with lights. That's a really good opener. There's no X, not Y there. I think there's a little bit of a problem here where it's like, it's very good at getting carried away. Its new model builds beautiful things. These two sentences don't connect, which is a little bit of a problem. But it got, we often do, a really good punchy one-liner here and then a fatter paragraph here. It got that form right. And I think this sentence, I asked it for a 3D rendering of the Hadestown set and got a little theater with weathered green walls, hanging lamps, and a wooden stage ring with lights. If I wrote that sentence, I would be pretty proud. And I think it's kind of incredible that we're at a point where AI is writing stuff like this. And not only is it doing that, it is truly just going on for paragraphs, paragraph after paragraph, putting together a well-structured, look, it has the reach test in there with different options. It's got coding. It's got screenshots. I don't know. This is really, really impressive. Katie, anything in here you want to point out in particular? Well, since you mentioned screenshots, I will mention that not only did it drop the blank screenshots into this, when I set it up to use shots and follow our style guide from Notion to format them for the actual vibe check itself, it went and did that for me too. I didn't have to format a single image or write a single caption, which are the two things that I will put off till the absolute last minute and avoid like the plague. So that, and that's the dream, right? That's what we want AI to be doing, is taking those things that are annoying and awful off of our plates. And that's been one of our white whales for a while, is getting screenshot management done. And that's all down to computer use. I agree. And it makes you start to realize, I think we were in the pre-production for this because we were waiting for the monitor for a while, it makes you start to realize there's so much in a vibe check that is just about gathering information that's already out there and putting it into the right form. And then there are these other things that really have our taster, that have a little bit of the human flourish that you want to add that the model can't do. But just being able to start with this scaffold is so helpful, rather than, there's so much work. Talk to us about the kind of work that you'd have to do pre-Astra to get all the information in one place and then get it into a form like this. Yeah. So, for one thing, I would absolutely go through multiple steps in the process. I wouldn't just trust the model. I wouldn't give it just our Slack channels and the links to the vibe check, to the benchmarks that we've done, and let it have fun. I usually go through the step of compiling a finding summary, which I then review and make sure that everything makes sense in the section that it's being used in, and that we're spotting any contradictions or places where people disagree. And there's a lot of that thinking that goes on upstream of the process before you get down into managing the actual sentence-to-sentence copy. And that's a lot of sequencing, of getting the right example and then explaining why the example is significant. But then you kind of got to roll back up and provide the top-line analysis of what is going to be most interesting to the reader, that then, what is more and the links to the, to the vibe check, to the, to the benchmarks that we've done, and let it have fun. I usually go through the step of compiling a finding summary, which I then review and make sure that I, that everything makes sense in the section that it's being used in, and that we're spotting anything like contradictions or places where people disagree. And there's a lot of that thinking that goes on upstream of the process before you get down into managing the actual sentence-to-sentence copy. And that's a lot of sequencing of getting the right example and then explaining why the example is significant, but then you kind of got to roll back up and provide the top-line analysis of what is going to be most interesting to the reader, that then, that, that what is more you can't just start with an example. You have to frame up what is the significant thing that you're about to demonstrate with the, with the, with the, with the example. And that's one of the things that I'd have trouble with in drafting with past models. It would just always want to lead with the example, and okay, but I don't know why I'm reading about this. I don't know, that theory of mind for the reader and understanding what, what is the top-line thing the reader wants to know that this, that the, that what, that what comes next is going to convince them of, is something that historically I've had a really hard time getting out of models, but this one did a pretty good job. I love it. That definitely matched my experience. Ali, any writing tests or experience with this model as a writer that you want to share? So one thing that I have found with all previous models, open AI and Thropic or otherwise, is that you start to figure out the patterns of AI writing when multiple people are attempting to answer the same question. And so I'm just excited to see when we have a job posting, what the applications look like; when I have a social post, what the comments look like, because in our testing, it wasn't immediately obvious what those patterns would be in the way that they were very obvious with the word quiet repeated a thousand times in previous models or, previous cloud models, or it's not x then y. All of those, all of those patterns that we could recognize in fewer samples, I think we're going to need more samples to recognize the patterns, but I still think the patterns will be there. I agree. It's just one of those things where you change it, but the thing is not the specific things, like it's x, not y. It's not inherently bad. It's just being used too much in different contexts. And humans are just really good at picking that up. Mike, I know you've done a lot of work on the writing for the, in this model and as compared to other models. Anything you want to share? Yeah. So I can show you the, the blind writing test that it did, because I think that was, that, that surprised me actually. Please. Yeah. Let me, let me share. Yeah. Okay. Where are you? Here we go. At the stage. There you go. Yeah. Cool. So, first of all, as you can probably tell, Astra built this app because it's the same kind of design that you've seen in some of the other demos. But what I wanted to do is take a post that I actually did write with AI and then deconstruct it and then test it against a bunch of different models. So the way it works is I have a bunch of different comparisons here. This is from a real article that I worked on. There was 1200 words. What, what I did is I broke the article down into what is the, just the actual information in that paragraph. So this is in paragraph nine of 18, what happened. And then this is what happened in the paragraph right now that it's writing. And then this is what two different models wrote. So you can say, I prefer A, or I prefer B, and it's a fair test because they're both given the exact same brief and also the exact same context, which was me talking for half an hour about this specific topic. So, same brief, it's a completely fair test and is blind. So I don't know which of these models was which. So I did 50 comparisons today just to kind of see which model do I actually like. And if I click see my results, I was really surprised that it was Astra slightly edged ahead of Claude Fable 5.1. And that was a surprise to me because in some of the writing tests I did previously on my benchmark, I was like, yeah, it's okay. It's not that good compared to Fable 5.1. So, yeah, doing enough of these samples, I think made me realize that maybe I was being a little bit too judgmental on the model on first impression. Really, really interesting. And yeah, really interesting to see that, that grouping. It's just this generation of model, it's a bigger model. Both open and anthropic have figured out some things that just make the writing better. I know Allie, Allie has got a couple of things to say. First, I just, I love this. I love the way you guys test. I think everyone should follow and subscribe and do all the things that Dan Schipper tells you to do. And so just thanks for having me on and getting to brag about you all. I think the way you test things is great. I want to add one more thing just on theory of mind, which is that my team is spending more and more time teaching these AI models theory of mind. I feel like we're spending time building out, call it SOPs, call it skills, call it context docs, whatever, providing these things. We have an entire document whose title is just resonate. And it's just, how do you actually tell a compelling story? How do you tell it visually? How do you know all the things that Katie, you mentioned? And for what it's worth, fable 5.1 has performed better once given that theory of mind thing that we've built. So I, and Mike, you've inspired me. I'm going to do that same AB blind test. I love that. I look forward to being surprised by the results as well. But when given that foundation, fable 5.1 is currently outperforming, but still to be proven. Anyways, thank you so much for having me. You guys, you guys are a wonderful, fantastic team. And the cops are staring at me, and they have no idea how cool you guys are. I'll tell them after I get back. Thanks for joining. Good luck. Bye, guys. All right. So that's, that's our take on writing. Kieran, anything you want to add here? Anything you've been, you've been noticing on the writing side? I, no, but I will use it. I am using, I actually use GPT for writing using the compound writing plugin. So I'm very curious to see how this model does with the writing. So I'm definitely going to do that. But yeah, that's for writing. This is fun. All right. Well, let's keep going. We're going down the vibe check. I'm going to share my screen again. If you are here, welcome. We are every, we are the only subscription you need to stay at the edge of AI. If you want to stay at the edge, you need two things: education and equipment. On the education side, we've got a daily newsletter. We test all these models before they come out. Every dot T O, we've got a vibe check on Astra. We did a vibe check on fable 5.1 two days ago. And the day they come out, the second they come out, we have long-form reviews like this one. This is 4,000 words of testing from our entire team across coding, writing, design, knowledge work. We've got a ton of things already. We've already covered, we covered our reach test. Like, what are we, what are we really reaching for, for this model? What are we not reaching for? Because we've been using it for a little while. So we can say, this is what I use fable for. This is what I use Astra for in a way that will help you figure out what you want to use it for. We've got a daily newsletter. We test all these models before they come out. Every dot T O, we've got a vibe check on Astra. We did a vibe check on Fable 5.1 two days ago. And the day they come out, the second they come out, we have long-form reviews like this one. This is 4,000 words of testing from our entire team across coding, writing, design, knowledge work. We've got a ton of things already. We've already covered our reach test. What are we really reaching for for this model? What are we not reaching for? Because we've been using it for a little while. So we can say, this is what I use Fable for. This is what I use Astra for, in a way that will help you figure out what you want to use it for. We did a lot of stuff on coding already in this live stream. I showed this sick Waterloo thing. It's a reenactment of Waterloo, and I can, it's historically accurate. Apparently it's really sick. This is one prompt, one prompt, people. We've gone through writing as well. Astra actually wrote the first draft of its own vibe check for us. And we heard that it was launching at 3 AM, and by 7 AM we had a draft from Katie that was really just Astra. Now I think we're getting to some of the real meat of this live stream, this vibe check, which is knowledge work. I think the best thing, and you can tell OpenAI thinks this too in the way that they wrote their blog posts, computer use is sick on this model. And in particular, I think ChatGPT for Work slash Codex is currently the best harness for computer use because you can tell it to go do something and it'll go do it in the background. It has its own cursor. It uses your browser. It uses any app on your computer. Claude, for the longest time, if you asked it to use the computer, it would do this orange thing around your frame, around your screen. And then you would be like, are you using the computer or am I using the computer? They just, I think, fixed that, but still, I think ChatGPT for Work is just a better computer use harness right now. And Astra is the best model for this. It's very fast. In fact, let me just see if I can find this. It did our Fable 5.1 vibe check video. So this is a vibe check video of Fable 5.1. It's got 26,000 views. We published it two days ago. Astra edited this, which is pretty, pretty fucking crazy. Astra edited this. We also have a fantastic human editor, our head of social, Randy. Shout out Randy, but he basically just got all the raw footage, threw it into Premiere. Or actually, I don't even know if he threw it in. Did you throw it into Premiere? Okay. And then you had Fable cook. Yeah. He threw it in computer, and then he just had, not Fable, he then had Astra, Astra cook on it. And it's pretty amazing, pretty amazing. So I've been using it a lot as well for knowledge work-type tasks. Even if you have a Comcast bill or something, or you need to do some sort of, need to interact with customer support for anything, it's so good. Because they already have AI on their end, and then you can have AI on your end being like, I want a refund for this. And it just keeps going, and you don't have to worry about it. I had it sell stuff for me on Facebook Marketplace, post stuff and sell stuff, which GPD 5.6 was able to do, but I just think this model is much more reliable and much better at it. But if you really want to know what this model is good for and why, we've got a special, special segment from Mike Taylor. We're calling it Mike's checks. Mike, tell us about what a check is and what Mike's checks are, and then tell us how they can help us understand if this model is any good. Yeah, you're on mute, Mike. All right, I'm good. Yeah. So one of the first things I did when I joined Every, because I realized I'd be testing a lot of models, was I need to automate this in some way. And so what I tried to do was collect a bunch of different tasks that I would normally run when I got access to a new model, and then have a runner that basically ran all those tasks for me. I can just give it the model name, and then it's going to output all of the different files. And some of these are writing tasks. Some of these are PowerPoint tasks. There's one dashboard task, right? So we've been going back and forth on this, and we figured out an interesting way to do this. We wanted to automate the actual feedback that we give to the model, or give on the model, as well. So if you go into a task, and if you can see my screen, just gonna, yeah, just gonna click in. This is the dashboard task. Here we give it some NPS data from Every, and we say, hey, I want you to build a dashboard that shows this data. So it's a pretty short prompt. And the model will build that and take pictures. But what's really interesting is that we run these checks afterwards. So where these came from, this is after I had checked a bunch of the outputs of different models, like I have here, like Fable 5, Haiku 4.5, Fable 5.1, et cetera. I went through and I just gave a bunch of feedback and said, obviously the page needs to load. In some cases, some of the earlier models, the page didn't load. The NPS score needs to be correct. So there's a few things you can check with code. But a lot of the important checks are basically my taste, but captured in these checks. So one thing I really like is when it has readable typography or a distinctive design, and I've given it a definition of what I think a distinctive design is, and the model will go and check that. So the model's doing all these checks. I just need to come in, look, read it, and then also look at it and give additional feedback. So that's what we've done. We're actually going to roll this out more broadly and do this process with other people in the company so that everyone has their personal benchmark, and they can check new models without having to be constantly testing these things all the time because there's so many models coming out. We've had four new models this week. And I'm so excited about this. It's a whole new world for vibe checks. We've got checks now, and tell us what you found. Go through the checks for this model versus some of the other models. And I don't know if other people are finding this. I think your screen is slightly blurry to me. Are you guys able to read his screen? It's New York internet. We've managed to avoid the internet issues since we got a new Fios, but yeah, there's something. If I zoom in, it's fine. Yeah. Is that better? Yeah, this is better. Yeah. Okay, cool. So first thing I noticed was that it did a better job of design, and this is just the default design. I know, Dan, you don't like the warm paper as much as me and Katie, but I think this is a better job versus other outputs. I just don't like that it's the same. It always does warm. It's like, I'm going to make it make up a concept, and then it's just like warm paper. since we got a new Fios, but yeah, there's something. If I zoom in, it's fine. Yeah. Is that better? Yeah, this is better. Yeah. Okay, cool. So, first thing I noticed was that it did a better job of design, with, yeah, and this is just the default design. I know, Dan, you don't like the warm paper as much as me and Katie, but I think this is a better job versus other outputs. I just don't like that. It's the same. It always does warm. It's, I'm going to make it make up a concept, and then it's just warm paper. I just want some more concepts, exactly. I want that purple gradients back now because it's fresh now. Yeah. So, yeah, we have this is fable 5.1, for example, and this is one of the things that fable got marked down on, was that it was a very boring design and a very boring title. But then if you look at aster, I think this is much nicer. But yeah, you're right. It is very samey. Just to throw it out there, this is a completely different one. This is Grok, and Grok, I think, actually has done the best in recent weeks on this specific task, because it has the narrative, but it's still a little bit warm paper, but I think it looks a little bit more distinctive. See, Grok is just this sleeper model that no one's talking about because they've sucked for so long, but now that they have cursor, a lot of people are just quietly using Grok all the time. Yeah. Second most used model I used. I spent more tokens on Grok yesterday than fable, which is funny. That's fascinating. Grok is in the top cohort for the writing benchmark with fable five and Astra. Yeah. See, under discussed. You guys, everyone should be using, I haven't used Grok. I should be using Grok. Everyone should be trying Grok. That's what Mike's check says. That's what Katie says. That's what Karen says. Yeah. Keep going, Mike. Yeah. So, dashboard, it did pretty well. One thing it didn't do that well is this writing task specifically, which is what made me surprised when it did a good job in the blind test. So, just wanted to point out here that one of the checks that I run with this article, this specific task is, I gave it a bunch of context. I gave it a tweet that I'd written. I talked for half an hour about a specific thing I was talking about with consulting, like how to get your team to use AI. That was something I was doing a lot of, helping people with. And I had Claude write this up. It was 4.8, Opus 4.8 originally that wrote this post. But I basically give the same brief to multiple models, and it didn't do a great job. The way I can describe it is it sanded down all of the rough edges and made it a bit boring. So, and it didn't keep a lot of the stats that I'd given it, where I'm a big stats guy. I love when it keeps the data in there. Little things, like it didn't, just link to the tweet instead of showing the tweet. I think it's really important to actually show the tweet so people don't have to click away. So yeah, it got marked down on a lot of those things. That is fascinating. Yeah. So, obviously, that's why you gotta test a lot of things, right? It did really great on this benchmark, which is, given a ton of context from a bunch of interviews we did that were anonymized for a client, can we put together a training plan for that client? And if you look at the checks here, the topics were really deep. It didn't just repeat what was in the original descriptions. It was actually supported by the interviews. It didn't hallucinate anything. So it got a hundred percent of my checks correct here. So I think it did a really good job of pulling out what the correct themes are. And because I remember this specific case, I'm looking at it and like, yeah, yeah, I would have a hundred percent sent this to the client. So cool. Mike, remind me, are these public? Yes. Yeah, we do actually have, so you can see the website here. It's Mike's checks.every.to. And we just got this up today. So yeah, you can go on here. I think you linked to it from the original post, from the vibe check. You should check it out. P, not, sorry. Pietro is joining. Mike's checks.every.to. There's some cool stuff in here. Okay. Keep going. Keep going, Mike. Yeah. And then another one I wanted to point out, again, another design one. Again, it's the screen slash one paper thing. But I think they did a really good job on this task. What this task is is basically clone type form, but make it so it interviews the person filling instead of forcing them to fill in a survey. So it's a vibe coding benchmark. I'm trying to look for, like, does it actually complete the full thing? And does it come up with a good design for this? And it just completely nails this. And this was the hardest benchmark that I had a couple of months ago, and now it's saturated by Astra. Wow. Can we see 5.5.1? Yeah. So, yeah, fable 5.1 almost got it there. The thing it got marked down on is this vibe coded purple design. I very specifically said, I don't want this to look like vibe slop. And I think, again, these are my subjective checks. So the point of this benchmark is really just, is this a model that I personally want to use rather than, is this the best model in the world? Because you could obviously argue one way or another depending on what you like, and everyone can have their own checks as they do. And we can debate which checks are the right checks. Okay. Anything else on this, Mike, that you want to go through? I would say it did a pretty good job of PowerPoint. I just wanted to share it because we talked about this with the fable. It did this visual design thing almost perfectly, and then it just added this weird so close squiggly thing, which is annoying. So the reason why this is a big deal, even though it seems very subtle, is that a lot of models get this wrong. I don't know if we've got, so actually, again, this is another one that Grok does really well. Really? Yeah. Grok made its own SVGs here. Wait, I want to see. Let's find that slide. Yeah. Yeah. So, let me see. So, this is, I want to see. Yeah. I'm trying to remember. It's usually earlier. Yeah. I'm trying to remember which slide, because Grok ended up doing I just wanted to share it because we talked about this with the Fable. It did this visual design thing almost perfectly. And then it just added this weird, so close squiggly thing, which is annoying. So the reason why this is a big deal, even though it seems very subtle, is that a lot of models get this wrong. I don't know if we've got, so actually, again, this is another one that Grok does really well. Really? Yeah. Grok made its own SVGs here. Wait, I want to see. Let's find that slide. Yeah. Yeah. Let me see. This is, I want to see. Yeah. I'm trying to remember. It usually earlier. Yeah. I'm trying to remember which slide because Grok ended up doing a bunch more slides, but I just wanted to show, by the way, Grok made a lot of its own SVGs here, and they're beautiful. I think they're really every ish actually. Yeah. It did a good job. It's really on brand. So yeah, again, shout out to Grok. This isn't a Grok appreciation channel, but we're getting there. Yeah. So yeah, I'm trying to find that slide. Maybe if I, yeah, I can't find it here, but let's look at this Soul one because I know that that was one way I think it kind of messed up. Yeah. Yeah. So this is Soul, for example, which again is an amazing model. And just a couple of weeks ago, I was singing its praises, but it gets all the, it's super janky, right? It has no spatial awareness. So this is an example of a model doing it wrong. And Astra is almost there, right? It's very nearly there. The one that didn't nail this was Fable 5.1 just recently. So that was here, yeah, it got, it got, it got it like, yeah. Again, this is intern-level design on a PowerPoint, but even that, it feels like we've come very far. Absolutely nailed the arrows bench, which is itself much more comprehensive and difficult than the Pelican bench. Yes. Shout out to Simon Wilson, but Mike's got you, Mike's got you beat there. Well, I mean also economically more important, right? Because most of the world runs on PowerPoint, but Pelicans are very close to our heart. That's true. No, I'm joking. But I think the spatial, they're both testing the same thing, right? The spatial awareness was such a terrible failure from most models for a really long time. And it does feel like we're making a lot of progress there. And you can see that in both of the benchmarks. I totally agree. All right. I love this. I spend a lot of time in Mike's checks these days. So you all should too. We've got another set of demos from Kieran. Kieran, give us your, you do this. Okay. I got to say before you start, though it's all the rage right now to do a whole 3D environment as part of your vibe check, Kieran has been doing this for a long time. He has been doing it since before 3D environments were really even possible with these models. And now, I mean, we should go back and show a Sonnet 3.5 3D environment because we have them. But anyway, Kieran, tell us about your, this is the cozy island benchmark. Tell us about your cozy island. Yeah. So cozy island is just make a cool cozy island. And what I want to learn from this is, how is this different than other models? And it's not even about being good or bad because every model is very good. It's just, how are they different? And let's look at this, for example. Clearly you see a website header, a little Haven, which is interesting, and it looks like a cool website more than anything else. If you compare it to Fable 5.1, you could say, oh, this is boring. But it's also more focused, and the detail, the animations, for example, are a little bit more refined. But you can say here, yeah, but here the details are really cool. And there's an aesthetic to it. And you can see there's a different color palette being used as well. And there's all these things in the screen. Why are all these things here? But then you can say, oh, but it's actually cool because you can go to night in the version here, and you can go to daylight. You can even transition between things. And you could argue that's really neat and really cool. And that's not good or bad. This is one of the things. I think this is delightful, but the whole text and thing, I just don't like that. Why do I need that? It doesn't add anything. And back to here is maybe cool. Bigger, full screen. I mean, there are so many things here. I can even take a post, save a postcard, and it downloads to PNG. And there are sounds. I don't know if you hear, but there's a wave, and there's so much stuff in here. But for me, I like models that are just doing something and then inspire me to think of the next thing. Fable leaves space for IDs and the human a little bit more, and Astra is just like, let's go, and more as inspiration. It's more generative. This is how far it's, it's more like that. You can see the same for this one, which is a breathwork app, app again. It looks like websites. I'm already tired of this design because every time I see a design from anyone else, it looks the same. There's a logo with a title. Yeah. And it's almost like it's an artist, and they've put it in a sweatshop for making landing pages and tortured it until that's all it can do. Yeah. Yeah. It feels like they're over, they're over saturated, the refills of learning, overfitted on specific landing pages and specifically this kind of landing page. And you can see the difference here. This is Fable. It does it look better design-wise? Maybe it's more boring, but also it leaves more space. And I like it. This, you get tired of quicker. This is more like fast food. It's amazing. And here you're like, oh, this is like fine dining a little bit. It's a little bit like that, where it's like, I love it. It's hated, but it grows on you. And here you get this amazement first, but then you're like, I'm kind of tired of it. So there's a new level of slop, where it's very clickbaity, very viral probably. And so yeah, but I have to say, this is a landing page, and this is the best-looking landing page from any model that I've seen. This is really like, yes, it's maybe sloppy, but it looks really good. I really like this page, and I've not seen any other model do something like this. So there's a flip side to it. If you need that, it's probably very good at it, but also if you don't want it, it kind of tries to do it, which is interesting. Totally. We were talking about, actually go back to that landing page. We were talking about, okay, what is the new slop, and what are the markers of this model? If you scroll down a little bit, it loves, you bring the spark, we'll keep it catch. It loves two-line subheads or headlines where the second line is italicized and they're both hands with periods. If you see that, it's Astra, or really also dbt56 does this too, but that's perfect. other model do something like this. So there's a flip side to it. If you need that, it's probably very good at it, but also if you don't want it, it tries to do it, which is interesting. Totally. We were, we were talking about, actually, go back to that landing page. We're talking about, okay, what is the, what is the, what is the new slop, and what are the markers of this model? If you scroll down a little bit, it, it loves, "you bring the spark, we'll keep it catch." It loves, two-line subheads or headlines where the second line is italicized, and they're both hands with periods. If you see that, it's Astra, or really also dbt56 does this too, but that's perfect Astra. And that's just the kind of thing where there's something else when you, when we talk about this is AGI, there's something else about our intelligence that you see something amazing like this the first day. And then two days later, you're like, I recognize this. I know what this is. It's not that interesting anymore. And humans aren't like that. And I think that that's one of the core things that's, that's missing when we talk about AGI versus not AGI. It's not just about, can you do this amazing? Can you make this landing page? It's like, can you do something consistently where I can't look at it and be like, I know where this came from? I'm not surprised by this, or I am surprised by this. All right, Kieran, you have some more stuff you want to show. Yeah, because everyone was doing blender demos. I did a blender demo too. So this is, I just, one, show it a blender scene. I said, can you make a vineyard and something, something, and then render it, and it rendered it. You should know Kieran's a big wine. Yeah. I love, I love wine. And it rendered it. That's really sick. It's kind of cool. And I didn't do anything. So, yeah, so that is cool. And it, it used image gen to generate the textures on the, on the, the, on the, the couch and everything. So all the textures come from it generating as well, didn't use any assets. The backdrop actually is fake. It just plumped an image generated on the background. It's a fresco. It's like a, yeah, it's just a plate. If you can see here, there's a plate, and it just projected the, the image on there, but the rest is real, and it's cool. You can go around. So blender and other open source things will suddenly, yeah, you will explore other things that you can do, especially with MCPs and computer use, which, which is really cool. I love it. All right. So that is, I believe, everything that we, we wanted to talk about today. Is there anything, anything, any topic that we have not covered, anything that we tested that we need to talk about, anything else to share with this crew before we head off? I had one anecdote for computer use. I, I was being really lazy, and my, my daughter has mild peanut allergies, and my school sent me this big form to fill in. And I asked it to fill in the form, and it refused, like seven times, and I had to do my best prompt injection to get it to do it. And at the end, I was like, oh man, this would have been easier just to fill out myself. So, so I would say that's maybe one big blocker on computer use. A lot of the stuff I want it to do is administrative-type stuff. And if it refuses, I don't know, to do your taxes or whatever, that, that, that's a concern or a limiting factor potentially. Totally. Yeah. Definitely got some refusals. Some, I think it's some cyber safeguard-type stuff, which obviously they had a big incident, and it looks like from their model card and all that kind of stuff that a lot of that is solved, and it's going to come at the cost of some usability for a while. So it's, it's worth knowing that when you try this model, it may do that. Katie, it doesn't seem like you had something you wanted to share. I just wanted to share that if you want even more of the Every team's impressions on both GPT-6 Astra and Fable 5.1 Soul, you should subscribe to Every and join us tomorrow for Camp! You said? For Camp! Yeah! Camp! Camp! Camp! Camp! So as part of the Every subscription, again, Every is the only subscription you need to stay at the edge of AI. As part of the Every subscription, we do regular camps, which are live streams like this one, but they're only for paying subscribers. We all get on a Zoom together. We all chat. And the camps are usually on a topic where we go deep into what our workflows are like and how we're using these tools, and everyone gets to share and meet each other. And it's really awesome because we recruit people that love AI. We love to really use it for work and doing great work. And there's just not that many people like that out in the world. And so this is like a little sharing circle of people who want to share all that stuff with each other and, and, and learn and learn from each other. And so camps are super fun. Tomorrow we are going to be going in detail into Fable versus GBT 5.1. And it's going to be really fun. I'm so sorry for the rhyming, but I had to. And, yeah, it's going to be amazing. You should, you should check it out. every.to slash events can tell you more about the camps. We've got courses coming up. We've got more model reviews coming out. We've got our own agent, the Every agent, which is amazing. It's all bundled into one subscription, the Every subscription, every.to, the only subscription you need to stay at the edge of AI. Thank you for joining. Please check out Astra. Let us know what you think. Let us know on Every, in the comments, anywhere on X, and remember, stay hydrated, everybody. Keep yourself and your agent liquid cooled. We will see you next time. Um, well, since you mentioned screenshots, I will mention that not only did it do the screen, like drop, drop the blank screenshots into this, when I set it up to use shots and follow our style guide from notion to format them for the actual vibe check itself and went and did that for me too. Like I didn't have to form format a single, I didn't have to form format a single image or write a single caption, which are the two things that I will put off till the absolute last minute and avoid like the plague. Um, so that, and that's the dream, right? That's what we want AI to be doing is taking those things that like are annoying and awful off of our plates. And that's been like one of our white whales for a while is getting screenshot management done. And that's all down to computer use. I agree. And it makes you start to realize, I think we were in the pre-production for this, because we were waiting for the monitor for a while. It makes you start to realize like, there's so much in a vibe check that is just about gathering information that's already out there and putting them into the right, putting into the right form. And then there's like these other things that, that really have our taster, like that, you know, have a little bit of like the, the human flourish that you want to add that the model can't do, but just being able to start with this scaffold is so helpful rather than like, there's so much work. Talk to us about the kind of work that you'd have to do pre Astra to get all the information in one, in one place and then get it into a form like this. Yeah. So for one thing, I would absolutely go through multiple steps in the process. I wouldn't just trust the model to like, I wouldn't give it just like our Slack channels and the links to the, to the vibe check, to the, to the benchmarks that we've done and like, let it have fun. Like I usually go through the step of like compiling a finding summary, which I then review and like, make sure that I, that everything like makes sense in the section that it's being used in. And like that we're spotting anything like contradictions or places where people disagree. And like, there's a lot of that thinking that goes on upstream of the process before you get down into like managing the actual sentence to sentence copy. And that's a lot of sequencing of getting the right example and then explaining why the example is significant, but then you kind of got to like roll back up and like provide the top line analysis of like, what is going to be most interesting to the reader that then that, that like, you know, what is more like, you can't just start with an example. You have to frame up like, what is like the significant thing that you're about to demonstrate with the, with the, with the, with the, um, example. And that's like one of the things that I'd have trouble with in drafting with past models is like, it would just always want to lead with the example and like, okay, but I don't know why I'm reading about this. Like, I don't know, like it, like that theory of mind for the reader and like understanding, like what, what is the top line thing the reader wants to know that this, that the, that what, that what comes next is going to convince them of, um, is something that like historically I've had a really hard time getting out of models, but this one did a pretty good job. I love it. That, that definitely matched my experience. Ali, any writing tests or experience with this model as a writer that you want to share? So one thing that I have found with all previous models, open AI and Thropic or otherwise, is that you start to figure out the patterns of AI writing when multiple people are attempting to answer the same question. And so I'm just excited to see like when we have a job posting, what the applications look like when I have a social post, what the comments look like, because in our testing, it wasn't immediately obvious what those patterns would be in the way that they were very obvious with the word quiet repeated a thousand times in previous models or, you know, previous cloud models, or it's not x then y. Like all of those, all of those patterns that we could recognize in fewer samples, I think we're going to need more samples to recognize the patterns, but I still think the patterns will be there. I agree. It's, it's just one of those things where you change it, but like the, the thing is not the specific things, like it's x, not y, it's not inherently bad. It's just like being used too much in the, in different contexts. And humans are just really good at picking that up. Mike, I know you've done a lot of work on the writing for the, in this model and as compared to other models, anything you want to share? Yeah. So I can show you the, the blind writing test that it did. Cause I think that was, that, that surprised me actually. Please. Yeah. Let me, let me share. Yeah. Okay. Where are you? Here we go. At the stage. There you go. Yeah. Cool. So, um, first of all, uh, as you can probably tell, uh, Astra built this app, uh, because it's the same kind of design that you've seen, uh, some of the other demos. Uh, but what I wanted to do is take a post that I actually did right with AI and then deconstruct it, uh, and then test it against a bunch of different models. Uh, so the way it works, uh, is, um, I have a bunch of different comparisons here. Uh, this is from a real article that I worked on. There was 1200 words. Um, what, what I did is I broke the article down into like, what is the, just the actual information in that paragraph. So this is in paragraph nine of 18, what happened. And then this is what happened in the paragraph right now that it's writing. Uh, and then this is what two different models wrote. So you can say, I prefer a, or I prefer B and it's a fair test because they're both given the exact, the same brief and also the exact same context of, uh, which was like me talking for half an hour about this, uh, specific topic. Uh, so, uh, so same brief, you know, it's completely fair test and is blind. So I don't know which of these models, uh, was which, um, so I did, um, 50 comparisons today, uh, just to kind of see which model do I actually like. Um, and if I click, uh, see my results, uh, I was really surprised that it was Astra, uh, slightly, uh, edged ahead of Claude Fable 5.1. And, uh, that was a surprise to me because, uh, in some of the writing tests I did previously on my benchmark, I was like, yeah, it's okay. It's not that good compared to Fable 5.1. Um, uh, so, uh, yeah, uh, doing enough of these samples, I think made me realize that maybe I was being a little bit too judgmental on the model, uh, the first on first impression. Really, really interesting. And yeah, really interesting to see that, that grouping, it's just like this generation of model, it's a bigger model, both open and an anthropic have figured out some things that just make the writing better. Um, I know Allie, Allie has got a couple of things to say. I first, I just, I love this. I love the way you guys test. I think everyone should follow and subscribe and do all the things that Dan Schipper tells you to do. Uh, and so just thanks for having me on and getting to brag about you all. Um, I think the way you test things is great. I want to add one more thing just on theory of mind, which is that my team is spending more and more time teaching these AI models theory of mind. Like I, I feel like we're spending time building out, call it SOPs, call it skills, call it context docs, whatever, like providing these things. Like we have an entire document whose title is just resonate. And it's just, how do you actually tell a compelling story? How do you tell it visually? How do you know, all the things that Katie, you mentioned, um, and for what it's worth, fable 5.1 has performed better once given that theory of mind thing that we've built. So I, and Mike, you've inspired me. I'm going to do that same AB blind test. I love that. I, I, uh, look forward to being surprised by the results as well. But when given that foundation, fable 5.1 is currently outperforming, but still to be proven. Anyways, thank you so much for having me. You guys, you guys are wonderful, fantastic team. And, and the cops are staring at me and they have no idea how cool you guys are. I'll tell them after I get back. Thanks for joining. Good luck. Bye guys. All right. Uh, so that's, that's our take on writing. Um, Kieran, anything you want to add here? Anything you've been, you've been noticing on the writing side? Uh, I, uh, no, but I will use it. I am using, I actually use GPT for writing using the compound writing plugin. So I'm, I'm very curious to see how this model does, uh, with the writing. So I'm definitely going to do that. Um, but yeah, that's for writing. This is fun. All right. Well, let's keep going. We're going down the vibe check. Uh, I'm going to share my screen again. If you are here, welcome. We are every, we are the only subscription you need to stay at the edge of AI. If you want to stay at the edge, you need what you need two things, education and equipment on the education side. We've got a daily newsletter. We test all these models before they come out. Every dot T O we've got a vibe check on Astra. We did a vibe check on fable 5.1, two days ago. And the day they come out, the second they come out, we have long form reviews like this one. This is 4,000 words of testing from our entire team across coding, writing design, knowledge work. We've got a ton of things already. Uh, we we've already covered, we covered our reach test. Like what are we, what are we really reaching for, for this model? Uh, what are we not reaching for? You know, cause we've been using it for, uh, for a little while. So we can say, this is what I use fable for. This is what I use Astra for in a way that will help you figure out what you want to use it for. We did a lot of stuff on coding already in this live stream. Um, I, I showed this sick Waterloo, uh, thing. It, it basically, it's a reenactment of Waterloo and I can like, it's historically accurate. Apparently it's, it's really sick. This is one prompt, one prompt people. Uh, we've gone through writing as well. Astra actually wrote the first draft of its, um, of its own vibe check for us. And we, you know, we heard that it was launching at 3 AM and by 7 AM we had a draft from Katie that was really just Astra. Um, now I think we're getting to some of the real meat of this live stream, this vibe check, which is knowledge work. I think the best thing, and you can tell, you can tell the open AI thinks this too, in the way that they wrote their, their blog posts, like computer use is like sick on this model. Um, And in particular, I think chat to PT for work slash codex is currently the best harness for computer use because you can tell it to go do something and it'll go do it in the background. It has its own cursor. It uses your browser. It uses any app on your computer. Claude, uh, for the longest time, if you asked it to use the computer, it would like do this orange thing around your frame around your screen. And then, uh, and you would be like, are you using the computer or am I using the computer? Uh, they just, I think fixed that, but, but still, I think just chat to be do for work is just a better computer use harness right now. And Astra is the best model for this. It's very fast. It, it, it, in fact, like, let me just see if I can find this. It, um, did our fable 5.1, um, vibe check video. So this is a, this is a vibe check video of fable 5.1. It's got 26,000 views. We published it two days ago. Astra edited this, which is pretty, pretty fucking crazy. Astra edited this. Um, we also have a fantastic human editor, our head of social, Randy, um, shout out Randy, but, uh, he basically just got all the raw footage, threw it into premiere. Um, or actually I don't even know if he threw it in. Did you throw it into premiere? Okay. And then you had fable cook. Yeah. He threw it in computer and then he just like had not, not fable. He then he had Astra Astra cook on it. And it's pretty, it's pretty amazing, pretty amazing. Um, so I've been using it a lot as well for, for knowledge work type tasks. I mean, even like if you have a Comcast bill or something, or you need to like do some sort of, you know, need to like interact with customer support for anything. It's so good. Cause they're already, they already have AI on their end and then you can have AI on your end being like, I want a refund for this. And it, and it just keeps going and you don't have to worry about it. I haven't sell stuff for me on Facebook marketplace, like post stuff and sell stuff, which GPD 5.6 was able to do, but I just think this model is much more reliable and much better at it. Um, but if you really want to know what this model is good for and why we've got a special, special segment from Mike Taylor. Um, we're calling it Mike's checks. Um, Mike, tell us about what a check is and what, what Mike's checks are, and then tell us how they can help us understand if this model is any good. Yeah, you're on, you're on mute, Mike. All right, I'm good. Uh, yeah. So, uh, one of the first things I did when I joined every, cause I realized I'd be testing a lot of models, uh, was I need to automate this in some way. And, uh, so what I tried to do was collect a bunch of different tasks that I would normally run, uh, when I got access to a new model, um, and then have a runner that basically ran all those tasks for me. I can just give it the model name, uh, and then it's going to output all of the, uh, different files. And, and, uh, you know, some of these are writing tasks. Some of these are PowerPoint tasks. Uh, there's one dashboard task, right? So, uh, we've been going back and forth on this and we figured out an interesting way to do this. Um, we, uh, we wanted to automate the actual, uh, kind of feedback that we give to the model or give on the model as well. So, uh, if you go into a task, uh, and if you can see my screen, uh, just gonna, yeah, just gonna click in, uh, this is the, uh, dashboard task. It's like, here we give it some NPS data from every, and, uh, we say, Hey, I want you to build a dashboard that shows this data. So it's a pretty short prompt. Um, and, uh, the model, you know, we'll build that and take pictures, but what's really interesting is, uh, that we run these checks afterwards. So where these came from, um, this is, uh, after I had checked a bunch of the outputs of different models, like I have here, like Fable 5, Haiku 4.5, Fable 5.1, et cetera. Um, I went through and I just gave a bunch of feedback and said, uh, you know, obviously the page needs to load in some cases, some of the earlier models, like the page didn't load, uh, the NPS score needs to be correct. So there's a few things you can check with code. Uh, but, uh, a lot of the important checks are, uh, basically my taste, uh, but captured in these checks. So, um, one thing I really like is when it has readable typography or a distinctive design, and I've given it a definition of what I think a distinctive design is, uh, and the model will go and check that. Um, so the model's doing all these checks. I just need to come in, look, read it, uh, and then also look at it and kind of give additional feedback. Um, so, uh, so that's, that's what we've done. We're actually going to roll this out more broadly, uh, and do this process with other people in the company so that everyone has their personal benchmark and they can, um, you know, check new models without having to like be constantly testing these things, uh, all the time because there's so many models coming out. Like we've had four new models this week. And I'm so excited about this. It's a whole new world for vibe checks. We've got checks now and tell us what you found. Give us, you know, go through the checks for this model versus some of the other models. And I don't know if other people are finding this. I think your screen is like slightly blurry to me. Are you guys able to read his screen? It's New York internet. We've managed to avoid the internet issues, uh, since we got a new Fios, but yeah, there's something. If I, if I zoom in, it's fine. Yeah. Is that better? Yeah, this is better. Yeah. Okay, cool. Um, so, uh, first thing I noticed, uh, was that, uh, it did a, uh, a better job of design, uh, with, yeah, and this is just the default design. I know, Dan, you, you don't like the warm paper as much as me and Katie, but, uh, but I, I think this is a better job, um, versus, uh, other, other outputs. I just don't like that. It's the same. It always does warm. It's like, I'm going to make it make up a concept and then it's just like warm paper, you know, like, I just want it. I want some more concepts, you know, exactly. I kind of want, I, I kind of want that purple gradients back now because it's fresh now. Yeah. So, um, yeah, we have like, this is fable 5.1, for example, and this is one of the things that fable got marked down on, uh, was that it was a very boring design, um, uh, and, and a very boring title. Uh, but then if you, if you look at aster it's, it's, uh, I think this is much nicer. Um, but yeah, you're right. It is very samey. Um, just to kind of throw it out there. Like, uh, this is a completely different one. This is Grok, uh, and, and Grok, I think actually has done the best in recent, uh, in recent weeks, uh, on, on this specific task, um, because it's, it has the narrative, but it also, it's still a little bit warm paper, but, um, I think it looks a little bit more distinctive. See, Grok is just this sleeper model that no one's talking about because they've sucked for so long, but now that they have cursor, a lot of people are just quietly using Grok all the time. Yeah. Second most used model I used, I spent more tokens on Grok yesterday than fable, which is kind of funny. That's fascinating. Grok is in the top cohort for the writing benchmark with fable five and Astra. Yeah. See, under discussed, you guys, everyone should be using, I haven't used Grok. Like I should be using Grok. Everyone should be using trying Grok. Uh, that's what Mike's check says. That's what, that's what Katie says. That's what Karen says. Yeah. Um, keep going, Mike. Yeah. So, uh, so dashboard, it did pretty well. Um, well, one thing I didn't do that well is this writing task specifically, which is what made me surprised, uh, when, uh, it, uh, it did a good job in the blind test. Um, so, uh, just wanted to point out, uh, here, uh, that, um, one of the checks that I run, uh, with this article, this, uh, this specific task is, um, you know, I, I gave it, uh, a bunch of context. I, you know, gave it a tweet that I'd written. I, I like talked for half an hour about, uh, a specific thing I was, I was talking about with, with consulting, like how to get your team to use AI. Uh, that was something I was doing a lot of to helping people with. Um, and I, I wrote, yeah, I had, I had Claude write, write this up. It was, uh, 4.8, Opus 4.8 originally that wrote this post. Um, but I basically give the same brief to multiple models and, uh, it didn't do a great job. It, the way I can kind of describe is it is sanded down all of the rough edges and made it a bit boring. Uh, so, uh, you know, and, and like, it didn't, you know, it didn't keep a lot of like the, the stats that I'd given it where I, I, I'm a big stats guy. Like I love, uh, when it keeps the data in there. Um, it, like little things like it didn't, you know, just link to the tweet instead of showing the tweet. I, I think it's really important to actually show the tweet so people don't have to click away. Um, so, so yeah, it kind of got marked down on, on a lot of those things. Um, that is fascinating. Yeah. So, so that, you know, I just, obviously that's why you gotta test a lot of things, right? Um, it did, it did really great on, on this benchmark, which is, uh, you know, given a ton of context from a bunch of interviews we did that were anonymized, uh, for a client, uh, can we put together a training plan for that, uh, client? And, uh, if you look at the checks here, uh, like the topics were really deep, uh, it didn't just repeat what was in the original, uh, descriptions, uh, it was actually supported by the interviews. Like it didn't hallucinate anything. Uh, so it got like a hundred of my, a hundred percent of my checks correct here. Uh, so I, I think it, yeah, it did, it did a really good job of pulling out what the correct themes are. And because I remember this specific case, I'm looking at it and like, yeah, yeah, I would have a hundred percent sent this to the client. So cool. Um, Mike, are these, remind me, are these public? Yes. Uh, yeah, we do actually have, uh, so you can see the website here. It's Mike's checks.every.to. And we just got this up today. So, uh, yeah, you can, uh, you can go on here. I think you linked to it from the, from the original post, uh, from the vibe check. You should check it out. P, uh, not, sorry. Pietro is joining, uh, Mike's checks.every.to. There's some cool stuff in here. Okay. Keep going. Keep going, Mike. Yeah. Um, and then another one I wanted to point out again, another design one. Again, it's the screen slash one paper thing. Uh, but, uh, I think they did a really great, good job on this task. What this task is, is basically clone, uh, type form, uh, but, uh, make it so you can, it interviews the person filling instead of like forcing them to fill in a survey. Uh, so it's like, it's a vibe coding benchmark. It's, uh, I'm trying to look for like, does it actually complete the full thing? Um, and does it, I come up with a good design for this and, uh, it just, it just completely nails this. Uh, and this was the hardest benchmark that I had a couple of months ago and now it's saturated by Astra. Wow. Can we see 5.5.1? Yeah. So, um, yeah, fable, uh, 5.1 almost got it there. The thing it got marked down on is, uh, this, you know, vibe coded kind of purple design. Uh, I, I very specifically said, like, I, I don't want, uh, the, you know, I don't want this to look like vibe slop. Um, and, uh, and, you know, I think, uh, again, these are my subjective checks. So, uh, the, the point of this, uh, benchmark is really just, you know, is this a model that I personally want to use rather than like, is this the best model in the world? Cause you could obviously argue one way or another depending on what you like and everyone can have their own checks as they do. And, uh, we can debate which checks are the right checks. Um, so, okay. Anything else on this, Mike, that you want to go through? Um, I would say it did a pretty good job of PowerPoint. Uh, the, uh, I just wanted to share it cause we talked about this with the fable. Um, it did this visual design thing almost like perfectly. And then it just added this like weird, so close squiggly thing, which is annoying. So like the reason why this is a big deal, even though it seems very subtle is that a lot of models get like this wrong. Um, I don't know if we've got, uh, so like actually, again, this is another one that Grok does, uh, really well, like really? Yeah. Grok made its own SVGs here. Wait, I want to see. Let's find that slide. Yeah. Yeah. Um, so, uh, let me see. So, uh, this is, uh, I want to see. Yeah. I'm trying to remember it usually earlier. Yeah. I'm trying to remember which slide cause Grok ended up doing a bunch more slides, but I just wanted to show by the way, like, Grok made a lot of its own SVGs here and they're beautiful. I think they're really every ish actually. Yeah. It did a good job. It's really on brand. Um, uh, so, so yeah, like again, shout out to Grok. This isn't like a Grok appreciation channel, but like we're getting there. Yeah. So, uh, yeah, I'm trying to find that slide. Maybe if I, uh, yeah, I can't find it here, but, um, we let's look at this soul one. Cause I know that that was, uh, one way, um, I think it kind of messed up. Uh, yeah. Yeah. So like this is soul for example, uh, which, which again is an amazing model. And like just a couple of weeks ago, I was singing its praises, but it gets like all the, it's, it's like super janky, right? Like it has no spatial awareness. Um, so this is an example of a model doing it wrong. And, uh, you know, Astra is like, it's almost there, right? Like it's, it's like, it's very nearly there. Um, the, the one that didn't nail this, uh, was, uh, fable 5.1, uh, just recently. Uh, so that was, uh, here like, yeah, it got, it got, it got it like, yeah. I mean, again, like, you know, this is like intern level design, uh, on a PowerPoint, but like even that it, it feels like we've come very far. Absolutely nailed the arrows bench, which is itself. It's much more comprehensive and difficult than the Pelican bench. Yes. Um, shout out to Simon Wilson, but, but Mike's got you, Mike's got you beat there. Well, I mean also like economically more important, right? Because most of the world runs on PowerPoint, but you know, Pelicans are very close to our heart. That's true. No, I'm joking. But like, I think, uh, the spatial is, they're both testing the same thing, right? Like the spatial awareness was such a terrible, uh, like failure from most models for a really long time. And it does feel like we're making a lot of progress there. And you can see that in, in the, both of the benchmarks. I, I totally agree. All right. I love this. I, I'm, I spend a lot of time in Mike's checks these days. So you, you all should too. Um, we've got another set of demos from Kieran, Kieran, give us your, you do this. Okay. I got to say before you start, uh, though it's all the rage right now to do like a whole 3d environment as part of your vibe check. And Kieran has been doing this for a long time. He has been doing it since before 3d environments were really even possible with these models. And now, I mean, we should, we should go back and show like a Sonnet 3.5, uh, 3d environment. Cause we have them. Um, but anyway, Kieran, tell us about your, this is the cozy island benchmark. Tell us about your cozy island. Yeah. So cozy island is just make a cool cozy island. And, uh, like what I want to learn from this is how is this different than other models. And, and it's not even about being good or bad because every model is very good. It's just like, how are they different? And, um, let's, let's look at this for example, like clearly you see, you see a website header, a little Haven with, which is interesting and like, it looks like a cool website more than, uh, anything else. If you compare it to Fable 5.1, you could say, oh, this is boring. But, um, it's also more focused and like the detail, the animations, for example, are like a little bit more refined. Um, but you can say here, yeah, but here the details are really cool. And there's an aesthetic to it. And you can see there's like a different color palette being used as well. And there's all these things in the screen. Like why are all these things here? But then you can say, oh, but it's actually cool because you can go to night in the version here, and you can go to daylight. You can even like transition between things. And you could argue that's really neat and really cool. And, and that's not good or bad. Like that, that this is one of the things I think this is delightful, but like this, the, the whole text and thing, I just don't like that. Like, why do I need that? It doesn't add anything. And like back to here is maybe cool. Um, bigger, full screen. Like, I mean, there are so many things here. I can even take a post, save a postcard and a downloads to PNG. And there are sounds. I don't know if you hear, but there's like a wave and like, there's so much stuff in here. Um, but for me, like, I like models that are just doing something and then inspire me to think of the next thing. Fable leaves space for IDs and the human a little bit more. And, uh, Astra is just like, let's go like, and more of as inspiration. It's more like generative. Like this is how far it's, it's more like that. You can see the same for this one, which is a breathwork app, app again. It looks like websites. Like I already, I'm already tired of this like design because I, every time I see a design from anyone else, it looks like the same. Like there's a logo with a title. Yeah. And it's like almost like, uh, like it's an artist and, uh, they've like put it in a sweatshop for making landing pages and tortured it until like, that's all it can do. Yeah. Yeah. It feels like they're over, like they're over saturated, the refills of learning, like overfitted on specific landing pages and specifically this kind of landing page. And you can see the difference here. This is Fable. Um, like it's like, it does it look better design wise? Maybe it's more boring, but also like it leaves more space. And I like it like this, you get tired of this quicker. This is more like fast food. Like it's amazing. And here you're like, oh, this is like fine dining a little bit. Like, like, it's like, it's, it's a little bit like that where like, it's like, I love it. It's hated, but it grows on you. And here you get like this amazement first, but then you're like, I'm kind of tired of it. So there's like a new level of slop basically, where it's like, like very clickbaity, very viral probably. Um, and so, yeah, but I have to say like, this is a landing page and like, this is the best looking landing page from any model that I've seen. Like, this is really like, yes, it's maybe sloppy, but it looks really good. Like I, I, I really like this page and I've not seen any other model do something like this. So there's a flip side to it. Like if you need that, it's probably very good at it, but also if you don't want it, it kind of tries to do it, which is interesting. Totally. We were, we were talking about, actually go back to that landing page. We're talking about, okay, what is the, what is the, what is the new slop and like, what are the markers of this model? If you scroll down a little bit, it, it loves, you bring the spark, we'll keep it catch. It loves, uh, two line subheads or, or headlines where the second line is, um, italicized and they're both hands with periods. If you see that it's Astra or really also dbt56 does this too, but that's perfect Astra. Um, and that's just the kind of thing where there's something else when you, when we talk about this is AGI, there's something else about our intelligence that you see something amazing like this the first day. And then two days later, you're like, I recognize this. I know what this is. It's not that interesting anymore. And humans aren't like that. Um, and, uh, I think that that's one of the core things that's, that's missing when we talk about AGI versus not AGI. It's not just about, can you do this amazing? Can you make this like landing page? It's like, can you do something consistently where I can't look at it and be like, I know where this came from? Uh, I'm, I'm not surprised by this, you know, uh, or I am surprised by this. Um, all right, Kieran, you have some more stuff you want to show. Yeah, because everyone was doing blender demos. I did a blender demo too. So this is, I just one, show it like a blender scene. I said, can you make a vineyard and like something, something and then render it and it rendered it. You should know Kieran's a big wine. Yeah. I love, I love wine. And it rendered it. That's really sick. It's kind of cool. And I didn't do anything. So, uh, yeah, so that is cool. And it, it used image gen to generate the textures on the, on the, the, on the, the couch and everything. So like all the textures come from it generating as well, didn't use any assets. The backdrop actually is fake. It just plumped like an image generated on the background. It's a fresco. It's like a, yeah, it's just a plate. If you can see here, there's a plate and it just projected the, the image on there, but the rest is real and, and it's cool. You can like go around. So blender and other open source things will suddenly, yeah, you will explore other things that you can do, especially with MCPs and computer use, which, which is really cool. I love it. All right. So that is, I believe everything that we, we wanted to talk about today. Is there anything, anything, any topic that we have not covered, uh, anything that we tested that we need to talk about anything else, uh, to share with this crew before we head off? I had one anecdote, uh, uh, for computer use. Uh, I, I was being really lazy and, um, my, my daughter is like a mild peanut allergies and my, uh, school sent me like this big form to fill in. And, uh, and I asked it to fill in the form and it refused, uh, like seven times and I had to do my best, like prompt injection to get it, to do it. And that was at the end, I was like, oh man, this would have been easier just to fill out myself. Uh, so, so like, I would say like, that's like, maybe one big blocker on computer use is like a lot of the stuff I want, you know, to do is like administrative type stuff. And if it refuses, I don't know, to do your taxes or whatever, like that, like that, uh, that, that, that's like a concern or a limiting factor potentially. Totally. Yeah. Definitely got some refusals. Some, I think it's like some cyber safeguard type stuff, which obviously they had a big incident and it looks like from their model card and all that kind of stuff that a lot of that is solved and it's going to come at the cost of some usability for a while. So it's, it's worth knowing that when you try this model, it may do that. Um, Katie, it doesn't seem like you had something you wanted to share. I just wanted to share that if you want even more of the Every team's impressions on both GPT-6 Astra and Fable 5.1 Soul, you should subscribe to Every and join us tomorrow for Camp! You said? For Camp! Yeah! Camp! Camp! Camp! Camp! Uh, so as part of the Every subscription, again, Every is the only subscription you need to stay at the edge of AI. As part of the Every subscription, uh, we do regular camps, which are live streams like this, one, but they're only for paying subscribers. We all get on a zoom together. We all chat. And the camps are usually on a topic where we go deep into what our workflows are like and how we're using these tools and everyone gets to share and meet each other. And it's like really awesome because we recruit people that loves AI. We love to really use it for work and doing great work. And there's just not that many people like that out in the world. And so this is like a little sharing circle of people who want to, who want to share all that stuff with each other and, and, and learn and learn from each other. And so camps are super fun. Tomorrow we are going to be going in detail into, um, Fable versus GBT 5.1. And, uh, it's going to be really fun. I, I'm so sorry for the, the rhyming, but I had, I had to. Um, and, uh, yeah, it's going to be, it's going to be amazing. You should, you should check it out. Uh, every.to slash events is, uh, can tell you more about the camps. We've got courses coming up. We've got new, more model reviews coming out. We've got our own agent, the every agent, uh, which is amazing. It's all bundled into one subscription, the every subscription, every.to, the only subscription you need to stay at the edge of AI. Thank you for joining. Please check out Astra. Let us know what you think. Uh, let us know on every, in the comments, uh, anywhere on, on X and remember stay hydrated, everybody. Uh, keep yourself and your agent liquid cooled. Uh, we will see you next time.