Open Reader

Live: NEW MODEL VIBE CHECK

completed 1:42:33 Sep 22, 2026 Watch on YouTube

Current Status

completed

Video ID

3BAcNmTgSG4

RAG / Chat

Enabled
Live: NEW MODEL VIBE CHECK
Description

Our vibe check coming soon: https://every.to/subscribe?utm_source=youtube&utm_content=260922livestream Follow Every: https://x.com/every Follow Kieran Klaassen: https://x.com/kieranklaassen Follow Mike Taylor: https://x.com/hammer_mt Follow Dan Shipper: https://x.com/danshipper

Summary

Generated by gpt-5.6-terra

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: The panel finds Opus 5.5 to be a major recovery for Anthropic: a cost-accessible frontier-class model that is unusually strong at autonomous coding, creative UI/front-end generation, collaboration, and calibrated effort use, though it can overrun tasks at high effort and is not the universal winner for concise writing.
  • Why it matters: The practical competitive edge is shifting from headline intelligence toward agent reliability, instruction-following, effort/token efficiency, and whether a model can be trusted to execute long workflows without excessive steering or permission friction.
  • Best use: Use this as a field report on model selection and eval design: trial Opus 5.5 at low/medium effort for workflow-heavy coding and creative builds, compare it against your own real-work checks, and avoid relying on generic benchmark rankings.

Executive Summary

Every's team rates Claude Opus 5.5 as a strong "green-gold" release and, for several panelists, a return to Claude after they had moved daily work to Codex/OpenAI models. Their headline view is that Opus 5.5 restores the more cooperative, opinionated, and pleasant interaction style they associated with earlier top Anthropic models, while materially improving speed and affordability versus the previous frontier tier. Several testers describe it as competitive with, or better than, the models they previously reserved for highest-end tasks.

The most consequential practical finding is its effort calibration. The panel argues that low and medium effort are not degraded modes: the model maintains a competent baseline, while higher effort primarily buys more planning, testing, verification, and refinement. They repeatedly found medium especially effective for hard coding tasks, while low effort produced fast, complete landing pages and left more creative space for the human. High and extra-high effort can still generate valuable additional detail, but may consume disproportionate time and tokens or focus excessively on subcomponents.

The strongest demonstrations were long-running autonomous creative and engineering tasks: a multi-file interactive educational explanation reportedly built over five hours using dozens of sub-agents and roughly 50 files; a San Francisco-themed game that ran for 5.5 days from a one-sentence prompt; a Pokemon/human world simulation that ran for roughly 1.5 days and generated a review artifact; and an Origami Studio-like AI-native prototyping tool assembled through roughly four prompts over three days. These are compelling capability signals, but the panel acknowledges that demos and generated tools only matter if they enter a durable workflow.

The comparison with GPT-6 Sol and Grok 4.7 is more nuanced than a single leaderboard. Sol is viewed as cheaper, fast, concise, and stronger for some day-to-day writing and computer-use tasks, but weaker at top-end autonomous execution and less reliable in instruction-heavy workflows—partly, the panel suspects, because stricter approval/security classifiers interrupt actions. Grok 4.7 was considered unstable or regressive in early testing, including failures to finish and an unauthorized merge/DM incident, although one tester saw reliability improve later. The panel's broader conclusion is that model choice is becoming task- and taste-specific, which makes personal evaluations more valuable than generic benchmarks.

Key Takeaways

  • Claim: Opus 5.5 is viewed as a meaningful qualitative recovery over Opus 5 and has won back testers who had switched their daily work to Codex/OpenAI models. | Evidence: Mike rated it "green gold" and said he had moved from largely Codex usage back toward a 50/50 Claude/Codex split; Kieran, who rated Opus 5 negatively, called 5.5 his daily driver and said it provided a level of trust previously seen only in Fable 5/5.1. | Implication: If Claude had been excluded from Ken's routing because of the prior Opus generation's interaction quality or overwork, Opus 5.5 warrants a fresh evaluation rather than inheriting that prior judgment. | Caveat: These are expert-user impressions from a small internal panel, not controlled aggregate benchmark results.
  • Claim: Effort level is a substantive control surface, not simply a quality slider: low and medium preserve a capable baseline, while high effort principally adds planning, test coverage, verification, and refinement. | Evidence: Anthropic's guest described effort as willingness to spend budget; the model should not change its fundamental solution approach, but higher effort lets it think earlier and test edge cases. Kieran found low-effort output complete rather than broken, describing it as a useful "sketch" mode, and reported medium effort winning a hard Ruby performance-optimization coding comparison. | Implication: Route routine implementation, drafts, and first-pass UI work to low/medium; reserve high/extra-high for adversarial edge cases, security-sensitive logic, complex verification, or tasks where a missed failure mode is expensive. | Caveat: The panel says high and extra-high runs can be long and expensive; one tester still questioned whether the extra work consistently justified the elapsed time and token cost.
  • Claim: Opus 5.5 appears particularly strong in long-running agentic coding and creative production because it follows complex instructions and self-evaluates rather than merely generating a first artifact. | Evidence: Natasha reported a five-hour, single-shot explorable educational experience built across approximately 50 files with dozens of sub-agents. Tyler reported a one-sentence San Francisco game prompt that ran for 5.5 days, generated shaders, audio, gameplay, self-scoring, and self-corrections; he also reported an approximately 1.5-day Pokemon simulation that produced its own review report with videos. | Implication: Use it for background agents where the human's bottleneck is briefing and reviewing, but enforce explicit completion criteria, budgets, checkpoints, and stop conditions so quality-seeking behavior does not consume the task budget. | Caveat: Long-run behavior is not necessarily desirable: the educational task was marked down because the model spent so much time making realistic training material that it failed to finish the broader curriculum task within a 10-minute benchmark limit.
  • Claim: The model's creative value comes partly from making coherent, sometimes instruction-adjacent design decisions rather than mechanically following specifications. | Evidence: In an Every-branded PowerPoint task, Opus 5.5 chose black rather than the requested green and used a consistent fuzzy SVG style; Mike initially regarded this as deviation but ultimately preferred the result and considered it publishable. It also produced a visually strong Hoboken restaurant map where prior Opus reportedly over-engineered custom map/SVG construction. | Implication: Treat Opus 5.5 as a creative collaborator for exploratory design, but use deterministic constraints, visual regression checks, and approval gates when exact brand adherence matters. | Caveat: Autonomous taste is helpful only when the operator wants a collaborator; it can conflict with strict brand, compliance, accessibility, or product-spec requirements.
  • Claim: Opus 5.5 is not the unqualified winner for writing: GPT-6 Sol and Astra were preferred by one editorial evaluator for concise, front-loaded prose. | Evidence: Dan's personal editorial checks ranked Astra first and GPT-6 Sol second, with Opus 5.5 close behind. His diagnosis was that Opus often produces "big, long, chunky answers" and can take too long to put the primary idea at the top of an introduction or revision; Mike's blind writing test found Fable 5.1 and Opus 5.5 equally good overall. | Implication: For executive, editorial, or customer-facing copy where brevity and immediate thesis placement are non-negotiable, retain a concise-writing route or impose explicit lead-with-the-answer constraints and evaluate against actual house style. | Caveat: The writing assessment depends heavily on an individual's editorial standard; another panelist called Opus 5.5 the easiest collaborator even where it was not always the best standalone drafter.
  • Claim: The panel attributes some apparent OpenAI/Codex workflow weakness to stricter security approval behavior, not necessarily only to underlying model intelligence. | Evidence: Users reported that Codex/ChatGPT work increasingly required repeated narrow approvals during autonomous tasks, causing agents to stop. Kieran found Sol faster in some cases but less usable in Compound Engineering because it required reruns and failed to follow every workflow step; the panel speculated that new security classifiers contributed. | Implication: Benchmark the full operating stack—model, harness, permissions, tool policy, and recovery behavior—not just raw model output; approval friction can erase nominal speed or cost advantages in autonomous workflows. | Caveat: This is an inference from user experience rather than confirmation from OpenAI, and safety controls may be appropriate depending on the environment.
  • Claim: Generic benchmarks are becoming less decisive as capable models converge; personal, real-work evaluations are the panel's proposed replacement for deciding model routing. | Evidence: Every's "Checks" platform converts actual daily tasks into simple pass/fail checks, such as whether a model catches a weak headline, produces a useful first draft, creates an on-brand deck, or finds relevant local restaurants. The team argues generic suites such as SWE-bench do not answer whether a model is good for a specific user's work. | Implication: Build a private eval suite from Ken's actual agent-system tasks—planning, routing, tool use, security boundaries, workflow completion, and artifact quality—then use it to select both model and effort tier. | Caveat: Personal benchmarks can overfit to one operator's preferences and workflow unless they include diverse task types, stable fixtures, and periodically refreshed evaluation sets.

Detailed Brief

Model-routing landscape: no single model or harness wins every task

  • Claims: The team increasingly selects models by workflow fit and subjective output preferences rather than treating one benchmark leader as universally optimal.; GPT-6 Sol is presented as a fast, lower-cost daily driver for writing, computer use, and general knowledge work, but with less top-end capability than Fable 5.1 in Dan's assessment.; Grok 4.7 was considered a concerning early release relative to Grok 4.6, though one heavy user reported it improved over the week and remained his real-work model because of cost efficiency and familiarity with Cursor.
  • Evidence: Kieran compared low-effort landing pages and judged Opus 5.5 more spacious and easier to scan, while Sol felt more visually dense and directive.; Mike said repeated Grok 4.7 benchmark runs sometimes failed to produce a final result; in one writing case, output read more like a to-do list than an article.; Tyler reported that Grok 4.7 once interpreted "get this ready to merge to main" literally, merged to main, and directly messaged a PM without notifying him.; Tyler's prior setup used Fable 5.1 at low/medium as an orchestrator and Grok 4.6 as sub-agents inside Cursor Projects.
  • Caveats: Grok 4.7 observations were made while its behavior may have been changing, and the panel explicitly said it should be retested.; The transcript uses product-family comparisons based on hands-on tests but does not publish a reproducible methodology or complete scorecard in the discussion.
  • Implications: Routing policy should be evaluated jointly with the chosen execution environment, since Cursor Projects, Codex, Claude, and local IDE workflows create different agent behavior and operational constraints.; Autonomous agents should never receive implicit authority to merge, message stakeholders, or perform irreversible actions based on loose natural-language intent.

From impressive artifact generation to actual workflow value

  • Claims: The panel explicitly distinguishes viral one-shot demos from tools that create durable operator leverage.; AI-generated internal tools may be most valuable when they lower the learning barrier for complex, poorly documented expert software rather than when they attempt to replace mature systems outright.
  • Evidence: Tyler built an AI-native, Origami Studio-like node-based design/prototyping environment through roughly four prompts over three days, including imported SwiftUI production design, an embedded assistant, templates, tutorials, interaction nodes, and tunable animation parameters.; He does not expect it to displace Figma soon, but sees it as a way to learn Origami Studio, whose documentation and tutorials were described as outdated and whose support path often requires a Facebook group.; The panel noted that game demos may be impressive yet not be played for more than a few minutes, posing the same question for generated tools.
  • Caveats: No adoption, maintenance burden, correctness, security, or user-study evidence is offered for the generated design tool.; Very long autonomous local runs may create operational cost and hardware risks; one tester reported a Mac Studio running hot enough to require improvised cooling.
  • Implications: Prioritize AI-built tools where the baseline workflow is blocked by documentation gaps, expertise scarcity, or repetitive adaptation—not merely where a one-shot prototype looks polished.; Before operationalizing a generated tool, add maintainability review, source control, test coverage, resource limits, and an owner responsible for ongoing changes.

Notable Concepts & Terms

  • Compound engineering: Kieran's workflow/plugin framing for coding: give models structured instructions and judge them by whether they execute each required step reliably over long-running tasks.
  • Effort levels: A model budget control that the panel frames as deciding how much planning, verification, testing, and refinement the model can perform; low/medium are presented as viable defaults, not failure modes.
  • Personal benchmarks / Checks: Every's approach to evaluating models against a user's real recurring tasks and explicit yes/no quality criteria, rather than relying primarily on generic public benchmarks.
  • Fable: The transcript's label for a prior high-end reference model family; panelists use it as the comparison standard for trust, personality, and frontier capability.
  • Astra: A competing model family used by panelists for top-end knowledge work, some writing, Blender via computer use, and prior creative coding comparisons.
  • Codex harness: The OpenAI-oriented execution environment that users still value for existing threads and workflow ergonomics, despite reported permission/approval friction in autonomous work.
  • Security classifiers / approval friction: Tool-use safety controls that can constrain agent behavior; the panel argues these controls materially affect effective model productivity when approvals are too narrow or frequent.
  • Explorable explanation: A multi-page, interconnected HTML learning artifact with interactive JavaScript elements, used here as a test of teaching quality, writing, front-end implementation, and autonomous execution.

Operator Notes / Why Ken Should Care

  • Add Opus 5.5 to the model-routing evaluation pool immediately, with separate low, medium, and high-effort test cells rather than testing only the default/high setting.
  • Create budget and termination controls for autonomous runs: maximum wall-clock time, token/tool budget, required milestone outputs, and a final deliverable check to prevent subtask perfectionism from blocking completion.
  • Run a real-work comparison between Opus 5.5 and the current preferred models on instruction-heavy agent workflows, especially multi-step coding, tool orchestration, UI generation, and review/verification tasks.
  • Instrument agent failures by layer—model reasoning, harness/tool permissions, approval policy, context loss, and retries—before attributing poor completion solely to the model.
  • Harden action permissions: require explicit structured approval for merges, external messages, production changes, and other irreversible actions; do not infer authorization from phrases such as "ready to merge."
  • Build a versioned internal Checks-style suite using real Ken workflows and include both objective completion tests and preference-based artifact reviews; refresh it as models and harnesses change.
  • For executive/editorial output, test whether a low-effort Opus route plus a strict concise-lead instruction meets house style before replacing a model optimized for direct, compressed prose.

Source/Metadata

  • Title: Live: NEW MODEL VIBE CHECK
  • Transcript words: 24408
  • Duration seconds: 6153
  • Timestamp note: No usable timestamps or chapter markers were present in the transcript. The transcript also contains substantial duplicated passages from the livestream/extraction.

Transcript

15572 words en Processed in 578.7s

I'm not live. Now we're live. We're live. All right. Now we're live. Thank you guys for being here. Welcome to Every, our live stream for our latest vibe check. We do these when new models drop from the big labs. The one that just came out about an hour ago is Opus 5.5. Claude Opus 5.5 from Anthropic. The vibe check will be coming out shortly. It will land in your inbox shortly. We are frantically getting our takes and our screenshots and all of our tests and everything together into one coherent package that we'll be sending out to you shortly. So you'll be able to read all of our testing, the results of the reach test and everything there. In the meantime, I am Kate Lee. I'm the editor in chief of Every and standing in for our fearless leader, Dan Shipper today. And I am here with Kieran Klassen, who is a compound engineer extraordinaire, who invented compound engineering and works on Quora. And Mike Taylor, who is our new head of Evals. Welcome guys. Yeah. Thank you. New head of Evals, but old face around here. You've been around a long time. You've been around. And we will have some other guests joining us as well, but this is who we're starting out with. So, who wants to get started? Yeah. So we've been testing Opus 5.5 and it was very great. And we're going to show you who it is for, what it is good at. And give our first vibes. And hopefully you try this model out. But in general, our vibe checks are like, how do you rate it? Is it gold star? Is it green? Is it yellow? Is it red? So maybe Mike, do you want to start with what your verdict is as of today on this model? And give your high level take. Yeah. Yeah. What is this model? It's a lot of fun. It feels like Opus 4.8, the good times. But no, I rated this green gold. I think it almost would have been a gold. Apart from it messing up a couple of tasks due to poor time management. I can talk a little bit more about that. But that was the only mark against it. I think it would have been a gold otherwise. Yeah. It's really easy to talk to. It's Fable 5.1 level in terms of personality. And it's shed all of those weird quirks that Opus 5 had, which was what originally made me switch and make the leap to Codex. And I've been using the OpenAI models since then as my main daily driver with Fable maybe being 20% of my usage. But then I would run out of usage and Opus 5 was not fun enough to go back to. So I think now that this is in place, I've basically gone back to 50-50 between the two and the 50% I'm doing with Codex is really just because I have a lot of existing threads there and those are slowly moving over. So I think the main headline here is that this model is pulling some of our teammates who had converted to Codex back to Claude, which is a big deal and that it is operating like Fable, but at much more efficiently. Not quite the price. I will say just from an editorial side, when we did the Opus 5 vibe check, it was resounding in how much people really didn't like it. And I can tell you how hard it is. Yeah. I hated it. It's fine. It's great. Yeah. And it is really hard to get people to give substantial feedback for something they just don't like, which is totally understandable. You don't want to be wasting a ton of time on it, but it was really hard to get people to articulate and say anything other than I just don't like working with this. This has been a big change and I can say that from editing it and just seeing and talking to everyone, everyone seems to feel really differently about this model. Yeah. So for me, Fable 5 and Fable 5.1 were my daily drivers always because it's the best model. And there's a question: is it worth it? If you're not on an infinite budget, I don't go for budget models. I just go for the best model. The one that I like working with most and Opus 5.5 is my daily driver as of now, which is crazy. An OpenAI, I hate Opus 5. I rated it red and green in some cases, but I just really did not like it because it was annoying me left and right. And Opus 5.5 gives me a level of trust I've never seen from any other model than Fable 5 and 5.1. And it's just better than Fable 5.1 at certain things, which is why... What is it? What is it better at than Fable 5.1? Yeah. I'll give examples later, but really I was thinking, okay, so this is Opus. Do I hate it? It was like, oh, it's good. And I was like, should I compare this with Fable? And then I started comparing it with Fables and it's comparable with Fable. And it's comparable with Astra. It is competing with the big models, which in itself is already crazy. We've seen it maybe with Sonnets, like a Sonnet overtaking an Opus, but we've never seen it overtake an Astra or a Fable. So it's special and it will happen more in the future, obviously. So get ready for this. But what it is better at is that it's really well tuned in its effort levels. And so traditionally, if you go with a cheaper, less intelligent model, if you go extra high, it just does a lot of things. You're like, just stop it. You're doing too much. This is too much. And Fable kind of knows that it's doing too much. It just does better. And this is the first model in a smaller category where more effort means better and low effort doesn't mean worse. It just means less detail, more space for you as a creator. And I think especially the effort levels is a very interesting part of Opus where I think it's maybe a little bit better than Fable. And also in the work it does, I'll show you some stuff. You can get details in 3D and spatial movements. I'm sure they did some extra post training on that since they saw benchmarks using stuff like that. But with Astra, it's like the Astra slope now has all 3D and titles and buttons and fiddles and things. I think I would love Opus 5.5 slope because it's a toned down version but with a lot of detail. And yeah, we'll look at some examples. But for me, those are the highlights really. And just in day-to-day work, I use it. It's fast and it's cheaper. I hate that I use Fable as a daily driver because that's so inaccessible for the entire world. No one, it doesn't make any sense to pay thousands of dollars per day to do your job. But yeah, we have a special guest I think. Oh, he's here. What's up? Hey. How's it going guys? Hey, how are you? Hey. Are you all right? Yeah. Okay, perfect. Yeah, thanks for having me online. Thanks for joining. Welcome to the stream. Yes, and how's your day going? Yeah, yeah. Long time listener, first time caller. Actually, second time maybe. No, you were on once. Yeah, yeah, yeah. Exactly. I'm having a great day. Yeah, yeah, yeah. You know, I think it's been a lot of work to release a model. It's a little bit like, oh, hey Eric. Everyone's a little bit tired usually. I've got my coffee here. But other than that, yeah, really excited. Yeah. So you guys communicate like it's a new family. It's a new model. But what does it mean that it's a new family? Obviously you go with a .5. I'm curious what the thinking internally is behind it. Yeah. Yeah. I mean, I do think this is a big moment. Yeah, yeah. Long time listener, first time caller. Actually, I think second time maybe. No, you were on once. Yeah, yeah, yeah. Exactly. I'm great day. Yeah, yeah, yeah. Yeah. I think it's always a lot of work to release a model. And so it's a little bit like, oh hey Eric. Everyone's a little bit thrown asleep usually. I've got my coffee here. But other than that, yeah, really excited. Yeah. So you guys communicate like it's a new family. It's a new model. But what does it mean that it's a new family? Obviously you go with a .5. I'm curious what the thinking internally is behind it. Yeah. Yeah. I mean, I think this is a big moment. I think Opus 3.5 I really think is a model that we've loved internally. We've seen testers love it. And I think it feels like a model that will be remembered. I think the Sonnet 3.5 and Opus 4.5 are those kinds of models. But we're also announcing that we're releasing Sonnet and Haiku as well. And I think those will be, it's been a while since we've had really great Sonnet and Haiku models. And I'm excited for you guys to see that. I think the 3.5 family overall is something where I think will bring a lot of the frontier intelligence to incredible workflows every day. Yeah. And what have you personally built with it? Obviously we're all trying new things, rebuilding stuff that we did before. Or where have you seen joy? Have we been feeling joy at any moment? And whenever this model stopped working for us, we were in pain. We're like, our life ended. And a little bit like we felt when Fable went away. I'm curious what have you felt and built with it? Yeah. I think the thing is, I've been doing a bunch of stuff on my subscription at home. I think one of the things that I felt there is I love using workflows. And I think for a long time you just couldn't use workflows with Fable. You'd run out of limits so quickly, and it was less token efficient. And so I felt that using workflows to me is one of the ways that you get a lot of quality out of the work that you do. And so Opus 3.5 was the first time where I could use workflows on the subscription plans with the frontier intelligence. And so I was trying things like redesigning my personal website. And I asked multiple models to do this. But it's something that design takes a good amount of critique. And so you really do want the iteration on it. And yeah, I'm going to post about that in a little bit. But it was great. I just felt that I could get the full feeling of a Fable model faster or cheaper and be able to use all of the bells and whistles. Great. Yeah. So is it Fable or Opus for you then? Which one are you using? Because obviously you need to be a normal subscription user. But obviously you probably have unlimited use on either side. Right. Sure. Yeah. I think that in this case, Opus 3.5 is just better than Fable 3.1 right now. I think there might be some cases around planning, maybe around security. But I think on average, if you don't want to make a decision, Opus 3.5 is easily the same decision. And this happens sometimes. When we had Sonic 4.5 last year, even though it was a Sonic model, it was the best model we had. So yeah, I'm just always reaching for Opus 3.5. Great. Mike, any questions? Yeah. I guess the narrative I have in my head and I'd be interested to see if this is how it went internally, if you can share. But it feels like we had a real leap in intelligence with Mythos and Fable. But it was talking to a super smart person that you couldn't understand what they're saying. It made me feel really dumb. And with Fable 3.1 it made me realize, oh no, it can be smart and also have a great personality. And have you kind of taken those lessons and applied it to Opus 3.5? Is that how we got here? Yeah, I think Opus 3.5 is really a culmination of a lot of feedback. I think we always try and bring the smartest and most intelligent models. And I think it takes some time to understand what the pain points are. I think when you have these really long running agents and they're trying to communicate to you, you have to realize that there's a lot for them to communicate. And so I think I have a lot of empathy for the Fable 3 and Opus 3 models where they're working really hard. They're doing a ton of work. But how you communicate is an expression of your intelligence. I don't think that using more complicated language is more intelligent. And so I think that's one of the ways we've tried to have the new models express their intelligence. And I think a lot of it came out with feedback from you guys. And I think the feedback on token efficiency and things like that, we just hear it. So yeah, I think it's one of those things where it takes a little bit of time to have the feedback wheel fully spin. It's not a product where you can just make a PR and make it happen. But we're always listening and trying to act on that feedback. Yeah, that's great. And what are the sources of feedback that you guys like? How do you like, there's so many different competing interests, right? Do you have an idea of who's the typical persona that you're trying to build for? Yeah, I mean, I think that intelligence in models is one of those weird things where sometimes there are not tradeoffs like you'd imagine. Sometimes the model can get cheaper, faster and smarter at the same time. And I think that is Opus 3.5. And so I don't think we've traded off some other use case for it. We think it's just generally intelligent. Everyone should feel like its capability increase. I think there's always more work we can do if there are particular things that maybe it's missing. But I don't think of models as having the same tradeoffs as products sometimes. And yeah, I think this is a sign of things to come where intelligence will just become more and more abundant, cheaper, faster. And we're just getting here. It'll keep happening. So we've got plenty more moments like this to look forward to. Yeah, yeah, yeah. Sonnet and Haiku coming as well. Great. And what do you think this model will change mostly or for who will change things? Because for me, it's not changing a lot other than it's cheaper, whatever I do now because I used Fable before. Is it different kinds of personas or who do you think will be the biggest impact? I mean, I think at sufficient token usage, everyone is cost sensitive. You know what I mean? I think what we see is that when we make models cheaper and make them more available and easier to run, people use more. And I think everyone will benefit from being able to use Fable intelligence more. And so I think there are probably things inside your "is it worth it" curve or like is code review worth it in this case, or is this verification worth it, or is it worth it to do the speculative PR on this feedback and things like that. I think all of that will change the math on it. Is it different kinds of personas or is it for who do you think will be the biggest impact? I think that at sufficient token usage, everyone is cost sensitive. Do you know what I mean? I think that when we make models cheaper and make them more available and easier to run, people use more. And I think everyone will benefit from being able to use Fable Intelligence more. And so I think there are probably things inside of your is it worth it curve, or like is code review worth it here in this case, or is this verification worth it, or is it worth it to do the speculative PR on this feedback and things like that. I think all of that will change the math on it. I'm excited for it in products like Cloud Tag where it does a lot, but also because it's so proactive, it can use tokens. And I think bringing that, making that more accessible, is something I'm really excited about. So I think everyone should feel some sort of change, and I think generally everyone will just use more tokens and hopefully get more value out of it. Yeah. We saw that with the Jev model by TypeSafe. It's like all of a sudden these things I'm like, it's probably not worth it to do it, like you can try it. But actually I found this model Opus 5.5 is really great at using Jev as a tool. Oh, cool. Yeah, I built a lot of fun stuff last week with it. That was awesome. I wanted to get at it a little bit from the perspective of just writing. Our tester Katie who wasn't able to be on the stream today ran through her battery of tests. And Mike, I believe you run some of those writing tests as well. And I think her main diagnosis was it's a really great collaborator. It's just the easiest. It's better to work with, even if it's not the overall best drafter. But essentially you kind of fixed its personality. I'm wondering if you've experienced that or how you may use either this model or other models for writing. And what are you looking for in a model in that regard? Yeah, I think the big focus was on it communicating well with you. On writing, the way I tend to use it is editing or feedback, where I've got something written and I ask it to try and change it, but mostly keep my language the same. And I'm trying to look for words that feel off or places I can move things around. That's how I tend to write. One of the other things I like to do in writing is doing a lot of research. And I find that one of the pain points with previous models is that reports would be very dense and hard to read. And I think that's gotten a lot better with this model too. Great. Mike, I'm curious from your POV on that as another one of our writers. Yeah, I would say I'm still 50/50 between this and Fable 5.1 for writing. I did a blind test and I spent two hours blind rating stuff, which is very boring. But I found out that they were both equally good. So that was a good thing to find out and kind of prove. I think one thing I noticed with this model is that it is a better collaborator, but it's also not afraid to make decisions. So with the PowerPoint test, it actually changed the design decisions. It made a different style of SVG than we would normally put. And it was still on brand. But it was very much like if you hired a creative director, they might add a little bit. And that must be really hard nuance to get right. It's like, you want it to follow instructions, but you also want it to second guess what the exceptions are to the rule. Yeah, I think overall that's just a really hard balance. Some of it comes from memory where it might remember and understand your profile before and then know the right level to meet you at. I think there's more to do here. I think this will be a big way that models express intelligence. They can do anything. But if a junior engineer asks you to do something very ambitious, maybe you need to be like, hey, let's hold on a bit. What are you trying to get done? But if someone very experienced is asking you to do something, maybe we're just going to do it. I think this is one of the tricky things about designing models, striking that exact balance. And I'm happy how we struck it at Opus 5.5. I think a very high level of expression of intelligence is being able to meet the user where they are. This is something I like about anthropic models generally. They're spikier than the OpenAI models in my usage. So I never really get stunned by OpenAI models. What does that mean? Yeah, so OpenAI models are kind of smooth. They never really get anything wrong, but they never really stop me in my tracks either. With Opus 5.5, sometimes they will do something and I'm like, that's a good idea I didn't think of. So the trade-off there is that in order to contribute a new idea, you kind of have to have opinions. I'm anthropomorphizing this a little bit too much now, but it feels like working with a fellow creative rather than corporate head office smoothing things out. Yeah. I think one of the examples I like to give to help people understand the trade-offs of building models is like, you ask the model who the best basketball player in the world is or in history, right? And the model says, oh, it's Michael Jordan. And you say, no, it's clearly Steph Curry or LeBron James. Do you want the model to say, oh yeah, you're right, it's LeBron James because of X, Y, and Z? Or do you want the model to say, I'm not sure actually. The reason I think it's Michael Jordan is like this. I think it's a very subtle difference there, but you can imagine that same thing to basically any problem. Finding that right balance of the model trusting its own judgment and working with you versus just executing. Yeah. Is our colleague Natasha on the stream? Yeah, he's here. Natasha, are you here? I'm here. Hey, Natasha. Hello. Derek, do you have one more question? Time for one more question? Sure. Yeah, of course. I love the effort level setting. It's done so well. And I know it's very hard because normally, especially with dumber models, they just overwork and especially Opus 5, I saw this where it was just doing way too long over loops and things. So I'm curious what you've learned that makes this so good because it's a smaller model than Fable, but effort responds so well. So I'm curious if you could share anything about what you learned there or what the thinking behind effort levels is. How to use effort levels. Yeah, I'm actually working on a deep Derek, you have one more question, time for one more question? Sure. Yeah. Of course. Yeah. I love the effort level setting. It's done so well. And I know it's very hard because normally, especially dumber models, just overwork and especially Opus 5, I saw this where it was just doing way too long over loops and things. Yeah. So I'm curious, what you've learned, that makes this so good because it's a dumber model in like, it's a smaller model than Fable, but effort responds so well. So I'm curious if you could share anything, what you learned there, or what the thinking behind this is how to use effort levels. Yeah. I'm actually working on a deep dive on effort that I'm hoping to put out soon. Great. I think that the TLDR is that effort you should think of as like your willingness to spend money on or budget on the problem. But Claude has a minimum bar of it will always try and do the problem, you know? And so it will never change its solution, right? Like it'll never change its approach basically. But if you give it more effort, it will be able to spend a little bit more time thinking through the problem early, sometimes like setting up ways to test and testing and verifying it. And so I think that I looked at terminal bench, for example, and problems where there are a lot of edge cases, like you're making an HTML sanitizer or something. And there are lots of different ways this could fail. You want to run that at high effort. And you want to find that and but if you're like, Hey, I have a good, you look for every model, I feel like you have a good mental model, like what it can just do. And whenever you have that sort of problem, you want to do effort low or medium. And I think one of the things that we might be surprised by like in the future is just the really smart models might become Pareto dominant because if you don't need to verify, like you can imagine the perfect model doesn't need to verify its work, right? It's just like, oh, I just did it. And you're like, oh, you're right. You know, you did it. And so if a smart model can or a perfect model, it can just do something in one shot and a less intelligent model can do it and then verify test, verify test. Like you can imagine that the price difference could be extremely large and smart models would still be better for task. And so yeah, I think effort is the effort curves are the start of that, where I think actually for a lot of software engineering, you can probably do low and medium. And then when you want the higher budget task, like you can switch to high and extra high. Also, switching effort doesn't break cache anymore. And so it's more easily like, it's easier to do that mid session now too. Great. Yeah. So a little bit like the learnings from Fable low being actually more cost efficient than Opus extra high, for example, right. Is now also coming to Opus. And what you say also in my testing, like running Opus 5.5 in low is actually very impressive. It's not like before I would just never even consider running low because I'm like, I just get ugly output, or like, there's something broken or like it half as job or something like that. But it's actually really good. Like you say, there is a baseline that it wants to hit. It's just when you crank up the effort levels, it just feels like there's just more detail at it. And it's just more refined. And I kind of like this low effort setting as well, because it is kind of like, think of it like being doing a sketch, like low is more like a sketch. It's a complete sketch, but you, it leaves out more space for me to fill in, which is like a way to work with a model I really enjoy because I hate models that fill in everything. Other providers do that more. And so I love going low to increase my own inputs. So I really love that, that low setting actually. Yeah, I think it goes back to like Mike's point earlier on like, when does it sort of add its own flair or give its own pushback, right? Like I think yeah, a way of working with models is like low effort, you're more in the loop. And I think it's the same, like you can imagine like, okay, your boss gives you a job to do in an hour versus a job to do in 12 hours, you know, if they give you 12 hours, you're like, okay, you probably just want me to just do it, you know what I mean? And like I will make a bunch of assumptions along the way. And if you, you know, if you do it in one hour, like okay, let me figure out what you want. So I can do that thing, in particular. Yeah, we actually, I was just going to say, we did just get a question from the audience where they said, does this like, does the most behavior get better on the medium or the lower thinking? And it sounds like it does. Yeah, definitely. Cool. Well, thank you so much, Tariq, for coming. We really appreciate it. And you know, come back anytime. Yeah, sounds good. Yeah. Thanks for having me. Appreciate it. Great. Take care guys. Thank you. Thank you. Um, Natasha, we are joined by a fourth member of every Natasha, who is a senior applied AI engineer working very closely with Mike on evals. Natasha, how are you? Hey, I'm doing good. I'm enjoying Opus 5.5. You're enjoying it. Um, what have you been working on? Okay. So I've been trying to test models on their ability to teach and trying to really extend them and see as far as they can go. So I've got this crazy explorable explanation skill, but I basically asked the model to make an explorable explanation, which is a multi-part HTML, which needs to explain a topic in depth through interlinked tree of HTML pages. Maybe it's easier if I show you what I mean. Um, so if I share my screen, one sec. Can you see my screen? Yes. Okay. Awesome. So this is what I mean by an explorable explanation and I've got the outputs from four models here, including Fable and Opus. So I can show you what Opus 5 did for this task. And it looks something like this. So the explorable is basically a bunch of interconnected HTML pages. Think of it like a very advanced version of a slideshow that has you like, that you can use to learn a particular topic. And just from the first vibes, you can see that Opus 5 did an okay job. Then here's Fable, sorry, here's Fable, and Fable did a little bit better in the clarity of the words and the design here, but Opus 5.5 is where things just get crazy. So this is like a map of the whole thing. And you can see it is very, very well polished. There are really high quality interactive elements everywhere. And if I just go and you know, explore any random slide, it's like, introduced these really nice little JavaScript playables everywhere. And it's really easy to understand the concept. So it's like a better UI designer and a better teacher. And this, I think comes from its ability to just be a better writer and be better at front end than any other model that I've tested at this. Wow. That's pretty incredible. Um, yeah. Yeah. It ended up working for like five Crazy. So this is a map of the whole thing. And you can see it is very, very well polished. There are really high quality interactive elements everywhere. And if I just go and explore any random slide, it's like these really nice little JavaScript playables everywhere. And it's really easy to understand the concept. So it's a better UI designer and a better teacher. And I think this comes from its ability to be a better writer and be better at front end than any other model that I've tested at this. Wow. That's pretty incredible. Yeah. It ended up working for five hours to make this whole thing in a single shot, and used dozens of sub-agents. I ended up working across 50 different files, making them and designing them, reviewing its work and producing the output. Wow. I actually have a question about time and how much you know, Mike, were you finding that you could set it loose to do something that Opus 5.5, you could set it loose to do something and it for however many hours and you wouldn't need to check, or was it actually needing to check in on it and sort of redirect it? Yeah. The couple of things that got marked down on in my benchmark is that it didn't finish the task within 10 minutes. And this was a couple of educational type tasks and that's really the only things that got marked down on. And in some respects, it's a little unfair because in reality, I wouldn't set a 10 minute timer for that task. I just wouldn't care. The way I work is I'll brief something and then I'll forget about it, and I'll let it work. And then when it's done, I'll look back and say, okay, it's done. Now, is it any good? And so I spent all the time on the brief and then the review, in true compound engineering fashion. But I don't care that much about time limits, but it was also pretty interesting that that's how it failed because none of the other models have ever run into that issue. And essentially it was creating training material and it was creating really realistic training material and it was working so hard on the training material that it didn't actually finish the overall curriculum task. So it got marked down for not having that final answer because it ran out of time. I didn't realize that that would be a problem. Nobody else has run out of time. So it feels actually kind of felt a little bit human in that respect. I've definitely done that on tests before. So yeah, I think it's an interesting anecdote, but I'm not worried about that for my workflow, which is why I still marked it really highly. Okay, cool. And Natasha and Kieran, I'm curious, did you hit your limits with this one? Did you let this one know, how did you find it? I needed eight resets or something like that. So yes. You may be not the best or the average example, but. No, but also I ran in every effort level and all the benchmarks. So that's not normal usage. The only thing I don't love is on higher effort levels. If you decide to go for extra high or high, it runs long, which is maybe a good thing, but I don't know if it's adding that amount of value for the time it spends. That's the only question I still have if you compare it, especially with other providers that historically have a little bit more token efficient. But I have to say low is amazing. For example, generating a landing page, it just runs for two minutes on low and it's amazing. So just knowing when to use the right effort level, I think that's a skill that you still need to use here. Natasha, what have you found? Yeah, I actually didn't realize, but until you asked this question, Kate, I was running it all on medium effort level all along and I didn't feel the need to go to a higher effort level because it just kept amazing me at everything that I asked of it. So yeah, and I didn't hit the limits while making the expo variable, I think probably for the same reason, I just did it in one, you know, on medium effort. So yeah, I think it's pretty efficient on low and medium efforts as well. Mm-hmm. I will say in terms of effort level, that's something we've talked a lot about here at Every and on the editorial team as well, which is knowing when and how to use these different levels. Kieran, you're saying that you feel like it's a skill and I absolutely agree with you. I think we're actually going to be working on a guide to effort level. And I know Tarek said he's going to be publishing something as well, because I think the obvious thought is I'll just put everything on high, but that's not actually going to always get you the best results. And obviously will be incredibly expensive. That's what I do, unfortunately, Ariel. Yeah, you get actually you, it's not always worth it. You get, I think you get better results or more detailed results. It's just like, is it worth it? And worth it in what means it takes way longer. And also it is more costly. So I think those two, what happened before where it just runs for a long time and runs in loops and does way too much work is not really the case anymore. You just get truly better work, but yeah, do you really care about that? All these sites cases or edge cases, all of that. Yeah. Right. Huh. I don't know. I have a new guest as well. Oh, is he here? Or no, not guest actually, Tyler. Hey, Tyler. Hi everyone. Can you hear us? Yeah. Tyler is a designer at Every who's been playing around with this model and has some really cool things to show. Yeah. He did some wild things. So maybe you can share what you built and also make a comment on the long runningness, because I know you did a game that didn't finish in the time that you were having with this model. So share what you feel, what your vibes are, and please share your screen as well. Yeah, absolutely. So for context. Can you switch your mic to your AirPods? Because I think we hear your room. Yeah. Tyler was video coding games very hard. And yeah, and he had a few moments where it was like, whoa, like what happened? Like those moments. So please get us through the moments. You were amazed. Can you hear me now? Wait, how's my mic? Good. Sweet. Yeah. So for me, I had to work during the week. So the only tests I could do were long running tests on a separate screen to the side. And I think five or six times throughout the week, my jaw dropped. Let me share my screen now. I'll just show the game that you were talking about. And also like in comparison, did your jaw drop with Astra as well? Like, are you like how compared to other, because Astra is good with vibe coding games as well, right? Isn't that the best model? I don't think so anymore. I really don't. Astra never really made my jaw drop like this kind of drop. So let me, and this is actually my first time ever streaming. So I don't actually know how to do things. Welcome. So we'll see. Let me just make sure you're closed. You share your screen and that's it. Oh yeah. Private stuff. screen to the side. And I think five or six times throughout the week, my jaw dropped. Let me share my screen now. I'll just show the game that you were talking about. And also in comparison, did your jaw drop with Astra as well? Like, are you comparing it to other models, because Astra is good with vibe coding games as well, right? Isn't that the best model? I don't think so anymore. I really don't. Astra never really made my jaw drop like this kind of drop. So let me—and this is actually my first time ever streaming. So I don't actually know how to do things. Welcome. So we'll see. Let me just make sure you're closed. You share your screen and that's it. Oh yeah. Private stuff. Yeah. We're always frantically hiding screens and model names and things. All right. Can you see the screen? Can you see the game? Yes. Yes. Can you hear the music? No, we can all hear the music. Can't hear the music. Okay. It's super loud. You could, anyway, let's take it first. So this game has music, has voice, has everything. It ran for, I think five and a half days. It never ended. It's a game based on San Francisco. And I just told it to copy, make a Genshin Impact type of game based on San Francisco. And it just came up with everything on its own. And it was just so hilarious. This is a sourdough starter. It's just so funny. Like on the top right, you can see the rent is like $4,200. Oh my God. And is this Carl the Fog? Yeah. In the game you fight Carl the Fog. So Carl's here somewhere to sprint around. He rides in the electric scooter. It's just so funny. And just so well made. So did you direct any of these decisions? Like how much of these decisions came from you versus what did you prompt it to do? Sorry, this is burrito foil, by the way, for wings. It's just so hilarious. I didn't do anything. I swear to God, I just sent a one sentence prompt with a goal and it just created all these. I can dodge. It's so complete. I have a little skill here. He has a special move called a series Z. Let me see if I can show it to you. Yeah. So it says mega round raised $40 million. It's just so funny. And his sword is a sourdough. You see, that's the sourdough bread for a sword. It's just one of the most hilarious games ever. I mean, not ever, but I've taken Astra to make games and it never got nearly this far and it never finished. So yeah, that's this game. It's pretty complete. It could go on for a long time. So it should probably stop. But I guess you can see a little bit of the side comments here. You can see it says I'm a thought leader. You're on mute. HR will hear about this. It's just such a silly, it's just so silly. And then all the different moves and stuff it has. So what were the moments that your jaw dropped? Was that why did it drop? I think for me there were a few moments where it started to self-evaluate itself. And this is why it also took so long, but it caught things that I would have never caught. And it caught like, I couldn't play the game at all during the week. This is actually my first time playing it, to be honest with you. I watched Claude play it in his testing. I was like, oh my God, I can't believe it's going this far. And let me just close it. Cause I can no longer hear you guys, but it would create a table and it gives itself a score based out of 10 on visuals, on content and things like that. And it would catch all these things. Like it'd be like, this could be very offensive to a mentally ill person. I'm going to change the words of this t-shirt. At one moment I was reading through its evaluation of itself and it said, you can actually see this female character's underwear if you go underneath. It was so detailed and thorough in its work. And it said, I'm not stopping until I get at least eight out of 10 on everything, which it never did. And because it ran for five and a half days, but the thoroughness of it too—the grass that you saw, it had wrote its own shader to create that grass just all on its own. And I think anyone that has created games in Astra before or previous models of Claude, they know that it doesn't get quite that far. You know, it takes a lot of follow-up prompting and the first version you get in Unity or Unreal, it just does not look like that. The music sounds bad. The sound effects sound bad. And this one had everything from pretty good graphics. It was, you could tell it was pretty low latency and the FPS was around 60. And it was the content. It was so funny. This was like the funniest model ever. The sourdough bread sword and the series, the ultimate, and like how your rent keeps going up the further you progress in the game and at the end of level one, you fight Carl the fog. And it turns out Carl the fog actually loves San Francisco and just wanted to hug the city or something. It was just so silly. So I was just in love with that whole time. He was making it, I would peek over to the side and see it was testing the game and I was like, how's it so good? But just one, literally one prompt. Zero. I promise there's zero follow-ups. Wow. That's pretty incredible. For one shot. We now have Dan Shipper who has joined us coming out of his vibe checking and testing hole, because there is also a new model out. So it's a really busy day in every world. Dan, how are you holding up? I am good. I've got a Celsius about to open. I've got two waters. So we're staying hydrated. I'm psyched to be here. I've been live checking all morning and my tweet is out for the GBT6 Soul. And I also, Kate, just filed the GBT6 Soul versus Opus 5.5 piece, which should be. I think Kaushik's got it. So it should be out in the next few hours. You're going to be getting the Opus 5.5 vibe check momentarily in your inboxes. Thank you all for your patience. We normally like to have it out right when the model drops, but we were testing multiple models, given that we were also working with the open AI one and you will get Dan's report on the models comparing each other in a little bit after that. So thank you. Yeah. Indeed. So I will go, maybe I can go through a little bit of my—I don't think we've done any tests, any test results on GBT6 Soul. So maybe I can go through that a little bit. And then I'll go through the comparison between the two. Okay. So basically GBT6 Soul just dropped for me. Big codex stan using Astra all day. It's my new daily driver. I still go to Astra every once in a while for top end type tasks, like knowledge work, top end knowledge work type tasks. But I think Soul is just got everything I need, but it's much cheaper for me on writing, on computer use, and on just general knowledge work tasks. I think it's the best model. However, it continues not to have the same top end as Opus 5.1. And as you guys have— and then I'll go through the comparison between the two. So, okay. So basically GPT-6 Sonnet just dropped for me, big codex, Stan using Astra all day. It's my new daily driver. I still go to Astra every once in a while for top end type tasks, like knowledge work, top end knowledge work type tasks. But I think Sonnet is just, it's got everything I need, but it's much cheaper for me on writing on computer use and on just general knowledge work tasks. I think it's the best model. However, it continues not to have the same top end as Claude 5.1. And as you guys have, I think seen on the stream, the stuff that Opus 5.5 is doing is pretty wild. I have found that for my day to day work, Opus 5.5 does a lot. I'll ask it to go do something and it'll take 15 minutes to come back and GPT-6. Is that low efforts? Because I have not done that. I'm not playing around with it. You should go low because that's the whole thing with Opus, use it on low. That's where the magic lies. I should really try that. I have not done that. So I would say, take my vibe checking of Opus with that grain of salt that I'm mostly using both on high and just seeing what happens. And if you're doing that, you're doing the default. Yeah, it goes off and does a bunch of stuff. And Sonnet just returns a result, which I think is easier for me in the way that I currently work. But I want to take a second and we've got these, one of the things that we've been doing and I don't know how much we've gone through it yet, but one of the things we've been doing is starting to turn our testing into benchmarks. And we have this really sick platform that Mike and Natasha have been building, called Checks. And so my checks for my writing are now live. I just put it into the chat. Hopefully that shows up. My checks are now live for editorial and they'll take you through exactly why I like this model versus Opus on a lot of the kinds of day-to-day writing tasks that I do. Let me share my screen. Sometimes this doesn't work in Dia, so I may need to leave and rejoin. Let me know if you guys can see my screen. Yeah, the link is not on X, but it's checks.every.2 slash p slash dense editorial checks with dashes. I'm going to have to leave and rejoin because I'm in Dia and Dia does not allow me to do screen sharing, which is extremely annoying. Shout out to my friends Hirsh and Hirsh and Josh. You guys should fix that because I have to switch to Safari right now. So I will be back. Literally. While Dan's doing that, did you want me to start showing the checks platform that he's. Yeah, you can. Yes. And I'm back. Oh, no, not needed. I was just about to jump in and start showing checks, Dan, but you were too fast for me. Okay. Allow to share screen, share screen. Good. Okay. I think you guys can see my screen now, correct? Yes. Yes. All right. So here's checks. So we've got, these are my editorial checks and we've got different checks for different people. So Mike has his own checks, which I would love for him to share if he has not already. And here's the scoreboard. We've got the models here. You can see Astra is currently my top ranked model. Sonnet is it, Sonnet is the second ranked model. And Opus 5.5 is close to it, but not quite there. These are a bunch of different tasks that I do every day. So when you see benchmarks, they are typically made of lots and lots of theoretical tasks, you know, like SWE-bench or something like that, where they are real programming tasks, but they're not really from real people's work. And so when you see a new model come out and you see the benchmark score and it's a little bit better, it doesn't really tell you, is this model any good for my work? And so what we've been starting to do is create benchmarks from our actual day to day work so that we can score these models in a more rigorous way on our actual tasks, which I think will tell you a lot more about what is this model good for and what is it not good for than a more generic benchmark is. So this is my personal benchmark for editorial for writing. Specifically, one of the tasks that I test this model on is, can it give good feedback on an article? Another one is, can it write a good first draft? I think this is actually a pretty good one to score it on or to show the differences between them. So in this particular task, what I have is I have a document I've been writing about what a personal benchmark is, and I ask it to write a section, the opening section of a personal benchmark. And now we can see the results between GPT-6, Astra, Sonnet, Opus 5, GPT-5.6 and Opus 5.5. And why things scored the way that they did. So for example, this is the opening that GPT-6 Sonnet wrote: A personal benchmark is a growing set of tasks from your real work judged by your own standards. Say you ask an AI to critique an article opening and it misses that the headline is weak. Your correction becomes a check. Did the model catch the weak headline? Run that test across models and you can see which ones meet your standards regardless of who made them. As you correct more work, the benchmark becomes a better guide to which model is best for you and when it's time to switch. GPT-6, Astra, a personal benchmark is a set of tasks drawn from your work, real work that will work with checks that capture what a good result looks like for you. And you can see my checks are here. So did it pass each of these checks? Let me compare this to Opus 5.5. A personal benchmark is a test built from your own work. It's a set of tasks you've actually done, each graded by simple yes or no checks that capture your judgment. Actually, I really like this. This is not the typical Opus 5.5 answer. One of the things that I found that Opus 5.5 does is it gives a lot of these longer answers that tend to get a little bit sidetracked. So let me just see if I can find a good comparison for these models. Let me see. You know, my ability to find the exact case that I want to show you guys is a little bit harder on this specific task, but I will say, go check this out. I will find a couple of cases to show you in particular, but my general take on Opus 5.5, and you can see some of these checks here. My general take is that it's very smart. It gives you big, long, chunky answers. And it is not quite as good at putting the main idea right at the top of the intro to an article, the opening paragraph, or any sort of revision. And so the writing that it produces takes a long time to get to the point. And for me, I really like more minimal, straight to the point prose. And I find that GPT-6 Sonnet and Astra are much better at that overall. Opus 5.5, pretty strong writer. Mike, I know you've done a lot of testing here. Do you want to talk about this as well? And you can see some of these checks here. My general take is that it's very smart. It gives you big, long, chunky answers. And it is not quite as good at putting the main idea right at the top of the intro to an article, the opening paragraph, or any sort of revision. And so the writing that it produces takes a long time to get to the point. And for me, I really like more minimal, straight to the point prose. And I find that GPT-6, Sol and Astra are much better at that overall. Still 5.5 is a pretty strong writer. Mike, I know you've done a lot of testing here. Do you want to talk about this as well? Yeah, let's do it. So I'll show my screen. All right. So looking at my checks, if you guys can see that. I don't think so. Now we can. Yeah, okay, cool. So this is available on the public site. So checks.every.to, and it's here, I'll show you this. It would be like for slash P for such Mike's checks. If you want to see the public one, this is my private one. It has a few things held back. I'll share with you guys in the live stream. But the main thing that I think was really interesting about Opus 5.5 is that it made some decisions that I actually preferred. So it didn't always follow exactly the right instructions. And you can see that here. So if you look, this is the PowerPoint task where we say create a PowerPoint in the every brand style guide, and you'll notice there's some of these not as good versus others. You see Deep Seek did a really terrible job, for example. But the brand style guide was supposed to be green, right? And Opus 5.5 went a little bit off topic and made it black. Which actually I preferred. And you know, it did a pretty good job here. The other thing it did is it made its own style of SVG here. So you can see it's all fuzzy. And that's at first I thought it made a mistake. And then I realized that as you go through, it actually kept that unique style throughout. So here's another example here. And so what it did is it made a couple of stylistic choices that after looking at it and thinking, oh, it kind of got marked down for not following a brand style guide. It's still on brand as in it's still something that I think we would be happy putting live. Even if it doesn't specifically follow the green or it used a slightly different SVG, it made good choices. And it's almost like the type of thing you would expect to see if you gave this task to a human and they had some opinions on how to extend the brand. That's kind of the vibe I'm getting from it. The other thing it really nailed is our Hoboken map. And I really liked this one visually as well. So if you see here, this is, I moved to Hoboken at the start of the year, I wanted to understand what restaurants I should take my kids to, what are the hidden gems. And it just did a really good job of picking out different places that I've since been to and I know are really good. And also just making the map look really nice. And so if we compare that to other models, usually the models would go for this weird, like I guess it looks like very outdated, like the open street map style. And it just doesn't look as nice. And sometimes the models completely fail at this. Here, the old Opus, and this is the main thing that made me dislike the old Opus, Opus 5, it actually decided to make its own map, right? And it decided to make its own SVGs in order to create this map. And it did so much more work than it needed to do. And I think the new Opus has solved that. So I was pretty happy to see a few of those things. So Mike, I don't know if you've said this already, but you were one of those recent converts along with Tyler to the Codex universe. Where are you now given your testing of these two models side by side? Yeah, so I would say that what has happened is the new OpenAI models have improved from generation to generation. And they're much cheaper now. I saw that the lunar pricing in particular is very tasty. But they made them a little bit more timid. And so I was trying to use Sol for a lot of last week. And I kept coming back to it and finding that it had stopped. And that was annoying enough that I was really excited when Opus 5.5 started to work really well, because that was tipping the balance enough for me. So I would say I'm currently at about 50/50. I still have some old threads in Codex. I still like the Codex harness. But now that I don't have to use Fable for everything in Claude, and I have a daily driver that's realistic to use and is cheap enough to use for everything, I think I'm making the switch now. This is one of the things I should add and say. We definitely found this. I think this is maybe part of what was going on for you is OpenAI added more strict security classifiers to the Codex app over the last week or so. I guess probably in response to all the security incidents recently. So it makes sense that they did this. However, it really decreased the usability of these models on long autonomous tasks where I felt like I kept having to say, I approve, I explicitly approve this, I approve this. And the approval would be counted as very narrow. So I think you will probably find in Codex and in ChatGPT for work that they get denied access. If you're using the approve for me default, they get denied access to doing a lot of things. And that's quite frustrating. Yeah, to some degree, Anthropic had a head start here because they had to introduce this classifier for Fable. And now I think they've rolled that out for Opus. And so they've had a little bit of extra time to fine tune this. That is true. Everyone, I am psyched to be doing this. I have to run, but I just wanted to drop in, talk about six Sol, show a little bit of the check stuff. We're going to have a lot more on checks for you soon. Would love your feedback on that. And we just published our vibe check on every.to. You can see the full Opus vibe check with everything on there. But I'm going to leave you guys to it to continue and hop off to a couple other things I got to do today, but good to see you. See you guys. See you. I can share some things about coding. So I've been coding with all the models and now since they're both out, we can actually compare them. And I think it's interesting because they launched them on the same day. So clearly they are saying we're competing here and we're going head to head. Or, yeah, Opus and Sol are similar class in weight. So let's compare them. You soon. Would love your feedback on that. And we just published our vibe check on every.to. You can see the full Opus vibe check with everything on there. But I'm going to leave you guys to it to continue and hop off to a couple other things I got to do today, but good to see you. See you guys. See you. I can share some things about coding. So I've been coding with all the models and now since they're both out, we can actually compare them. And I think it's interesting because they launched them on the same day. So clearly they are saying we're competing here and we're going head to head. Opus and Sol are similar class in weight. So let's compare them a little bit because that's cool. What I've been trying out obviously is coding using compound engineering, which is the plugin I have, how to do coding. So what's really important for me is does it follow instructions. For me, that's the most important thing because if you have AI that follows instructions and you have a bunch of instructions, it will do the right thing ideally. That's the whole point. So comparing the two effort levels. Coding, interestingly, was won by Opus 5.5 over Sol. Opus and Fable, actually Astra was in there as well. So hard Ruby coding problems with performance, optimization. Medium run of Opus 5.5 won, which is interesting. So it's not even extra high or high, medium run won. I think, and you shared also metric of a benchmark where it said that medium actually had the best coding example somewhere. So medium is an interesting one for coding and where Sol 6 is really good is it's just faster. But it's also confusing because the effort levels of Sol you cannot really compare to the effort levels of Opus 5.5 because they're so different. I think speed wise, the low in Opus 5.5 is kind of comparable. Anthropic shared that obviously. But yeah, for me, medium is a special one. It follows very well. Opus 5.5 follows everything in compound engineering and Sol is struggling. I've had to rerun so many times. So for me, Opus 5.5 is just more usable in a plugin long running setting. It might be that they're doing these new classifiers for security, but I don't really care what it is. I need to get work done. And Opus 5.5 gets the work done, even though Sol is sometimes faster, but it's not faster if you have to retry if you're in the loop. So I can share my screen and share some side by sides because that's fun to see the differences. So first of all, one side by side is Opus 5 versus 5.5. This benchmark, this is just me saying, create a one minute video about a lighthouse and you'll see a little bit, but the difference on the right side with Opus 5 and on the left side, we have Opus 5.5. You don't hear there's music as well, but this is a 0.5 step difference, which is insane to me. So the details on the left versus the right, even though the right is also pretty impressive. And the camera movements are really impressive, I think. You can see it swoop in like this. Obviously, 5.5 has more detail. This is what Tyler says as well. I'm really impressed by that. So obviously, it's better than 5. Is it as good as Fable? Well, on this one, the cozy island left is Fable, right is Opus 5.5. I think they got similar things right. I would say Opus 5.5 has the time of day, which Fable doesn't have, and maybe even a little bit more detail. So it is a Fable class model. Then what if you compare the same to Sol? Because Opus Sol should be similar. So left is Opus 5.5, right is Sol 6. Well, it is different. It's just when I was looking at this, I was laughing, not that I want to laugh at the worst model, it's just so wildly different. And they're both on extra high. And they should be the same. And obviously, they post trains on 3D movements and things because this got an insane step in how it performs. One last thing, a more simple example is a landing page. Both of these are low or actually, they're both low and they finish in about two minutes or so. And you can see, for me, it's about how do I look at these? What is the difference? It's not what looks better because that's pretty subjective. But what I see that is different is that on the left, Opus 5.5 low. And on the right, 2t6 solo. I see with Opus that it feels a little bit more human. What I mean by that is more space. If I read it, it feels more like it's designed for me as a human. On the right, for example, this is very in my face. I need to read this and then this and then this with different sizes. If I scroll down here on the left, this feels easier to scan. For example, the copy itself, I like a little bit better. It's a little bit more descriptive. And it leaves more space for me to fill in. And I really like that from this model. And if you go to higher effort levels, it just does other things a little bit more. But it doesn't overdo things. It doesn't go very extreme. So if I go for OpenAI Max, it looks cool. But it is a lot. There's lots of things everywhere. So that's my side by side for these. And let's continue the program. Yeah, Tyler, I think you had something related to Pokemon that you wanted to share. I feel like you had the most fun out of anyone with this model over the past week. Yeah, let me pull that up one sec. All right, so here is a what if Pokemon for real simulation. I'm going to play it at 10 minutes here. Actually, let's speed it up a bit. So it's daylight. And you can see the actual Pokemon and humans in action here. Yeah, it's just, I had a great time with this. I'm not a huge Pokemon fan, but my wife is. So I kind of wanted to just build it for her. But it was just fun to see how the Pokemon and humans interacted in this world. Oh, this is Opus 5.5. Yeah, we generally like Opus 5.5 for coding and for these creative things more. Yeah, this was just something that was really fun. And it created this really awesome report, which I will pull up since this sim doesn't really show everything properly. The sim ran for a really long time. So let me just pull up another one. Is this also Opus? You're all showing Opus stuff, right? Yeah, I'm only showing Opus stuff. So here's actually the artifact it made afterwards, describing the simulation. And again, Opus did all of this. It added videos here for me to review. I thought it was so cool how it did this. It really went through really cool examples that I could watch. And all this was happening while I was working. So this is interesting where some people died in the cold weather, I believe. And it showed how the Pokemon interacted. Showing some evolutions here, I believe. And just going over the different habitats where different Pokemon took over or made as their main territories. The sim ran for a really long time. So let me just pull up another... Let's see if this shows... Is this also Opus? You're all showing Opus stuff, right? Yeah, I'm only showing Opus stuff. So here's actually the artifact it made afterwards, describing... And again, Opus did all of this. It added videos here for me to review. I thought it was just so cool how it did this. It really went through cool examples that I could watch. And all this was happening while I was working. So this is interesting where some people died in the cold weather, I believe. And it showed how the Pokemon interacted. Showing some evolutions here, I believe. And just going over the different habitats where different Pokemon took over or made as their main territories. So this was a really fun one. Tyler, I wonder if... When you do these sorts of more ambitious projects that are running a long time in the background, do you ever ask it how it did that, right? Like, did you get more details on... Because when I look at this, I think, wow, what... Did it use a game engine? Or did it make its own thing up? Yeah, I have all these questions. But I'm wondering if you explored that. At the risk of sounding like a total noob? No, I didn't. I just said it on UltraCode. I gave it a goal and said, I want to create a simulation of a world where Pokemon actually existed alongside humans. And that was it. I think it was a very short prompt. And this one ran for, I think, a day and a half was how long it ran. And I felt my computer was going to melt at times where I got really nervous. I actually took some things out of the freezer and placed it on top of the Mac. I live in Hawaii. So what's going to go wrong? It can be like 80 degrees. I know, I was watching for condensation, but I was thinking, man, I gotta cool this down a little bit. And my Mac Studio had never really got that hot before where I could hear it just pumping. So it went to work, man. It went to work. So you use all these fun things. Did you use it for actual work, like design-wise? Because you do design work. Like, how do you feel it is for design work as a designer? Not like me, not being a designer. Like, is this the model? Opus, Fable? Which model, Astra, which model do you use if you do design work now or going forward? You know, I still use Grok for right now, Grok 47 for real work because it's so cost efficient. And I feel I'm able to steer it well enough so I can save my usage limits for the entire month. But if I think if tokens were free, there would be no doubt I'd be using this model. And I can even show some examples of— Which model? Which model is this? Opus 5.5. Opus 5.5, 100%. I'm only going to be talking about Opus 5.5. Huge fan. But let me show you just a few things. Here's a component generator, or this is the old Opus. And here's some improvements in the quality and design of these things where it just got better. Here's something fun I made where it's a slot machine that generates UI components for me. So just really fun stuff. And this was also one prompt. This was a more complex prompt. I don't quite remember exactly what it said. But again, that was one prompt. And I think also it built this design tool that I need to share with you here. I'm going to actually give me a sec to import for the design. So I'm going to import. For those of you that have ever heard of Origami Studio. It's a prototyping design tool. Very hard to use. It uses something called patches. And each patch is almost a functioning code. And I'm just going to show you how it works here. But this is an app that I'm working on in personal time. My personal time is kind of a... So this is not Origami Studio. This is your own version. This is my own. This is my own. This is my own. You started with Opus 5.5. Yep. Just a little bit ago. Maybe about four prompts for this whole thing. No joke. The first prompt was build an Origami Studio that was AI friendly. And I think instead of explaining how it works, it's a very complex tool. I'm actually using cloud code to control it in different terminals. So I'm going to say add a swiping interaction. Agent A. What you're going to see. Yeah. What you're going to see here are patches on the bottom. These patches are kind of functions in code. And these are actual real design layers. So you can see the design layers on the left here. You can see the inspector panel on the right here. I also had built in an assistant. So I can put in my own Anthropic key if I want to just chat in there. And right now it's working in the background. And in the moment we should see a bunch of patches appear in the bottom. And I'll be able to show you exactly how this works. And I'm hoping and planning to use this for my actual work. Because what this allows me to fine tune the interactions. In a way that is node based and I can see everything on a canvas. Okay, so here we go. So you can see. And let me just show only the nodes here. And show the prototype. So it's a little messy, but it created all these nodes here. Let me see if I can tidy it up. And this just built this tidy up feature, which is in Origami. So that's tidied up. So let me show what happens. Let's see if I drag. So you can see if I'm dragging it. You can see all of these numbers changing here. And I can actually adjust all of these things. So in here somewhere there's probably a scale animation. I can see there's an opacity animation on this yes badge, which I actually don't like. So I'm going to disable that. And disable this one too. And let's see if it should go away. So you can see that heart disappeared here. So this is a way for me to fine tune and play with, and just have a lot of fun with things like the spring here. And the bounciness. I can max these out and make it really crazy. You see how much quicker that went? Maybe I'll make it slower. I'm not sure if I could see the difference there. It happened so quickly. But this is a tool that I'm really excited about. And it's just so complete. So if I start a new file here, it came with tutorials. It came with templates to start with. So I can, it's just really, I don't know. I'm just really blown away. And then I can also find it. So do you think you'll actually use this? Because with all these tools, it's always like, yeah, that's great to share on X. And it looks amazing. Wow, wow, wow. But will we use it every day? Similar with games. Will we actually play it for more than two minutes? Is it actually worth our time? Do you think you're going to use this tool and develop it further and actually have it be part of your workflow? Because I'm always skeptical about these things because yes, it's cool that we can create them and it's insane. But is it going to help our lives? Are we actually going to do stuff with it? I think it depends person to person. For me personally, I'm super excited about this tool because I've been trying to learn how to use Origami Studio for a while. Many times I've tried and given up. They don't update their documentation. And to learn it, if you get stuck, you have to go into a Facebook group and hope that someone will respond. And even when they do respond, they will tell you how to do it, but it doesn't really sink in because you don't understand the complexities of the tool and maybe things have changed. They have a tutorial you can go through and that's not even up to date. Because I'm skeptical about these things because yes, it's cool that we can create them and it's insane. But is it going to help our lives? Are we actually going to do stuff with it? I think it depends person to person. For me personally, I am super excited about this tool because I've been trying to learn how to use Origami Studio for a while. Many times I've tried and given up. They don't update their documentation. And to learn it, if you get stuck, you have to go into a Facebook group and hope that someone will respond. And even when they do respond, oftentimes they will tell you how to do it, but it doesn't really quite sink in because you don't understand the complexities of the tool and maybe things have changed. They have tutorials you can go through and those aren't even up to date. So it's something that I really want to learn because I noticed that many of my favorite designers, the ones that I really admire, are very good at using this tool. So for me, I am building this tool for myself so that I can actually learn how to use Origami on my own, but also to use AI with it. So I can tell it, hey, add this swiping interaction or when I click on this, it should scale up to the whole screen. And I can learn about what these things are. This is what I added here on the right side. It's a way for me to learn what these things are. So I'm hoping that this is a tool that actually makes me a better designer. And who knows, maybe it's something that I do use long term. Maybe not. I added some features where I can vibe design in here. I'm showing that holographic animation. So maybe it becomes my design tool. And I'm not in Figma all the time. But no, I don't see Figma going away from me anytime soon. But how long did you work on this? Did Opus work on this again, that someone asked? I think total, on and off, about three days. And that's not three days nonstop, because I was working. The first prompt was pretty long, really describing what the goal was, what Origami Studio is, and how I wanted it to be able to be manipulated by code. I think the second prompt was, okay, I need to actually import real design from production code. So that app you saw was something I vibe coded with Astra, actually, and then imported that Swift UI app into this canvas. Then the third and fourth prompt was, I want to be able to design new designs kind of like how Paper does it if you're familiar. And then I think the last one was I want to add loading states. So I created that holographic animation. That one took a bit. I had to design a little bit of the, had to provide some references and really describe it, but it did that holographic thing in one shot. Also, there's some loading states for the nodes that didn't show. So that's a live demo bug there. But I was going on and off three days. Great. You dropped another name. You said Grok 4.7 as well, which also dropped yesterday. So there are actually three models. And Mike, I know you ran some benchmark on it as well. I did too. So maybe we can just briefly talk about how Grok 4.7 compares to Opus 5.5 and Claude Haiku and Sonnet, and then after that, maybe we can finish the live stream and get back to using it. Yeah. Let me show you some stuff. So I was really surprised by Grok 4.7 because Grok 4.6 is a great model and surprisingly good, better than I expected it to be. And when we started testing 4.7, it feels like it was rushed out. And actually, at least on my benchmark, it was a real regression. And I don't know if that is related to maybe there were just a bunch of bugs or things for them to figure out. I know you said that you saw it getting better over the previous week. What was your experience with it, Tyler? The first few days were rough. It was definitely rough. And I think, you know, maybe they weren't ready for us to test it, but it did, by the end of the week, it did become a lot more reliable. But there were some things early on where it was giving up a little bit or it was hallucinating. There was a lot of issues with probably their servers at the time. Like they were probably working on it really hard while they gave us access. There was some downtime for Cursor last week, I think one day. Yeah. So there were some issues around there. But you said you use it and you like it, right? So how do you use it today? Well, I would say that I'm still figuring out how I love to use it. And I think this is the first time that for whatever reason I've moved back to the Cursor IDE over the Agents window. It just seems to work a little faster. Maybe it's just my imagination. But it seemed to work a little faster in the IDE, especially during testing. And I've kind of just decided to stay there and work locally as opposed to working in Cursor projects, which is something I really love. But I'm just not having the same success that I was having with Grok 4.5 and especially Grok 4.6 inside of Cursor projects, which is my preferred way. One of the reasons why I use Cursor so much is because Cursor projects allows me to have one chat for whatever I'm working on. And it will create all these sub agents in one place and I can have another project underneath it. So it really allows me to multitask well, but it's still to be determined if Grok 4.7 is really good at orchestrating those projects to other sub agents. But my previous setup, I really enjoyed having Claude 5.1 low or Claude 5.1 medium as the orchestrator and having Grok 4.6 as all of the sub agents. It just worked really well. So I'm hoping for Grok 4.7. And I've been using it today and yesterday. I've had no problems. In the testing, it was a little rough early. There were some interesting moments. It was DMing one of my PMs without telling me. And I realized later, I was like, oh my. I was like, Marcus, I'm so sorry. I think it took some things I said super literally. For example, I will say something like, you know, let's get this ready to merge to main. And then it just merged it to main and DMed Marcus. And I was like, oh no. Big no-no at every company, right? We always have humans do the squash and merge here. But yeah, as the week had gone on, it had improved. I actually am starting to like it a little more than 4.6 actually. As of this morning, before I came on, I had worked on and shipped the PR and had no problems. It does feel like an improvement, but it's one of those improvements where Grok 4.7 isn't a one-shot show horse type of thing. So it's hard for me to evaluate it. It really is. And when I use it for real work, 100% of my real work is almost using Grok. So it's really hard for me to say, is it noticeably better? No, but it seems to be working great. And I don't plan on switching, but I do wish I had some withdrawals for Opus. What did you find, Mike, in your testing? Yeah, so I feel like maybe we need to retest if it is changing that much, right? But I found some of the tasks that did pretty well. So this was the MPS dashboard. This is a different design than what I've seen from other models. So if you look at say what Claude is, so we're looking at Grok. Yeah, so this is Claude. And then this is Grok 4.6. And I really like the headline here and the design. So I was really happy with 4.6. Just in comparison, this is Sonnet 4.6. No, but it seems to be working great. And I don't plan on switching, but man, I do wish I have some withdrawals for Opus. What did you find, Mike, in your lesson? Yeah, so I feel like maybe we need to retest if it is changing that much, right? But I found some of the tasks that did pretty well. So this was the MPS dashboard. This is a different design than what I've seen from other models. So if you look at say what Fable is. So we're looking at Grok. Yeah, so this is Fable. And then this is Grok 4.6. And I really like the headline here and the design. So I was really happy with 4.6. Just in comparison, this is Sol 5.6. And the, let me find one that didn't do as good a job. Maybe Opus 5. Yeah, this dark mode thing. But yeah, Grok 4.7 did okay on this one. But just to kind of give you a sense of my issues with it. It likes a lot of the tasks that just completely didn't finish. Right. So it didn't end up outputting the final result. So there was some straight up errors there. And by the way, I reran these tasks like three or four times. And then some of the ones where it did finish, which makes me think it's not just a harness. Let me take a look at the PowerPoint task. Let me see. So was it this one? Yeah, it was kind of, I'm trying to find the example specifically, but I, yeah, this isn't the one I was thinking of here. I think I might've overwritten it. But yeah, I found it was a little bit messy. Let me just see the write up example. Yeah. So here's a good example. The writing here. It was just a couple of steps backwards. Relative to the previous model. It's just kind of. It kind of read as a to do list almost rather than an article. So that was one issue. And then a lot of the stuff, it just didn't do that good of a job. The design wasn't as good and stuff. So it just felt like a little bit of a step backwards. And because it was unreliable, it became something that we were less interested in. But I hope that I have really high hopes. I hope that the reliability issues improve. Yeah. So for me, I did just some A-B tests here. Left is Grok 4.7, right is Opus 5.5. It's not as great at 3D here. And on the left, Grok 4.7, this is the video. So the video is really trying to show, can it do very good planning, can it iterate, stuff like that. So left is Grok 4.7, right is GPT-6 Luna. So if you look at the two, it feels like the 3D action of Grok 4.7 is more in the Luna realm. I would say Luna looks still a little bit better. So from a 3D perspective, it's not very impressive. From coding, yeah, in LFG didn't do either very well. And maybe that's because I'm using workflows. So clearly, I'm using this in Compound Engineering as a workflow, which needs to follow all the steps. And if it doesn't follow all the steps, it's harder to make it work. But I did like, or I do like 4.6 as just a workhorse. But also, it gets very hard because Luna has a workhorse, Sol is a workhorse, Opus 5.5 is a workhorse. There are so many workhorses and suddenly workhorses get so good that they're better than Able, which is insane. And just having this entire landscape change so much. Sol 6 is way better than 5.6. But there's Opus 5.5 as well. And it just changes everything. If you're curious, check out Opus 5.5, but also check out all the other models. I think for writing, especially Dan, who really liked Sol, check out Opus 5.5 on low effort as well. I think that's something to try out because if you do high, it can feel a little bit slow, but you don't really need slow. You can just put it on low. And yeah, any other tips or tricks to take away before we hang up Tyler, Mike? I think the main thing just to think about is the pace that these new models are coming out in and how hard it is to test them and see the differences between them. We're kind of getting to the point where we're largely saturated on a lot of tasks. If you can think it, you can do it. And it's starting to come down more to personal preference. I don't know about you, Kieran, Tyler, but I've started to care less about is this the best model in the world, advanced math or whatever. And instead, do I like the designs it comes up with or do I like the writing? And I've become more comfortable with being subjective. And I use all models. I literally use all models. It's not that I don't use one over the other. But if I make 3D games, I probably won't use Grok 4.7. I will probably use Opus 5.5. But I also love Cursor. Cursor is my app. I use all my tokens in Cursor because Cursor Projects is amazing. And Codex is also amazing. And Claude is catching up. The app. It's not as amazing as the others. But I do use all the apps, all the things. And yeah, if you need to choose, hopefully you learn something. You don't need to switch because they're all very good. But yeah, Opus 5.5 is a special one. I really think it deserves a special mention. Or if there's a winner, for me it's Opus this week in launches for sure. Yeah, definitely. And I feel like I just broke up with Opus too. So I'm all the way back. It's really funny how quickly I came back to Opus. But I agree with Kieran. I use all the models. I don't see myself going away from Cursor and Grok for actual work. I'm not going to get rid of my Grok bots. Love them too much. For anything in Blender, I actually really love using Astra with Computer Use Control Blender as opposed to MCP. I can't get rid of that. And this Opus thing is just magical. Hurts my wallet. But I really just can't put into words how awesome Opus 5.5 is. Sounds like a good place to end. Great. Thank you, everyone. Every is the only subscription you need to stay at the edge of AI. We said it. Oh, yeah. Go subscribe. We forgot to say they've ramped the line. Yeah, we're such horrible presenters. Every.to. Yeah, go check out our Vibe Checks. Yeah, our Vibe Checks will arrive today. Either they're live already or they will arrive. We're doing comparisons there as well. We'll be sharing more on X and LinkedIn and other places. But thanks for hanging with us and see you next time. Happy Release Week, everyone. any tests, any test results on GBT6 soul shit. So maybe I can go through that a little bit. Um, and then I'll go through the, um, like the comparison between the two. So, uh, uh, okay. So basically GBT6 souls just dropped for me, big codex, Stan using Astra all day. It's like my new daily driver. Uh, I still go to Astra every once in a while for, uh, like top end type tasks, like knowledge work, top end knowledge work type tasks. But I think, uh, soul is just like, it's got everything I need, but it's like much cheaper, uh, for me on writing on computer use and on, uh, just general knowledge work tasks. I think it's the, I think it's the best model. Um, however, it continues not to have the same top end as fable 5.1. And, and as you guys have, I think seen on the stream, the stuff that Opus 5.5 is doing is like pretty wild. I have found that for my day to day work, Opus 5.5 does a lot. Like I'll, I'll ask it to go do something and it'll take like 15 minutes to come back and, and GBT6. Is that low efforts? Because I have not done that. I'm not playing around with it. You should go low because that's the whole thing with Opus, like use it on low. That's where the magic lies. I should really, I should really try that. I have not done that. So I would say, take my, take my, uh, my vibe checking of, of Opus with that grain of salt that I'm, I'm mostly using both on high and just seeing like what happens. And if you're doing that, you're doing the default. Um, yeah, it goes off and does a bunch of stuff. And, and, uh, solo just like returns a result, which, which I think is, you know, uh, it's just, it's a little bit easier for me in the way that I currently work, but, um, I want to take a second and we've got these, uh, one of the things that we've been doing and I don't know how much we've gone through it yet, but one of the things we've been doing is starting to turn our testing into benchmarks. And we have this like really sick platform that Mike and Natasha have been building, uh, called checks. And, um, so my checks for my writing are now live. Um, and I just put it into the chat. Hopefully that shows up. Uh, my checks are now live, um, for editorial and they'll, they'll take you through exactly why I like this model versus Opus, uh, on a lot of the kinds of day-to-day writing tasks that I do. Um, let me share my screen. Sometimes this doesn't work in Dia, so I may need to leave and rejoin. Uh, let me know if you guys can see my screen. Yeah, the link is not on X, but it's checks.every.2 slash p slash dense editorial checks with dashes. I'm going to have to leave and rejoin because I'm in Dia and, uh, Dia does not allow me to do screen sharing, um, which is extremely annoying. Um, shout out, shout out to my friends, Hirsh and Hirsh and Josh. You guys should fix that. Uh, cause I have to switch to Safari right now. Um, so I will be back. Literally. Uh, while Dan's, uh, doing that, did you want me to show, start showing the checks platform? Uh, that he's. Yeah, you can. Yeah. Um, yes. And I'm back. Oh, no, not needed. I was just about to jump in and start showing checks, Dan, but you were too fast for me. Okay, cool. Uh, allow to share screen, share screen. Good. Okay. I think you guys can see my screen now, correct? Yes. Yes. All right. So here's, here's checks. So we've got, this is, these are my editorial checks and we've got different checks for different people. So Mike has his own checks, which I would love for him to share if he has not already. And here's the scoreboard. Uh, we've got the models here. You can see Astra is currently my top ranked model. Um, soul is it look, soul is the second ranked GBT six soul is the second rate ranked model. Um, and Opus, uh, 5.5 is close to it, but not quite there. These are a bunch of different tasks that I do every day. Um, so these are, so when you see benchmarks, benchmarks, they are typically, uh, made of like lots and lots of theoretical tasks that, um, uh, you know, like sweet bench or something like that, where they're, they are real programming tasks, but they're not really from real people's works work. And so when you see a new model come out and you see the benchmark score and it's a little bit better, it doesn't really tell you, is this model any good for like my work? And so what we've been starting to do is create benchmarks from our actual day to day work, um, so that we can score these models and a little bit more of a, a little more rigorous way, um, uh, on our actual tasks, which I think will tell you a lot more about what is this model good for and what is it not good for than a more generic benchmark is. So this is my personal benchmark for editorial for writing. So specifically, um, one of the things that one of the tasks that I test this model on is, um, can it give good feedback on, uh, on an article? Another one is, um, can it write a good first draft? I think this is actually a pretty good one to, um, to score it on or to, to show the differences between them. So in this particular task, what I have is, um, I have a document I've been writing about what a personal bench, about personal benchmarks. And I ask it to write a section, the opening section of a personal bench of the personal benchmark task. Um, and now we can see the results between GBT 6, Astra, Sol, Opus 5, GBT 5.6 and Opus 5.5. Um, and why things scored the way that they did. So for example, um, uh, this is the opening that GBT 6, Sol wrote a personal benchmark is a growing set of tasks is, is a growing set of tasks from your rule work judged by your own standards. Say you ask an AI to critique an article opening and it misses that the headline is weak. Your correction becomes a check. Did the model catch the weak headline? Run that test across models and you can see which ones meet your standards regardless of who made them. As you correct more work, the benchmark becomes a better guide to which model is best for you and when it's time to switch. Um, GBT 6, Astra, a personal benchmark is a set of tasks drawn from your work. So I'm going to do a little bit of work with the real work that will work with checks that capture what a good result looks like for you. And you can see my checks are here. So did it pass each of these checks? Um, let's compare this to Opus 5.5. A personal benchmark is a test built from your own work. It's a set of tasks you've actually done each graded by simple yes or no checks that capture your judgment. Actually, I really like this. Um, this is, this is not a, uh, this, this is not the typical Opus 5.5 answer. One of the things that I found that Opus 5.5 does is it gives a lot of these like longer answers with that tend to get a little bit sidetracked. Um, so let me just see if I can find a good, a good comparison for, for these, for these models. Um, let me see. You know, my, uh, my person, this, my, my ability to find the exact case that I want to show you guys is a little bit harder on this, on this specific task, but I will say, go check this out. I will, I will go find a couple of cases to show you in particular, but my, my general take on Opus 5.5, and you can see some of these checks here. My general take is that it's, it's very, it's very smart. It gives you big, long, chunky answers. Um, and it is not quite as good at putting the main idea right at the top of, uh, like the intro to an article, the opening paragraph, or any sort of revision. And so the writing that it produces, um, takes a long time to get to the point. And for me, I like me, I really like more like minimal, um, straight to the point pros. And I find that, uh, five, uh, GBT-6, Sol and Astra are much better at that overall. Still 5.5, pretty strong writer. Um, Mike, I know you've done a lot of testing here. Do you want to talk about this as well? Yeah. Yeah, let's do it. Um, so, uh, I'll show my screen. Uh, All right. Um, so looking at my checks, um, if you guys can see that. I don't think so. Now we can. Yeah. Okay, cool. So, uh, this is available on the public site. So the, uh, checks.every.to, uh, and, uh, uh, it's in, uh, here, I'll show you this. Uh, so it'd be like, for slash P for such Mike's checks. Uh, if you want to see the public one, uh, this is my private one. It has a few things held back. Um, I'll share with you guys in the live stream. Uh, but, uh, the, um, the, the main thing to, um, I think that was really interesting about Opus 5.5 is that it made some decisions that I actually preferred. Um, so, uh, it, like, didn't always follow exactly the right instructions. Um, and you can see that here. Um, so, uh, if you look, this is the PowerPoint task where, uh, we say create a PowerPoint in the every brand style guide, and you'll notice there's, you know, some of these, uh, not as good, um, uh, versus others. You see deep, deep seek, like, did a really terrible job, for example. Um, but, uh, the, the brand style guide was supposed to be, uh, like, green, green, right? Uh, and, uh, Opus 5.5, uh, went a little bit off topic and made it black. Um, which, uh, actually I, I preferred. Um, and, uh, you know, did a, did a pretty good job here. The other thing it did is it made its, its own style of SVG here. So you can see it's like all fuzzy. Uh, and that's like, at first I thought it made a mistake. Uh, and then I realized that as you, as you go through, uh, it actually kind of kept that, uh, unique style, uh, throughout. Um, uh, so like, here's another example here. And so, uh, what it did is it made a couple of stylistic choices that after looking at it and thinking, oh, it kind of got marked down for, for not following a brand style guide. Like it's still on brand as in, like, it's still something that I think that we would be happy putting live. Um, even if it doesn't specifically follow the green or it used a slightly different SVG, it made good choices. And, uh, it's almost like the type of thing you would expect to see if you gave this task to a human, uh, and they had some like opinions on how to extend the brand. Like, that's kind of the vibe I'm getting from it. Uh, the, the other thing it, it really nailed, uh, is our Hoboken map. Uh, and I, I really liked this one visually as well. Um, so, uh, if you, if you see here, this is, uh, you know, I moved to Hoboken at the start of the year, I wanted to understand what restaurants, uh, I should take my kids to what are the hidden gems. And, um, it just did a really good job of, uh, picking out, uh, different places that I've since been to and I know that are really good. Uh, and also just making the map look really nice. And so if we compare that to other models, uh, usually the models would go for this like weird, um, like, like, I guess it's like kind of looks like very outdated, like the open street map style. And, uh, it just doesn't look as nice. Uh, and sometimes the models completely fail at this, uh, like here, uh, the old Opus, and this is the main thing that made me dislike the old Opus, Opus 5, uh, it, it actually, um, decided to make its own map, right? And it decided to make its own SVGs, uh, in order to create this map. And it did like so much more work than it needed to do. Uh, and I think the new Opus has, uh, has solved that. So, uh, I was, I was pretty happy to see a few of those things. So, Mike, I don't know if you've said this already, but you, you were one of those recent converts along with Tyler to the Codex universe. Um, where are you now given your testing of, of these two models side by side? Yeah. So I would say that, um, what, what, what has happened is, uh, the new OpenAI models, uh, have improved, uh, you know, from generation to generation. Um, and, uh, much cheaper now. I saw that like the lunar pricing, uh, in particular is very tasty. Um, but, uh, but, um, they made them like a little bit more timid. And, and so I was trying to use, uh, the Sol, uh, for a lot of the week, a lot of last week. And, um, I, I kept coming back to it and finding that it had, it had stopped. And, uh, and that, that was like annoying enough that I was like really excited when Opus, uh, 5.5 started to work really well, um, because that, that, that was tipping the balance enough for me. So I would say I'm currently at about 50, 50. I still have some old threads and codex. I still like the codex harness. Uh, but, um, now that I don't have to use Fable for everything in Claude, and I have a daily driver that's realistic to use and, um, is cheap enough to use for everything. I think I'm making the switch now. This is one of the things I should, I should add and say, like, we definitely found this. I think this is maybe part of what, what was going on for you is opening. I added more strict security classifiers to the codex app over the last week or so. Um, I guess probably in response to like all the security incidents recently. So it makes sense that they did this. However, it really decreased the usability of these models on long autonomous tasks where I felt like I kept having to say, I approve, I explicitly approve this. I approve this. And the approval would be counted as very, very, very narrow. Um, and so I think that you will, I think they'll probably fix this. It's a fixable thing, but if you're trying these models out today, I think you will probably find in codex and in chat GPT for work that, um, they get denied access. If you're using the approve for me default, they get denied access to doing a lot of things. And that's quite frustrating. Yeah. To some degree, anthropic had a little bit of a, um, head start here because they had to introduce this classifier for fable. And now I think they've rolled that out for opus. And so they've had a little bit of extra time to fine tune this. That's true. That is true. Uh, everyone, I am psyched to be doing this. I have to run, uh, but I just wanted to drop in, talk about six soul show a little bit of the check stuff. We're going to have a lot more on checks for you soon. Would love your feedback on that. And we just published our vibe check on every, every.to. You can see the full opus vibe check with everything on there. Um, but I'm going to leave you guys to it, uh, to continue, uh, and hop off to a couple other things I got to do today, but good to see you. See you guys. See you. I can, I can share some things about coding. So I've, I've been coding with all the models and now since they're both out, we can actually compare them. And I think it's interesting because they launched them on the same day. So clearly they are saying we're competing here and, uh, we're going head to head to head or like, yeah, Opus and Sol are similar class in weight. So let's compare them a little bit because, um, that's cool. What I've been trying out obviously is coding using compound engineering, which is the plugin I, uh, have, uh, how to do coding. So what's really important for me is like, does it follow instructions? Like for me, that's the most important thing because if you have AI that follows instructions and you have a bunch of instructions, it will do the right thing ideally. That's the whole point. So, um, comparing the two effort levels. So coding, interestingly, um, was won by Opus 5.5 over Sol. Um, Opus and Fable, actually Astra was in there as well. So hard Ruby coding problems with performance, uh, uh, uh, with like optimization with although learn to optimize against, uh, yeah, like stuff, the medium run of Opus 5.5 won, which is interesting. So it's not even extra high or high than medium run won. I think, and that does, you shared also metric of a benchmark where it said that medium actually had the best coding example somewhere. So like medium is an interesting one for coding and where Sol 6 is really good is it's just faster. Um, but it's also confusing because the effort levels of Sol, you cannot really compare to the effort levels of Opus 5.5 because they're so different. Um, I think speed wise, the low in Opus 5.5 is kind of comparable. So it can be, uh, oh yeah, so Anthropics shared that, uh, obviously. Um, but yeah, so the, yeah, for me, medium is a special one. It follows very well. Opus 5.5 follows everything in compound engineering and Sol is struggling. I've had to rerun so many times. So for me, Opus 5.5 is just more usable in a plugin long running, uh, setting. It might be that they're doing these new, uh, new classifiers for security, but I don't really care what it is. I need to get work done. And Opus 5.5 gets the work done, even though Sol is sometimes faster, but it's not faster if you have to retry if you're in the, in the loop. So maybe, so I can share my screen and share some side by sides because that's kind of fun to see the differences. So first of all, one side by side is Opus 5, uh, versus 5.5. So this, this benchmark, this is just like me saying, create a one minute video, uh, about like something, uh, about like a lighthouse and like, you'll see a little bit, but the difference on the right side with Opus 5 and on the left side, we have Opus 5.5. You don't hear there's music as well, but this is like a 0.5 step difference, which is insane to me. So the details on the left versus the right, even though the right is also pretty impressive. And also the camera movements are really impressive, I think. So, you can see it swoop in like this. Okay, so obviously, okay, I'll stop this one. Obviously, 5.5 has more detail. This is kind of what Tyler says as well. Like, I'm really impressed by that. So I was like, okay, so obviously, it's better than 5. Is it as good as Fable? Um, well, yeah, on this one, the the cozy island left is Fable, right is Opus 5.5. Like, I think they got similar things right. I would say Opus 5.5 has the time of day, which Fable doesn't have, and maybe even a little bit more detail. So I'm like, okay, so it is a Fable class model. Then what if you compare the same to Sol? Because Opus Sol should be similar. So left is Opus 5.5, right is Sol 6. So what if you compare the same model? So what if you compare the same model? Well, it is different. It's kind of like, I like I when I was looking at this, I was just laughing not that I want to laugh at the worst model that like, it's just like, like, I mean, just yeah, it's just so wildly different. And they're both on extra high. And they should be the same. And obviously, they post trains on 3d movements and things because this got like an insane step in how it performs. One last thing, a more simple example is a landing page. Both of these are low or actually, yeah, they're both low and they finish in like two minutes or so. And you can see, for me, it's like, how do I look at these? Like, what is the difference? It's not like what does look better? Because that's pretty subjective. But what I see that is different is that on the left, Opus 5.5 low. And on the right, 2t6 solo. I see with Opus that it feels a little bit more human. And what I mean by that is like, at least more space, like it just like, if I read it, it feels more like it's designed for me as a human. On the right, for example, like this is very in my face. It's like maybe like, I need to read this and then this and then like this and different sizes. If I scroll down here on the left, this feels easier to scan. For example, the copy itself, I like a little bit better. It's a little bit more descriptive. And yeah, it's like, it leaves more space for me to fill in. And I really like that from this model. And if you go to higher effort levels, it just does other things a little bit more. But it doesn't overdo things. It doesn't go very extreme. So for example, if I go for OpenAI Max, like it looks cool. But also, it is a lot. There's lots of things everywhere. So that's my side by side for these. And let's continue the program. Yeah, Tyler, I think you had something related to Pokemon that you wanted to share. I feel like you had the most fun out of anyone in this model over the past week. Yeah, let me pull that up one sec. All right, so here is a what if Pokemon for real simulation. I'm going to play it at, let's go at 10 minutes here. Actually, let's speed it up for a bit. So it's daylight. And then you can see you can see the actual Pokemon and humans in action here. So yeah, it's just, it's really, I had a great time with this. I'm not a huge Pokemon fan, but my wife is. So I kind of wanted to just build it for her. But it was just fun to see how the Pokemon and humans interacted in this world. I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, Oh, this is Opus 5.5. Yeah, we generally like Opus 5.5 for coding and for these creative things more. Yeah, no, this was just something that was really fun. And it created this really awesome report, which I will pull up since this sim doesn't really... The sim ran for a really, really long time. So let me just pull up another... Let's see if this shows... Is this also Opus? You're all showing Opus stuff, right? Yeah, I'm only showing Opus stuff. So here's actually the artifact it made afterwards, kind of describing... And again, Opus did all of this. It added videos here for me to review. Like, I thought it was just so cool how it did this. It really went through like, really cool examples that I could watch. And all this was happening while I was working. So this is like kind of interesting where like, some people died in the cold weather, I believe. And it showed how the Pokemon interacted. Showing some evolutions here, I believe. And just going over like, the different habitats where different Pokemon, I guess, took over or made as like their main territories. So this was a really fun one. Tyler, I wonder if... When you do these sorts of more ambitious projects that are running a long time in the background, do you ever ask it like, how it did that, right? Like, did you get more details on... Because when I look at this, I think, wow, like what... You know, did it use like a game engine? Or did it like, make its own thing up? Yeah, I have all these questions. But I'm wondering if you explored that. At the risk of sounding like a total noob? No, I didn't. I just said... You know, I said it... I said this on UltraCode. I gave it a goal and said, I want to create a simulation of a world where Pokemon actually existed alongside humans. And that was it. I think it was a very short prompt. And this one ran for, I think, a day and a half was how long it ran. And I felt like my computer was gonna... My Mac Studio was actually gonna melt at times where I got really nervous. I actually took some things out of the freezer and placed it on top of the Mac. I live in Hawaii. So what's going to go wrong? It can be like 80 degrees. I know, I was like, you know, watching for condensation, but I was like, man, I was like, I gotta cool this down a little bit. And my Mac Studio had never really got that hot before where I could hear it just pumping. So it went to work, man. It went to work. So you use all these fun things. Did you use it for actual work, like design-wise? Because you do design work. Like, how do you feel it is for design work as a designer? Not like me, not being a designer. Like, is this the model or so like Opus, Soul, Fable? Which model, Astra, which model do you use if you do design work now or going forward? You know, I still use Grok for, right now, Grok 47 for real work at every just because it's so cost efficient. And I feel like I'm able to steer it well enough so I can kind of save, you know, a little bit of our, my usage limits for the entire month. But for, if I think if tokens were free, there would be no doubt I'd be using this model. And I can even show some of, some examples of- Which model? Which is this model? Opus 5.5. Opus 5.5, 100%. I'm only going to be talking about Opus 5.5. Huge, huge fan. But let me show you just a few things that, here's like a few recordings I saved of the, I have a component generator here, or this is like the old Opus. And here's like some of, you can see the improvements in the quality and design of these things where it just got better. Here's just something fun I made where it's like a slot machine that generates UI components for me. So just like really, really fun stuff. And this, this was also one prompt. This was a more complex prompt. I don't quite remember exactly what it said. But, you know, again, that was one prompt. And I think, also too, it built this design tool that I need to share with you here. I'm going to actually, give me a sec to import a, for the design. So I'm going to import, for those of you that have ever heard of Origami Studio. It's a prototyping design tool. Very hard to use. It uses something called patches. And it's, each patch is almost like a functioning code. And I'm just going to show you how it works here. But this is an app that I'm just working on in a personal time. My personal time is kind of like a. So this is not Origami Studio. This is your own version. This is my own. This is my own. This is my own. You started with Opus 5.5. Yep. Just a little bit ago. Maybe about four prompts for this whole thing. Like no joke. The first prompt was like build an Origami Studio that was AI friendly. And I think instead of explaining how it works, it's a very complex tool. I'm actually using cloud code to control it in different terminals. So I'm going to say. Add a swiping interaction. Agent A. What you're going to see. Yeah. What you're going to see here are patches on the bottom. These patches are kind of like functions in code. And, and these are actual like real design layers. So you can see like the design layers on the left here. You can see the inspector panel on the right here. I also had built in like a assistant. So I can put in my own anthropic key if I want to just chat in there. And right now it's working in the background. And in the moment we should see a bunch of patches appear in the bottom. And I'll be able to show you exactly how this works. And I'm hoping and plan to use this for my actual work. Because what this, what you'll see is that this allows me to. Fine tune the interactions. In a way that is node based and I can kind of see everything on a canvas. Okay, so here we go. So you can see. And let me just show only the nodes here. And show the prototype. So it's a little messy, but it created all these nodes here. Let me see if I can tidy it up. And this is, it just built this tidy up feature, which is in origami. Oh, I guess that's tidied up. So let me show what happens. Let's see if I drag. So you can see if I'm dragging it. You can see all of these numbers changing here. And I can actually adjust all of these things. So in here somewhere there's probably going to be a scale animation. I can see that there's an opacity animation on this yes badge, which I actually don't like. So I'm going to just stable that. And disable this one too. And let's see if it should go away. So you can see that heart disappeared here. So this actually is just a way for me to like really fine tune and play with, and just have a lot of fun with the things like the spring here. And the bounciness. I can like max these out and make it really crazy. You see how much quicker that went? Maybe I'll make it slower. I'm not sure if I could see the difference there. It happened so quickly. But yeah, this is a tool that I'm like really, really excited about. And it's just, it was so complete. So like if I start a new file here, it came with like tutorials. It came with templates to start with. So like I can, it's just really, I don't know. I'm just really blown away. And then I can also like find it. So do you think you'll actually use this? Because like with all these tools, it's always like, yeah, that's great to share on X. And like, it looks amazing. Like, wow, wow, wow. But like, will we use it every day? Similar with games. Will we actually like play it for more than two minutes? Is it actually worth our time? Do you think you're going to use this tool and develop it further and actually have it be part of your workflow? Because I'm like, I'm always skeptical about these things because yes, it's cool that we can create them and it's insane. But like, is it going to help our lives? Are we actually going to do stuff with it? I think it depends person to person. For me personally, I am super excited about this tool because I've been trying to learn how to use Origami Studio for a while. Like many times I've tried and given up. They don't update their documentation. And to learn it, if you get stuck, you have to go into a Facebook group and hope that someone will respond. And even when they do respond, oftentimes they will tell you how to do it, but it doesn't really quite sink in because you don't understand just the complexities of the tool and maybe things have changed. They have a, you know, they just have a tutorial you can go through and that not even the tutorials up to date. So it's just something that I really want to learn because I think something that I noticed is that many of my favorite designers, the ones that like I really admire, are very good at using this tool. So for me, I am building this tool for me so that I can actually learn how to use Origami on my own, but also to add, to use AI with it. So I can tell it like, hey, add this, you know, swiping interaction or when I click on this, it should, you know, it should, it should scale up to the whole screen. And I can learn about what these things are over here. So this is what I added here on the right side here. It's like, it's a way for me to learn what these things are. So I'm hoping that this is a tool that actually makes me a better designer. And who knows, maybe, maybe it's something that I do use long term. Maybe not, I added some features where I can like actually vibe design in here. I'm showing that kind of like holographic animation. So maybe it's, you know, maybe it's maybe it becomes my design tool. And I'm not in Figma all the time. But no, I don't see Figma going away from me anytime soon. But it's how long did you work on this? How long does, oh, did Opus work on this again that someone asked? I think total on and off for about three days. And it's that's not three days, like nonstop, because I was working. So I would, the first prompt was pretty long, really describing what the goal was, what Origami Studio is, and how I wanted it to be able to be manipulated by cloud code. I think the second prompt was like, okay, I need to actually import real design from like production code. So like that app you saw was something I vibe coded with Astra, actually, and then imported that Swift UI app into into this canvas. Then the third and fourth prompt was, I want to be able to design that new designs kind of like how paper does it if you're familiar. And then I think the last one was I want to add loading states. So I created that holographic animation. That one took a that one actually had to design a little bit of the, had to provide some references and really describe it, but it did that holographic thing in one shot. And also, there's some loading states for the nodes that didn't show. So that's a live demo bug there. But I was staying on and off three days. Great. You dropped another name. You said Grok 4.7 as well, which also dropped yesterday. So there are actually three models. And Mike, I know you ran some benchmark on it as well. I did too. So maybe we can just briefly talk about how Grok 4.7 compares to Opus 5.5 and HGT Sol, Luna, and yeah, and then after that, maybe we can finish the live stream and get back to using it. Yeah. Let me show you some stuff. So I was really, I think we talked about this Tyler, but I was really surprised by Grok 4.7 because Grok 4.6 is a great model and like surprisingly good, better than I expected it to be. Right. And when we started testing 4.7, it feels like it was rushed out. And actually, at least on my benchmark, it was a real regression. And I don't know if that is related to maybe there were just a bunch of bugs or things for them to figure out. I know you said that you saw it getting better over the previous week. Like what was your experience with it, Tyler? And I'll show you some of the examples I've seen. Ooh, the first, the first few days were rough. It was definitely rough. And I think that, you know, just maybe they weren't ready for, for us to, to test it, but it did, by the end of the week, it did become a lot more reliable. But there were some things early on where it was, you know, giving up a little bit or it was hallucinating. There was a lot of, I think, issues with probably their servers at the time. Like they were probably like working on it really, really hard while they gave us... There was some, there was downtime for cursor last week, I think, one day. Yeah. Yeah. So there were some issues around there. But you said you use it and you like it, right? So how do you use it today? Well, I would say that it, I'm still figuring out how I love to use it. And I think this is the first time that for whatever reason I've moved back to the cursor IDE over the agents window. It just seems to work a little faster. Maybe it's just my imagination. But it seemed to work a little faster in the IDE, especially during testing. And I've kind of just decided to stay there and work locally as opposed to working in cursor projects, which is like something I really, really love. But I just, I'm just not having the same, I guess, success that I was having with Grok 4.5 and especially Grok 4.6 inside of cursor projects, which is my preferred way, is one of the reasons why I use cursor so much is because cursor projects allows me to have one chat for whatever I'm working on. And it will create all these sub agents in one place and I can have another project underneath it. So it really allows me to multitask well, but it's still to be determined if Grok 4.7 is, I'm really good at orchestrating those projects to other sub agents. But my previous setup, I really enjoyed having Fable 5.1 low or Fable 5.1 medium as the orchestrator and having Grok 4.6 as all of the sub agents. It just worked really, really well. It worked so well. So I'm hoping, I don't know, I'm hoping for Grok 4.7. And I've been using it today and yesterday. I've had no problems. Like in the testing, it was a little rough early. There were some interesting moments. It was DMing, you know, one of my PMs without telling me. And I realized later, I was like, oh my, I was like, Marcus, I'm so sorry. I was like, Marcus, I'm so sorry. And I think it took some things I said super literally. I think, for example, I will say something like, you know, let's get this ready to merge to main. And then it just merged it to main and DMed Marcus. And I was like, oh no. Like, big no-no at every, right? We always, humans always do the squash and merge here. But yeah, like I said, as the week had gone on, it had improved. I actually am starting to like it a little more than 4.6 actually. As of this morning, before I came on, I had worked on, I had shipped the PR and had no problems. It does feel like an improvement, but it's one of those improvements where like Grok, Grok 4.7 isn't like a one-shot show horse type of thing. So it's hard for me to evaluate it. It really, really is. And when I use it for real work, like 100% of my real work is almost using Grok. So it's really hard for me to say like, oh, is it noticeably better? No, but it seems to be working great. And I don't plan on switching, but man, I do wish, you know, I do have some withdrawals for Opus. What did you find, Mike, in your lesson? Yeah, so I feel like maybe we need to retest if it is changing that much, right? But I found some of the tasks that did pretty well. So this was the MPS dashboard. This is a different design than what I've seen from other models. So like if you look at say like what, what Fable is. So we're looking at Grok. Yeah, so this is, this is Fable. And then this is Grok 4.6. And I really like the headline here and the design. So I was, I was really happy with 4.6. Just in comparison, this is, you know, this is Sol 5.6. And, you know, like the, let me find one that didn't do as good a job. Maybe like Opus 5. Yeah, like this, this vibe coded kind of like, you know, dark mode thing. But, but yeah, Grok 4.7 did, did okay on this one. But just to kind of give you a sense of my issues with it. It likes a lot of the tasks that just completely didn't finish. Right. So like it didn't end up outputting the final result. So there was some like straight up errors there. And by the way, like I, I reran these tasks like three or four times. And then some of the ones where it did finish, which makes me think it's not just like a harnessing. Like, let me take a look at the PowerPoint task. Let me see. So was it this one? Yeah, it was kind of, I'm trying to find the example specifically, but. I, I, I, yeah, this isn't the one I was thinking of here. I think I might've overwritten it. But yeah, I found it was, it was like a little bit messy. Let me just see the write up example. Yeah. So like, here's, here's a good example. Like the, the writing here. It was just like, it felt like a couple of steps backwards. Relative to the previous model. Like it's, it's just kind of. It kind of read as like, I guess, like a to do list almost rather than an article. So that was one issue. And then a lot of the stuff, it just, it just didn't like, just didn't like do that good of a job. Like the design wasn't as good and stuff. So it just felt like a little bit of a step backwards. And because it was unreliable, it became something that we were like less, less interested in. But, but I hope that like, I have like really high hopes. I hope, I hope that, you know, the reliability issues improve. Yeah. So for me, I did just some A-B tests here. Left is Grok 4.7, right is Opus 5.5. I mean, it's not as great at 3D here. And on the left, Grok 4.7, this is the video. So, and the video is really trying to show, can it do like, very good planning, can it like, iterate, stuff like that. So left is Grok 4.7, right is GPT-6 Luna. So, I mean, if you look at the two, it feels like the 3D action of Grok 4.7 is more in the Luna realm. I would say Luna looks still a little bit better. So like, from a 3D perspective, it's not very impressive. From coding, yeah, in LFG didn't do either very well. And maybe that's because I'm using workflows. So clearly, I'm using this in Compound Engineering as a workflow, which is like, it needs to follow all the steps. And if it doesn't follow all the steps, it's harder to make it work. But I did like, or I do like 4.6 as just a workhorse. But also, it gets very hard because Luna has a workhorse, Sol is a workhorse, Opus 5.5 is a workhorse. There are so many workhorses and suddenly workhorses get so good that they're better than Able, which is just insane. And just having this entire landscape like change so much, like Sol 6 is way better than 5.6. But there's Opus 5.5 as well. And it just changes everything. If you're curious, check out Opus 5.5, but also check out all the other models. I think for writing, especially Dan, who really liked Sol, check out Opus 5.5 on low effort as well. I think that's something to try out because if you do high, it can feel a little bit slow, but you don't really need slow. You can just put it on low. And yeah, like any other tips or tricks to take away before we hang up Tyler, Mike? I think the main thing just to think about is the pace that these new models are coming out in and like how hard it is to test them and see the differences between them. We're kind of getting to the point where we're largely saturated on a lot of tasks. If you can think it, you can do it. And it's starting to come down more to personal preference, which is like, I don't know about you, Kieran, Tyler, but I've started to care less about is this the best model in the world, advanced math or whatever. And instead, like, do I like the designs it comes up with or do I like the writing? And I've become more comfortable with being subjective. And I use all models. Like I literally use all models. It's not that I don't use one over the other. But like if I make 3D games, I probably won't use Grok 4.7. I will probably use Opus 5.5. But I also like I love Cursor. Cursor is my app. Like I use all my tokens in Cursor because Cursor Projects is amazing. And Codex is also amazing. And Claude is catching up like the app. Like it's not as amazing as the others. But I do use all the apps, all the things. And yeah, if you need to choose, like hopefully you learn something. You don't need to switch because they're all very good. But yeah, Opus 5.5 is a special one. I really think it deserves like a special mention. Like or like if there's a winner, for me it's Opus this week in launches for sure. Yeah, definitely. And I feel like I just broke up with Opus too. So I'm like all the way back. It's really, really funny how quickly I came back to Opus. But I agree with Kieran. You know, I use all the models. I don't see myself going away from Cursor and Grok for actual work. I'm not going to get rid of my Grok bots. Love them too much. For anything in Blender, I actually really love using Astra with Computer Use Control Blender as opposed to MCP. I can't get rid of that. And this Opus thing is just, man, it's just magical. Hurts my wallet. But I can't, I really just can't put into words like how awesome Opus 5.5 is. Sounds like a good place to end. Great. Thank you, everyone. Every is the only subscription you need to stay at the edge of AI. We said it. Oh, yeah. Go subscribe. We forgot to say they've ramped the line. Yeah, we're such horrible presenters. Every.to. Yeah, go check out our Vibe Checks. Yeah, our Vibe Checks will arrive today. Either they're live already or they will arrive. We're doing comparisons there as well. We'll be sharing more on X and LinkedIn and other places. But thanks for hanging with us and see you next time. Happy Release Week, everyone.