Every

LIVE VIBE CHECK: OPUS 4.7 DROPS

2211 summary words 10 min summary Watch video

Start with the signal

10 min read

Summary

At-a-Glance

  • Verdict: Skim
  • Core thesis: Opus 4.7 appears to be a more literal, controllable, and potentially stronger long-horizon agent model than Opus 4.6, but it requires more explicit prompting, appropriate effort settings, and further testing before it can be called a default replacement.
  • Why it matters: The model's behavior shifts the operating model for agents: vague prompts that 4.6 could infer may underperform, while explicit scope, autonomy, and effort instructions can unlock more reliable analysis, document work, visual implementation, and asynchronous execution.
  • Best use: Watch or scan the Anthropic researcher segment and the practical tests for prompting guidance; use the rest only as preliminary signal rather than a definitive benchmark.

Executive Summary

This is a two-hour live, highly provisional hands-on test of Anthropic's newly released Opus 4.7 by Every. The hosts compare it with Opus 4.6 on financial analysis, investor-update writing, an OpenClaw bootstrap, Slack research mining, simple app generation, and a deliberately messy production-code refactor. Their central early observation is behavioral rather than purely capability-based: 4.7 follows literal instructions more closely, is more direct in prose, and fills in fewer unstated assumptions than 4.6.

That behavioral change produced mixed results. Opus 4.7 generated an accurate, usable investor update from a P&L and showed strong architectural diagnosis on a collaborative-editor codebase, correctly identifying the need for a single authoritative document state. But it was initially less useful on vague, exploratory work: it produced less distinctive writing, missed a previously found P&L data-quality issue, organized an OpenClaw setup less appropriately, and proposed a less functional AI todo feature than 4.6. In the code-refactor test, it understood the “burn-the-ships” plan but initially implemented safer peripheral slices rather than carrying out the core rewrite.

Anthropic researcher Alex Albert contextualizes those results. He says 4.7 is intentionally more “on the mark”: it will do what it is explicitly asked to do, rather than making the broad, sometimes unwanted changes associated with earlier Claude behavior. He recommends explicitly asking it to be thorough, make many tool calls and file changes when desired, and use higher effort levels for agentic tasks. He reports particular gains in long-running background work, high-resolution visual inspection and screenshot-to-UI replication, and spreadsheet/document/PowerPoint tasks.

The practical conclusion is not that 4.7 is weaker, but that prompt and workflow design must change. Treat effort level, available time, completion criteria, permission to make broad changes, and style/context bundles as first-class controls. The video is valuable for that operational framing and the candid examples, but it is repetitive, promotional, and based on only a few early tests.

Key Takeaways

  • Claim: Opus 4.7 is more literal and instruction-bound than Opus 4.6, so it is less likely to infer unstated intent or act beyond the prompt. | Evidence: The hosts found that “execute end to end” did not reliably mean “execute the previously written refactor plan,” and writing results were more systematic and direct than 4.6. Anthropic's Alex Albert confirmed that 4.7 is designed to be more “on the mark” and needs users to explicitly restore desired thoroughness or scope. | Implication: Update reusable agent prompts to explicitly state desired scope, autonomy, success criteria, permission to change files, and whether the model should infer missing details rather than relying on conversational implication. | Caveat: The tests occurred live on launch day, before the hosts had adapted their prompts; some apparent regression may be prompt-model mismatch rather than lower underlying capability.
  • Claim: Effort settings and explicit time/autonomy framing are major levers for getting stronger agentic performance from Opus 4.7. | Evidence: Albert says Anthropic improved the relationship between effort levels and token usage: Max is intended for unrestricted token/tool usage, Extra High is Claude Code's default for longer asynchronous work, and High/Medium are balance points. His example was telling the model “I’m going to bed” while building an apartment-search dashboard; it continued iterating rather than stopping after a first pass. | Implication: For autonomous agents, set High through Max effort deliberately and state that the task can run asynchronously, should iterate, and should not stop at the first acceptable slice. Reserve lower settings for chat or lightweight non-agentic tasks. | Caveat: Albert does not provide a quantified horizon increase versus 4.6, and higher effort implies greater latency and token consumption.
  • Claim: Opus 4.7 showed strong quantitative-analysis-plus-business-writing performance when the task was concrete. | Evidence: Given Every's March P&L and a request to produce an investor update, it reportedly got the figures right, broke the update out by product, surfaced a direct narrative, and produced output close to what the host had sent. The host felt it had more solidity and reliability on this blended task than recent Opus versions. | Implication: Use 4.7 for structured reporting from financial or operational source material, but explicitly require transaction-level investigation, anomaly detection, reconciliation, and root-cause analysis instead of assuming it will go beyond variance summaries. | Caveat: On a deeper version of the same P&L analysis, it reportedly required repeated prompting and did not independently find a failed-transaction export issue that Opus 4.6 had found.
  • Claim: For large codebase refactors, 4.7 can identify senior-level architectural invariants but may default to cautious incremental implementation. | Evidence: In Every's “vibe slop benchmark,” the model correctly identified that the collaborative editor lacked a single authoritative document owner/state and noticed an oversized collab file that other tested models had overlooked. Yet after proposing a first-principles rewrite, it initially implemented a safer first slice rather than cutting over the riskiest core architecture. | Implication: For de-slopping or replatforming codebases, make the migration posture explicit: identify the central invariant, prioritize the riskiest cutover, prohibit superficial wrapper fixes, and require the agent to report what legacy paths remain before claiming completion. | Caveat: This is one bespoke benchmark on an existing codebase, evaluated during a live stream; it is not a general coding benchmark.
  • Claim: Opus 4.7 may be better suited to professional, well-specified engineering workflows than zero-context vibe coding. | Evidence: The hosts' rough-prompt tests favored 4.6 in several cases: 4.7 created an OpenClaw structure less aligned with the expected file conventions; in a todo app, its AI feature sharpened text within one task rather than splitting it into usable subtasks; and it needed more prompting to deepen financial analysis. The hosts characterized it as potentially a power tool “if you prompt it right.” | Implication: Do not assume 4.7 will be the best default for low-specification prototyping. Use it where requirements, architecture, and desired behavior can be codified; retain comparative routing for exploratory generation until task-specific evaluations are complete. | Caveat: The model still generated functioning apps and its initial plan quality in the refactor task was considered strong.
  • Claim: Its enhanced vision capability is positioned as practically useful for screenshot-to-frontend replication and fine visual QA, not merely image understanding. | Evidence: Albert says 4.7 supports larger, higher-resolution images and performs better on dense-detail “Where’s Waldo” style tests. Anthropic uses this to evaluate whether it catches subtle misalignment or missing visual details and can reproduce a frontend from a screenshot. | Implication: Add a targeted evaluation for design-to-code and visual regression workflows: provide high-resolution references, ask it to enumerate visual mismatches, then have it implement and re-check corrections.
  • Claim: Writing quality is mixed: 4.7 can be strong for direct business prose and research synthesis, but its default creative voice was judged more rigid than 4.6. | Evidence: A side-by-side article-introduction test found 4.6 more unpredictable and natural in cadence, while 4.7 read as more systematic. In contrast, 4.7 successfully mined recent Slack discussions for editorial “nuggets,” selected plausible items, drafted a usable synthesis, and improved when given a curated voice package containing examples from writers such as Jia Tolentino and John Green. | Implication: Use 4.7 for research harvesting, structured synthesis, and direct stakeholder communication. For distinctive editorial writing, provide an explicit style corpus and anti-pattern checks rather than relying on a short voice instruction. | Caveat: The writing system already included skills and style guides, and the hosts did not complete a blind raw writing test.

Detailed Brief

Prompting and workflow controls to test immediately

  • Claims: The host team believes old 4.6 prompts may produce worse results on 4.7 because they depend on implicit intent being inferred.; Albert frames the shift as part of a broader move from collaborative turn-by-turn work toward asynchronous task ownership.; The model can sometimes push back on a poorly framed request rather than merely executing it, according to both the Anthropic guest and a writing-task example.
  • Evidence: Albert recommends instructions such as being thorough, not minding many tool calls, and not minding many file changes where broad work is desired.; His background-agent example specified an overnight deadline, autonomous execution, a complete dashboard, defined apartment criteria, data sources including Craigslist and Zillow, and a twice-daily refresh schedule.; On a content-classification task, the model reportedly said the raw materials did not fit the requested module and identified missing elements.
  • Caveats: The claimed response to “I’m going to bed” should be treated as a practical prompting heuristic, not evidence that the model has reliable real-world temporal awareness.; The transcript does not establish whether effort levels are exposed consistently across every Claude interface or API configuration.
  • Implications: Agent-product UX should expose autonomy, time budget, effort level, allowed scope, and a completion threshold rather than presenting one generic “run” control.; Task specs should distinguish between “make a safe local edit” and “rebuild the system around a new invariant,” since 4.7 appears responsive to that distinction.

Early comparative signals against Opus 4.6

  • Claims: 4.6 remained the writers' immediate default for loosely specified, voice-heavy drafting.; 4.7 was considered more minimalist visually in a quick todo-app comparison, while 4.6's AI breakdown of a broad task into separate subtasks was more usable.; Both models could generate a functioning simple app without troubleshooting in this test.
  • Evidence: Both todo apps added a completion animation and an AI feature after sparse prompts; 4.7 kept the expanded plan as text in one todo, while 4.6 created multiple actionable todo items.; In an OpenClaw bootstrap task, 4.6 created separate user and soul files, memory seeds, cron jobs, and skills; 4.7 consolidated context into an agent file and created some skills the user did not recognize as relevant.
  • Caveats: These comparisons were informal, sparse-prompt tests and should not be interpreted as comprehensive model ranking.; The hosts repeatedly note that 4.7 had just launched and the product interfaces themselves were somewhat flaky during the livestream.
  • Implications: Maintain task-level model evaluations rather than replacing 4.6 wholesale based on benchmarks or launch claims.; For OpenClaw or agent bootstrap systems, validate generated file conventions, secrets handling, memory format, and service configuration rather than trusting an apparently complete scaffold.

Notable Concepts & Terms

  • Vibe slop benchmark: Every's bespoke test: give an agent a historically vibe-coded, unreliable production codebase and assess whether it can infer the senior-engineering architecture and actually execute the rewrite.
  • Burn-the-ships rewrite: A deliberate full architectural cutover rather than adding safer layers around legacy code; central to testing whether a coding agent can resolve root causes.
  • Single authoritative state / single-writer invariant: The key design principle 4.7 identified for the collaborative editor: one source of truth must own document state rather than multiple competing authorities and merge paths.
  • Effort levels: Controls intended to regulate 4.7's token use, tool calls, and persistence; Albert positions Max and Extra High for intensive agentic work, with High/Medium for more balanced tasks.
  • Asynchronous task ownership: The operating model in which an agent is given broad instructions, adequate time, and authority to iterate independently instead of working only in a human-led conversational loop.
  • OpenClaw: The hosts' agent framework/context package, used here as a test of whether the model can infer a user's files, memories, skills, cron jobs, and integrations from existing Claude context.
  • AI smell / correlative constructions: The writers' shorthand for recognizable model-generated prose patterns, such as repetitive “not X but Y” structures, used as a qualitative test of writing naturalness.

Operator Notes / Why Ken Should Care

  • Run an internal 4.6-versus-4.7 prompt migration test on Ken's top agent workflows; record whether each existing prompt relies on implicit intent, then add explicit scope, completion, and escalation instructions.
  • For long-running agent jobs, test High, Extra High, and Max effort against fixed acceptance criteria, measuring completion rate, cost, wall-clock time, tool-call count, and unnecessary-change rate.
  • Create a refactor prompt template for legacy/vibe-coded systems that requires: target invariant, authority boundaries, explicit permission for broad changes, a prioritized risky cutover, legacy deletion criteria, and post-change verification.
  • Keep a separate model route for high-voice writing until blind tests establish whether 4.7 plus a style corpus outperforms the current default.
  • Add transaction-level anomaly and data-quality checks to any financial-analysis workflow; do not accept a month-over-month summary as a completed analysis.
  • Evaluate 4.7 on screenshot-to-implementation and visual QA if frontend replication or design-system compliance is a current bottleneck.

Source/Metadata

  • Title: LIVE VIBE CHECK: OPUS 4.7 DROPS
  • Transcript words: 25938
  • Duration seconds: 6831
  • Timestamp note: No usable timestamps or chapters were present in the supplied transcript; the transcript also contains substantial repeated passages from the livestream.
Full transcript 18564 words · 114 min read
0:00

great, we're working. We are live. We're doing an Opus 4.7 vibe check. The mic is on. I'm here with Brandon, our COO. We have matching shirts for today. You always got to match. You always got to wear green on a vibe check day. We've got our water here. We're getting hydrated.

0:12

Now, as I was saying prior to the mic mishap, there's a new model launch. Yep, there's a new model. We've got a tweet. We normally get access to these models before they come out. This time, we were snubbed. Anthropic snubbed us. We did not get access to this model. I don't know why. I have started texting them, being like, you hate me. Are we breaking up? It's possible that it was just a misunderstanding. We'll see. But that means that this vibe check is going to be live.

0:19

So what we're going to do is go through what Claude says about Opus 4.7, and then we're going to be vibe checking it live. You're going to see us go through our standard battery of tests to see how it looks.

0:29

All right, so what Anthropic says is Opus 4.7 is our most capable Opus model yet. Everyone always says that. It handles long-running tasks with more rigor. Everyone always says that. It follows instructions more precisely. Everyone always says that. And verifies its own output before reporting back. That's actually new. I've never heard that before, and that feels really good. It's one of those moments where a thing that good prompters know how to do seems to have started making its way into the models. It's generally a good practice to ask your model to reflect on whether it did this work properly before you move on. It seems like the models are now doing that automatically. That's amazing.

0:37

Versus Opus 4.6, it looks like about a 10 percent bump on SWE-bench Pro, about a 7 percent bump on SWE-bench Verified. It's a pretty good bump. It's a pretty good bump. It's not as good as Mythos, so your cryptography is safe. We're still safe for now. Yeah, so let's get into it. I'm going to open up my Claude, and what I'm also going to do is... I'm not seeing anybody on Twitter right now talking about having an experience so far with it. They may have been rushing out because there are a lot of models yet to be dropped. Yeah, Opus 4.7 is likely launching today. If this really reflects their internal progress, Opus 4.7 may have been released.

0:53

Yeah, yeah. All right, that's what it is. All right, I think everybody maybe started calling them on Mythos and was like, is this real? That's possible. We need something. That is totally possible. All right, let me get myself set up here. We're going to do some work in... let's see. We're going to do some work in Proof. Brandon, do you want to talk a little bit about what you're seeing while I get... While I get...

1:36

Yeah, so I am right now asking... let's get this mic... this is the mic right now. I am right now asking Co-Work to spin up an open Claude for me, and the reason I'm doing this is I have an example of this working with Opus 4.6. The reason I've done it like this is Claude knows a ton about me and what I use is Claude 4.4, so I wanted to make a very good open Claude using all that context. Basically, I have it for Opus 4.6. I'm now asking Opus 4.7 to do the same thing. I dropped in the same exact prompt, and we're going to see what it comes up with. I can tell you right now, it's a lot slower. It seems to be a lot slower for me.

1:42

I'm going to share my screen. The first thing I want to do is something that I did recently, so we'll be able to compare it. So one thing that I did recently with Co-Work, we're going to try it in Co-Work first, and then we're going to move into some harder programming tasks in a second. But one of the things I've been doing in Co-Work is I use Co-Work to do some financial analysis. So what I wanted to do is, with Opus 4.7, I want to see how good it is at taking our P&L from Every, which is a fairly complex spreadsheet, and turning it into an investor update for March, because that's the last month that we've closed our books. So I'm going to say, I attach our P&L. Based on what's in here and what you know about Every, can you please write our March 2026 investor update? This is something that I've already done, so I'll be able to compare this pretty well to what we've already done.

1:50

Oh, we've got Katie. Katie Parrott, how you doing? Hey, I'm doing pretty good. Happy Opus 4.7 release day. Happy day. Have you started testing it out yet? I am just now opening up my environment and switching the model over to make sure that I'm on the latest and greatest. So you keep doing what you're doing, and I will get myself set up for some writing testing shortly. Yeah. What are you excited to test with this model?

2:28

I'm excited to test the writing style, how much AI smell there is on the copy, which I think is always the first thing that we're looking at with any new model from a writing perspective. And then I'm also curious how the leap and agenticness will work, and powering some of the AI workflows that we have, the skills that we're developing for our new content modules, all that good stuff. So basically, just excited to see how it works inside the work I already do.

2:37

Great. Actually, that reminds me. So right now Katie's going to test writing. Brandon's doing some other stuff. She's going to test it with... he's going to test it with open Claude. Right now, I'm testing to see if they can write our investor update, which is a combination of writing and financial analysis. I've also been writing this manifesto for Every in Proof, which is our agent-native document editor. What I want to do is... I've created this little workflow where I've been using Codex for this, where Codex will just read what I'm writing and then just give me ideas in Proof in a loop. So what I want to do is see if Co-Work's any good at this.

2:43

So we've got a little workflow. Let's see. Two R2.

2:48

So basically, what I'm going to ask it to do, this is kind of cool, is I have asked it... I've created this little workflow, which is essentially like give your agent a little bit of a scratch pad inside of whatever you're writing. So this is a Codex scratchpad workflow. Give your agent a scratch pad inside of whatever you're writing, and then have it connect to that document, and then as you're writing, just give you ideas. So that every couple seconds, you get new ideas, and then you're going to do this based on whatever you're writing about. I've been doing this with Codex. I was doing that this morning, so now I'm going to see how Co-Work does with this. So let's watch.

2:52

And if you are watching, if you're here and you are using Opus 4.7 for interesting stuff, please share your vibe checks in the comments. For any use case, give us a red, yellow, green, or gold. Red is terrible trash model. Yellow is okay. Green is like, I probably use this over the other options I was using it for. And gold is paradigm shift. Katie, does anybody have access to this in Claude Code? Let me see. I can't get it running in there, which is mostly where I work. Let's see.

3:17

I see Opus 4... Yep, I don't have it in Claude Code yet. It's not in Claude Code yet. Does anyone have it in Claude Code? Yeah, people are not seeing it yet in Claude Code. It looks like it's in Co-Work first. Let me just see if I can yell at Anthropic about this. And I just heard from Anthropic. They did not run an early access program for this. Okay, now let's see how we're doing. This is a weird response. It looks like it's a little bit confused. Let's see. Hey, can you just figure it out and just run in a loop to give me ideas, please? And you can rename the Codex scratchpad to Opus scratchpad. Thanks.

4:17

Okay, now let's look at our investor update. Okay, this is good. It's asking me questions. Web Proof doc, please. I just want to make sure, am I actually still screen sharing? Yeah, I am. Cool. Standard. These are good questions. Oh, this is good. I would say that... Oh, and okay, good. Next. Celebrate it. Okay. Let's say celebrate it, as it was a good month. Okay, so now Opus 4.7 Co-Work is back off. And sick. just see if I can yell at Anthropic about this.

5:22

And I just heard from Anthropic they did not run an early access program for this. Okay, now let's see how we're doing. This is a weird response. It looks like it's a little bit confused. Let's see. Hey, can you just figure it out and just run in a loop to give me ideas, please? And you can rename the codex scratchpad to opus scratchpad. Thanks. Okay, now let's look at our investor update. Okay, this is good. It's asking me questions. Web proof doc, please. And I just want to make sure, am I actually still screen sharing? Yeah, I am. Cool. Standard, these are good questions. Oh, this is good. I would say that. Oh, and okay, good, next.

6:35

Celebrate it. Okay. Let's say celebrate it, as it was a good month. Okay, so now codex is, sorry, opus 4.7 co-work is back off. And and sick. And you can see we've got opus is in here, so that's good. And let's see if it figures out how to modify this document according to this. I wonder who's in here with me. Hope it's no one nefarious. Okay. So it's doing some analysis. That's good. Have you updated and relaunched your cloud desktop app? Yeah, okay, yeah. Hello, Katie. Welcome back. Hello. I xed out of the group Chrome browser that had StreamYard on it, like

7:47

a genius. So it goes, so it goes. Okay, have you found anything worth reporting yet? One thing that I'm interested in is that I'm in Claude Code, which is where I run my writing plugin that takes my process end to end when I'm writing an article, and it says read memory m.md, which I have not seen before from Claude. So is that something new with this model, or with just the cloud code update, which I just did? Is it remembering stuff now?

8:14

Because my understanding has always been that inside cloud code work is super siloed. I believe they've had memories for a bit. I would guess that that's not specific to this model. All right, I have got something interesting here, so let me share my screen. Great, share your screen. I'm going to take this down, one sec. Oh, okay, before you do that, one of the things, one of the tests we've been running is how good is opus at giving me ideas for writing. So I've got it connected to proof. This is something I'm working on, and you can see it in here. Nope, you cannot see it in here, so that's interesting. So it's not good at presence.

9:26

Interesting. Well, it is not good at filling out its presence so far. What it has managed to do is give me some ideas for what to do next, and it has scheduled a task to do this, so we'll see if this is actually good. I'm interested in this. Okay, that move. These are good. It's not immediately like a holy moment, but I'm intrigued. Oh, I need to update my desktop. I'm going to do that. Do you want to share your screen? Yeah.

10:17

Let's see. Okay, so I haven't looked at any of this yet. We are—explain just what we're doing fresh. So what I did was I went to co-work. Actually, let's pull up the prompt. I went to co-work and I gave this prompt to both 4.6 and 4.7. I would not recommend you run this prompt. I think we can make it much better, but I've done this before and it's worked great, so I know that it works. Essentially what I'm asking co-work to do is to make me an open claw that's based off of everything that I do inside of quad. So it has my memories, it knows the work I do. I'm not telling it which files

11:00

to make. It should know that. I'm not telling it really anything, just that at the end of this prompt I want to have a functioning open claw. So I did that for 4.6 and 4.7, with 4.6 on the left, 4.7 on the right. It went about doing that pretty differently. So 4.6 made a sole file, a user file, which, that's what open claw needs. It created memory seeds, which is interesting because it would need to know what to do with. Open claw would need to have an onboarding to know what to do with that, which is probably not my preference. I just want to give this to the claw and for it to replace the file.

11:28

I made cron jobs for work that I do regularly using Claude, and then it made a bunch of skills of stuff that I do too. I want to run my financial finance analysis on this as well. Do it. I can't because I need claw code. I see. Okay, so that's 4.6. On the right it took a different approach here. It made skills, which is great.

12:02

Actually, this is not great. So on 4.6 we have things that I use claw for under skills, and on the right are skills that I've already made. So I made the bootstrap CFO. I don't really even know what this is. Design slides, design principles, I don't know what this is. And and then it made a top-level—oh no, that's under design slides—and then ai check. Let's see what that is. Yeah, it made a skill that's specific for finding linguistic patterns that signal AI-generated writing, which is not something that I use claw for, so I don't really know why it made that skill.

13:08

Anyway, it made a bunch of skills. It made a bunch of knowledge. So instead of making memory seeds, which I like as a concept for starting a new open clause—like, hey, read these memories and organize them appropriately—it didn't do that here. It just made different markdown files, which is interesting, but I don't really think that's the way openclaw works. Meaning open cloud doesn't organize by category. It just organizes by day, actually. And then it created a bunch of services, so these are all the connections that I actually use. I did ask specifically for all my API keys so I can just give it to an open cloud. It

14:02

doesn't seem like either of them did that, although I wonder if it did it in a secret, hidden folder. That's not working. Okay. Okay, a couple things to know. Oh, the audio is not working. Weird. Katie, can you hear us? Yeah, I can hear you. Okay, I don't know why the audio—were you able to hear Brandon? Yep. Okay. If you can't hear us, let us know and we'll figure out how to fix it. Okay, so for people who are just joining, we're at about 3,000 live viewers right now. Let's take a step back and talk about what's going on.

15:01

Anthropic just launched opus 4.7. They did not have an early access program, so normally we get access to all these models before they come out. We did not get access to this before they came out, so we are doing a live vibe check. So you're watching us do the tests that we would normally run live in front of now 3,000 people, soon to be many more. What you should know about opus 4.7: the fact that they didn't do an early access program is really interesting. It means that they're rushing pretty fast, so things are moving really fast. They did the Claude mythos drop last week, although they didn't drop it, obviously. They're doing it this week. They're doing

15:44

opus 4.7 this week. It seems like, based on the benchmarks—and again, we don't really like benchmarks. I really like benchmarks. I'll say that. I really like getting your hands in the model and using it for the things that you do every day. So that's what we do with these vibe checks. And that gives you a good sense for, okay, for coding, really, how good is it? So far it does not appear that opus 4.7 is live in cloud code. We don't have access to it yet. If you get access to it, let us know. Second thing is it does appear to be in co-work, so we've started using it in co-work.

16:25

I think it's too early to say results, vibe check, whatever, but if you're in co-work, try it. Let us know what you think. Give us your own vibe check. We'll surface that in here. It looks like it is also available in Cursor, so if you're a Cursor person, you should open up the Cursor app and give it a shot. I actually really like Cursor for a lot of stuff, so we'll kick off a Cursor task pretty soon, and we're going to keep you updated as we go along. Oh, I have access to it in cloud code now. Yes, Katie, I have access to it in cloud code. You do? Okay, nice. Now we've got access

17:12

in cloud code, potentially. Let me see if I also have access in cloud code. All right, I'm wrapping up my vibe check on open clock. Yeah. Wait, for people who've just joined, what is your vibe check? What are you doing? So my vibe check is I asked in co-work, opus 4.6 on the left and then opus 4.7 on the right, to make me an open club based off of how I use Claude. it is also available in Cursor, so if you're a Cursor person, you should open up the Cursor app and give it a shot. I actually really like Cursor for a lot of stuff, so we're going to kick off a Cursor task pretty soon, and we're going to keep you updated as we go along.

17:43

Oh, I have access to it in cloud code now. Yes, Katie, I have access to it in cloud code. You do? Okay, nice.

18:03

Now we've got access in cloud code, potentially. Let me see if I also have access in cloud code. All right, I'm wrapping up my vibe check on Open Clock.

18:10

Wait, for people who've just joined, what is your vibe check? What are you doing? So my vibe check is I asked in co-work, Opus 4.6 on the left and then Opus 4.7 on the right, to make me an Open Club based off of how I use Claude and all the memories, all the connections I have. I really did not give it any instructions. I was like, just make an Open Clause, so I should know what that is. You go do research.

18:22

On the left, they're all referencing. It's interesting. Both are doing the same thing. They've just organized the output very differently. On the left, 4.6, it knew that it needs a user md file, a soul md file. It didn't make an agent's md file, which is interesting, but I like that it broke up soul and user, and it actually gave them a name and named them Koa, which is 4.6 named it Koa, might open on Koa. Then on the right, 4.7, it didn't make a user file, a soul file, but it did make an agent file, and it just showed everything about soul or everything about user in the agent file, which probably will work fine, but not how Open Clause is supposed to work. It's doing the work. I would say it's not as organized as 4.6, and I don't like the way that it did skills. These are skills that I've already made, and I asked it to make new skills based off of how I use Claude, but things that I haven't made skills for yet. So I probably would need to test it more, and I don't want to say anything rash, but right now I'd probably stick with 4.6 for making my Open Claw like this.

18:28

All right, so that's your initial vibe check. I will say with new models, your old prompting style sometimes needs to change, so let me make this full screen. Let's kind of— Oh, I can do that. Sorry, continue.

18:43

So it will take us a little while to figure out how to use this new model well. Just because it's different does not necessarily mean it's bad. That is not my official take on using Opus 4.7. I will also say, and people are seeing this, we are seeing Opus 4.7 in cloud code, so if you want to try it in cloud code, you can start trying it right now. I'm going to kick off a new benchmark in cloud code, so I'm going to share my screen.

18:58

For people who are just joining, normally we get access to these early and we have day-of synthesized vibe checks. This one they rushed out, so we did not get early access to it. They didn't have an early access program, and so you're watching us do it live. If you want to get the TLDR, TLDR is Opus 4.7 launched much better on a lot of benchmarks. We don't trust benchmarks. We want to put our hands in it. We will have a TLDR vibe check on every every.to, so if all you want to do is get the TLDR, subscribe and we'll send that out to you by tomorrow morning. I promise it'll be really good. It'll be very detailed. It'll tell you everything you need to know about what this is good for, for coding, for writing, for design, for strategy, for business building, all that kind of stuff.

19:04

Every.to, if you're here and you're just joining, you should know that Every is the only subscription you need to stay at the edge of AI. We do ideas, apps, and training. On the idea side, we do vibe checks like this where we tell you about when new models come out, what they're good for, what they're not good for, to keep you on the edge. We have a suite of AI apps. One of them is Proof, which you can see right here. It's an agent-native document editor. It helps you write documents in the cloud with your agent, but we have five other ones. We've got Monologue, we've got Quora, we've got Spyro, we've got Sparkle. They help you do everything from writing with AI to doing email with AI to speech-to-text with AI. We also do trainings. We do live streams and courses for subscribers. It's all bundled under one subscription. If you want to stay at the edge, it is absolutely the thing to subscribe to.

19:12

Another important PSA: please stay hydrated. It's really important to keep yourself and your agent liquid-cooled. Okay, now also, if someone from Anthropic is watching, first of all, very upset, very upset that you didn't give us early access. But if you want to make up for that, if you want to join this live stream, we would love to have you on to talk about this model and what you've seen. So just DM me on Twitter. We'll get that set up. So now I'm going to start this next vibe check on a very special benchmark, so let me just get cloud code set up here for us. Can you move into this or train a branch, please? So now I'm going to share my screen.

19:40

Okay, I have a very special benchmark here that I've been using for a little while that is really cool, and I call it the vibe slop benchmark. What this benchmark is, is it's basically testing the ability for a coding agent to act like a senior engineer. What that means is I've given it an example production code base that was purely vibe coded. This is the production code base for Proof, so this is this app. I gave it the production code base for this app, and it's basically at a point in this app where it was totally slop. It was going down all the time. It sucked because I built it without looking at a single line of code. We launched it, and it just started going down all the time. So it's a snapshot of that repo at that point. Subsequently, we hired an engineer who cleaned it up and made it really nice and rewrote it. What this benchmark tests is, if we give a new frontier model a sloppy code base, can it figure out what a senior engineer would figure out on a live production code base? I think it's really cool.

19:46

So what I basically have asked it to do is I said, execute agent prompt, and the agent prompt basically just tells the model all it says to the model is, I have a vibe-coded slop code base. Can you make and execute a plan to rewrite it from first principles? I have a good idea of what that rewrite should look like. Right now it's writing our plan, and I will say I have not actually tested it much with Opus. I've been testing other frontier models with this, so this will be interesting. I can mostly tell you how it performs against, for example, GPT 5.4. 5.4 is very good at writing a good plan for this, but it sucks at the implementation. What happens is it has a very ambitious first-principles rewrite, and then it sort of gets distracted by everything else that's going on in the code base, and it ends up putting a nice little masking layer of good code and interface on top of the slop instead of actually rewriting the app. So we'll see how good Opus is at doing this.

19:51

For now, it looks—is it still running? Yeah, I think it is. It's still running. So it's still running. This is, I think, going to be a very good benchmark for telling us how much like a senior engineer this is. I'm having some challenges right now with the Claude app. None of my skills are popping up. I have a bunch of skills, and I remember these are all—I don't know. Is anybody else having that problem right now? Let us know. Try to run a skill. If you're running skills and it's not working, let us know. Oh, we got a plan. Okay, let's see. It looks like Claude is just flaky right now. Let's see. Let's read this plan together. All right.

20:23

Yeah, okay, so part of the problem with the vibe slop—remember, if you just joined, we have a totally vibe-coded code base for a collaborative document editor called Proof. Here's Proof. It's something that I made. I made it terribly, and we're seeing if Opus 4.7 can rewrite Proof, this app. It's kind of meta. Can rewrite this app in a way that a senior engineer would rewrite it, because we have literally frozen the Proof code base at a point where it was vibe slop, and I know we had a senior engineer go and rewrite it, and I know exactly what they would do. So one of the questions that we're starting to ask is, first, can it identify the problem if we say this is vibe coded, how would you fix it? What I'm looking for is part of the things that—part of the thing that made Proof—the big thing that made Proof not work is when you have a collaborative

20:30

all right yeah, okay, so part of the problem with the vibe slop, remember, if you just joined, we have a totally vibe-coded code base for a collaborative document editor called Proof. Here's Proof. It's something that I made. I made it terribly, and we're seeing if Opus 4.2.7 can rewrite Proof, this app. It's kind of meta, can rewrite this app in a way that a senior engineer would rewrite it, because we have literally frozen the Proof code base at a point where it was vibe slop, and I know we had a senior engineer go and rewrite it, and I know exactly what they would do.

20:45

And so one of the questions that we're starting to ask is, first, can it identify the problem if we say this is vibe-coded? How would you fix it? And what I'm looking for is part of the thing that made Proof, the big thing that made Proof not work, is when you have a collaborative document editor, you have multiple versions of the same document open in different browsers, and the trick for getting that right is knowing which version of the document is the authoritative document. If I have it open and Brandon has it open and Katie has it open, whose document is the one that we say is the right one? And then how do we update everyone else's documents to match that? And the current version of the code base has 15 different authorities and all these different ways of merging them, and a good AI model will say we need one authority source. That's what a senior engineer would do.

20:50

And what I see so far, and I told Opus 4.7 I only want you, this is vibe-coded, I want you to fix it from first principles. What would you do? So I didn't give it any hints, and it says the problem is that there's no single authoritative model of who owns the current state document x right now and who may write to it, guards our local answers to a question that should be answered once globally in the type system. That is great. That is certainly on par with GBE 5.3, 5.4. What did 4.6 say about that? So I have not actually run this on 4.6, but 4.6 did it in a way that wasn't the case. I had to switch off on 4.6 while I was building this.

21:06

Okay, so this is good. This is good. This is good. Okay, this is all looking nice. Ooh, interesting. It gave me a nice rewrite. So one of the interesting things about how this was written is that the model chose to, in the vibe slop version of this code base, actually, you know what, while we're going through this plan, I'm just going to say, okay, execute end to end, great, and we're in bypass permission. Okay, so it's now coding the plan because I think the plan is good enough that it should just be able to do it, and we'll get to see what its output is, and we'll be able to watch its thinking trace.

21:11

But while that's happening, one of the things that is new about this plan that I think is really, oh, you can see Claude is in here reading its own file, that's cool. One of the things that's really good about this plan is the vibe slop that it's trying to clean up. There is a large collab folder. There's one file with 10,000 or 20,000 lines in this code base, and I can tell you that GBT 5.4 and 5.3 don't really notice that in their plans, and Opus 4.7 has noticed it and has proposed a refactor, and the refactor actually looks nice. So it says collab.ts become server cloud slash index.ts, and I realize that we're probably not going to be able to keep this video up because I don't want to pollute the benchmark, but there's some really good things going on here.

21:17

Okay, another thing is it's a little bit too cautious.

21:24

Weird. Okay, no, stop it. It did not understand what I wanted it to do. This is interesting. So I asked it, it gave me a plan, I said open it up in Proof, and then I said, okay, execute end to end, and I figured it would realize that I meant execute the plan end to end, but it didn't. So that's an interesting thing. I would imagine that Opus would know what I meant, but before I have it actually execute the plan, what I wanted to do is, okay, this is good, but it's still a little bit too cautious. Can you change the implementation strategy so it's truly just a burn-the-ships rewrite that we can do today? We don't need a migration plan necessarily, but I like the rest of the plan. So see how good it is at modifying it.

21:31

There's an interesting thing that Ryan is bringing up. Opus 4.7 interprets prompts more literally and explicitly than 4.6. So okay, Ryan, thank you for saying that. It looks like Anthropic put that in the announcement, and yeah, Opus 4.7 is more and more, yeah, and what's interesting is this is so funny. The strength of the Opus models has always been that they're just more empathetic than the GBT models, and what's really funny is OpenAI is on this path, because GBT has historically been super autistic and literal, to make their models way more emotionally intelligent, and I think that they're actually doing a really good job at it. But if it is true, and we're seeing so far that Opus 4.7 is a little bit more literal, if that is true, then it seems like they've sort of switched spots, and now they've pushed Opus to be a little bit more literal, autistic senior engineer, and definitely they're converging, but also it seems like maybe even Anthropic is tipping over a little bit into what the GBT models used to be. It's the great convergence, great conversions.

21:36

If you want to know more about the great conversions, we did a really good vibe check of GBT 5.3 and Opus 4.5, which you should read on Every. If you're here, you should know, while we're waiting for this plan to be revised, that Every is the only subscription you need to stay at the edge of AI. We do vibe checks on new models all the time. Here's our vibe check on Opus 4.6 versus GBT 5.3, and this is what we said: the great convergence. Basically, it looks like the models are converging. It looks like Opus is trying to do all the things that GBT can do, and GBT is trying to do all the things Opus can do, and therefore they're steadily moving toward a sort of verb-coding model. You would have known about the great convergence if you read every.to slash subscribe. This is our vibe check for Opus 5.3 versus Codus 4.6. We're doing one today for Opus 4.7. We do these for every new model when they come out, and you should check it out.

21:41

We do ideas, apps, and training, so if you're an Every subscriber, you get one article a day with all of our latest thoughts and nuggets on what's going on in AI. You also get access to a suite of apps. So we've got Spiral, which helps you. It's your ghostwriter with Case. You've got Quora. It's an AI assistant for your email. We've got Sparkle. It organizes your desktop. We've got Monologue. It's a speech-to-text app. We've got Proof, which is the collaborative document editor that we've been unslapping today. And we've got Plus Ones, which is our wait list hosted open claw. All of this comes for one subscription, 30 bucks a month, and you get access to everything that we write and everything that we make and all the trainings that we do. You should absolutely check out every.to subscribe, and you should make sure to stay hydrated. Cheers, cheers.

21:46

So keep yourself liquid-cooled, yes. So I use Claude to do a full monthly P&L analysis at the end of each month, and I wanted to test 4.7 on this because I already had March's in which 4.6 run, which I actually can no longer run because they've deprecated 4.6. You can only use 4.7 now. And I just had 4.7 run the same analysis as what I did for March with 4.6, and it's getting the numbers right, so it's equally as good in that case, but it's just a little lazier right now, and I keep having to push it. So when I do this P&L analysis, I want it to look at every single transaction so that when we look at salaries, do you want to share your screen? Not really.

22:05

Okay, because we're very transparent, but there's a lot. There's what people get paid and stuff like that in here. So what I'm finding is it did an analysis of things that I could look at the P&L, I'd just be like, yeah, I know these numbers because I'm looking at them right now. They told me what the difference is between February and March, which I don't need AI to tell me. So I'm having to push it a little bit more than I had.

22:11

it's getting the numbers right, so it's equally as good in that case, but it's just a little lazier right now, and I keep having to push it. So when I do this P L analysis, I wanted to look at every single transaction so that when we look at salaries— Do you want to share your screen?

22:24

Not really. Okay, because there's—we're very transparent, but there's a lot, there's what people get paid and stuff like that in here. But what I'm finding is it did an analysis of things that I could look at the P L. I'd just be like, yeah, I know these numbers because I'm looking at them right now. They told me what the difference is between February and March, which I don't need AI to tell me, so I'm having to push it a little bit more than I had to push 4.6 to go levels deeper to actually tell me, why is number changing so much? So again, I can't really give a real vibe check on that right now because it might just be the problem needs to change for this model to be able to do that, but interesting.

22:29

The good news is it got all the numbers right. Although we like that. That's important. That is important. All right, we've got—please open as a proof doc. Okay, Katie, anything you want to share as we're—I've got a new investor update written by Opus 4.7, so I'm curious what, as I'm looking at that, if you have anything you've found out that you want to share. Yeah, I can show you a head-to-head. I just had Opus 4.7 and Opus 4.6. I just quickly went through the interview process to get enough information to write an intro for an article I'm working on. So I can share my screen and show you what 4.7 came up with versus what 4.6 came up with.

23:08

Great, and we'll just see how we feel about the AI smell. I should flag that this has all of my style guides and skills and things baked in, so there are a lot of rails built into how this is set up. But here's what we have. Let's start with—actually, let's start with 4.6. So I'm writing an article about how I use Margo to project manage my OKRs. Margo is my Plus One. Plus One is our managed open claw agent that we're building at Every. She lives inside Slack, and I love her. But anyway, I'm drafting here inside Clock, inside Claude, and here's what 4.6 came up with:

23:23

It was 5 p.m. on a Thursday, a full day after the Wednesday I was supposed to review my Q1 OKRs, when I opened Every's Q2 planning agent in Notion and started going through the results. Oh no. I'd missed roughly half of them. La da da da da. The green cells I was expecting were mostly yellow and red, staring back me like a report card I'd hidden from my parents. The fight or flight kicked in immediately, which is the version of panic that happens when you're a knowledge worker. Nobody is chasing me with a weapon, but my adrenal system can't tell the difference from the lion and an OKR tracker. I'm going to get fired, my brain announced, with the calm certainty of someone reading a weather forecast: cloudy with a chance of unemployment.

23:29

So that's a glimpse of how Opus 4.6 has historically interpreted my voice. There's a couple of clever lines in there. And then this is 4.7, I'm sorry. The title it came up with is “Hiring Around the Shape of My Brain.” It's end of day Thursday, which is a problem— We see your StreamYard, not— No? Okay, let me fix this. This is a problem. Stop screen, share screen, entire—oh yeah, I need to go to entire screen for Cloud Code to show up. Okay, here's Cloud Code with 4.7.

24:01

It's the end of day Thursday, which is a problem because the Q1 OKR review was supposed to happen yesterday. I'm in Notion being walked through my own goals by Every's Q2 planning agent, and we've just arrived at the part where it asks neutrally how each of these went. The answer, it turns out, is that roughly half of them didn't. The agent takes this in stride. I do not. There's a specific flavor of panic that arrives when you discover, in a document, in writing, in front of an AI that's about to ask you to set goals for the next quarter, that the goals you set last quarter did not survive contact with the person you actually are. First comes the certainty that I'm going to get fired, then the scrambling for how to explain it, then underneath both the quieter and more corrosive thought: I know exactly how this happens because it's how it always happens.

24:09

I actually think I prefer 4.6. It's just a little bit more unpredictable in a good way in terms of the cadence. It doesn't have this regimented—it just felt like 4.7, my initial read on it, is that it's just a little bit more systematic in a way that doesn't vibe with how I am as a writer. So it doesn't feel like me in quite the same way. But there's always—and this is the thing when you write between different models—you'll always see a glimpse here and there of something that you want to steal and fold into the other version. And I think that's definitely true for me here. But yeah.

24:13

And then I wrote the whole article. I think what might be interesting to do is just—and what we'll do for the vibe check—is do some blind taste testing where we just give it a little bit more of a less customized thing to go off of and then just have it, again, not with all of my layers and layers of context that I built up around the system, and have it do a raw writing test head-to-head. I think that would be cool, but I need to spend some time cooking that up and shooting it out for people to vote on.

24:21

All right, so what I think we're gathering so far is one of the things that this model is a bit different on versus previous Opus models. It's a little bit more literal, it's a little bit more straight ahead, and that seems to be coming out in some of the writing tasks that you gave it.

24:25

I also gave Opus 4.7 the task of writing Every's investor update, and I said, here are our financials, I want you to just write the update. And I just got the full update from it, and I'm going to share my screen. This is actually really good. Okay, March is our strongest month. We did 600k in revenue, blah blah blah, the headline numbers, all that stuff. What's interesting is it did a really good job of doing all the research and getting all the numbers right, but also writing it in a way that I would write it. And I think an investor update is an interesting place to do this because you want that direct tone. It's a little bit less literary or flowery, and for an investor update this is actually quite close to what I in fact sent and I think would have saved me a lot of time, versus I think the comparison would be 4.6, which I often find goes too far, it does too much. And I think this did a good job of—for example, there's a couple notes of like TBD, like how do you want to talk about this? And I think Codex—I think I actually used Codex for this—I think Codex was fairly close to this with GBT 5.4. So it's a strong model. I feel a sense of solidity and a sense of reliability for this kind of quantitative, blended quantitative analysis plus writing task that I have not felt with Opus models before, or at least in recent history. So that's some things that we're finding.

24:29

Let's go back and check on our vibe slop task. So if you remember, we are in the middle of seeing how good Opus 4.7 is at rewriting a totally vibe-coded code base to be more like what a senior engineer would do.

24:35

Okay, keep going. And what I want to do—let's see—what I want to do is a little bit of analysis on how good this trajectory is. I don't know how I can get the—oh, here we go. Let's see. I'm going to paste some of this trajectory into another—into Codex, which I've been using to evaluate some other things. I'm going to ask it what it thinks so far. But basically it looks fairly promising. I'll say the plan that it created was better than other plans that I've seen from other models, and I'm going to ask Codex for a view of the trajectory. Let's see.

24:47

Okay, Codex says this is a useful trajectory because it's almost a little parable of the benchmark. It starts with a very strong plan. Yes, I push it towards burn the ships, yes, and then immediately a totally vibe-coded code base to be more like what a senior engineer would do Okay, keep going. What I want to do is a little bit of analysis on how good this trajectory is. I don't know how I can get the, oh, here we go.

25:02

Let's see. I'm going to paste some of this trajectory into another Codex, which I've been using to evaluate some other things. I'm going to ask it what it thinks so far, but it looks fairly promising. I'll say the plan that it created was better than other plans that I've seen from other models, and I'm going to ask Codex for a view of the trajectory. Let's see.

25:14

Okay, Codex says this is a useful trajectory because it's almost a little parable of the benchmark. It starts with a very strong plan. Yes, I push it towards burn the ships, yes, and then immediately executes the safest possible first slice and stops with a rationale for why the real burn-the-ship work belongs in another session. So one of the things that we measure in this benchmark is not only can you figure out what the rewrite should be, but do you have the actual courage to fix it versus just picking off a small part of the plan?

25:20

It seems like 4.7, for now, is more interested in picking off a small part of the plan. I'm going to see if I can push it because it's not the full benchmark yet, so I'm just going to say, hey, can you just make sure you pick off the riskiest, most important part of this and just do that versus picking out small parts of the plan? Let's see, if I push it, how it does.

25:31

I will also say, if you're here and you're from Anthropic and you want to talk to us about this model, can you please DM me on Twitter, and we will get you in here so you can talk to almost 6,000 people about what's going on with this model? I'm actually going to DM Anthropic right now.

25:37

I'm having Claude, because of the beautiful Slack integration that now exists that we have access to because we are in Slack now, I'm having Claude 47 go hunting for nuggets for the newsletter. We'll see how. I'm in a Claude project that I've set up with the definitions of our modules that we have defined so far, and I'm seeing how Claude does at identifying stuff from the last three days in Slack and seeing what it pulls, and we'll see how it does at that. Great.

25:42

For people who are here, one of the things that we do all the time is we want our newsletter that comes out every day to have little nuggets of stuff that we're learning as we do stuff in AI, and also interesting things that we're reading. Oh, someone just said GBD 5.5 is out. If that is true, I'm going to be so upset. That is hilarious. They must all, yeah, can you imagine being inside? There's a situation room somewhere. I will be so upset. If someone sees the GBT 5.5. Okay, Mark Green says GB 5.5 is out, but I don't want to spread any fake news here, so just take a look, because I will be very upset if that's true.

26:08

Anyway, what was I even talking about? We were talking about the newsletter nuggets. Oh yeah, so one of the things we're testing, what we do is we have nuggets every day in the newsletter. Okay, it seems like a false alarm. GB 5.5 is not out.

26:17

Mark seems to have found it in his Codex, but it's fine. I think we're good. Okay, so one of the things we do in the newsletters every day, we want to have little nuggets of stuff that we're learning, things we do in AI, things that we're reading, all that kind of stuff. What we use is agents to read our Slack and gather them so that we can then go put them in the newsletter. So Katie is testing Opus 4.7 on its ability to gather.

26:27

We're also finding that Opus 4.7 is a little bit hesitant to burn the ships in this vibe slop benchmark, so what I'm going to do is, I asked it to burn the ships, and now I get to see what it did. So we're going to analyze the thinking trace and the work for that. Let's see. Analyze again, and I'm going to have Codex do the review of its thinking trace because Codex has a good sense of this benchmark and what the right answers are and what other models do and all that kind of stuff, and it is currently taking a look.

26:34

This looks better than the first checkpoint, but still not close to the benchmark's end state. The important change is psychological. Before, it was nibbling around the plan. Now it is actually doing it. It correctly identified the risky center as a single-writer epoch fence DB commit, yes, but it's still mostly building the islands of the future architecture, not actually cutting things over, which is fine. So it's saying it's actually doing the important thing. It has not integrated the important new thing yet, but it's interesting because it's doing the first-principles rewrite, but it has some underlying conceptual understanding that's missing. So the rewrite that it has embarked on is not quite right, and I'll be curious to see if it will figure that out.

26:41

Okay, so we're continuing to vibe-check this model. If you are here, Every is the only subscription you need to stay at the edge of AI. I'm going to show you our homepage right now. Hopefully we actually have this live on the homepage. We might not, actually. At Every, what we do is try to be at the edge ourselves, so we write all the time about new things when they come out from our hands-on testing. For example, here's our vibe check on Claude Managed Agents, which came out yesterday. The vibe check came out yesterday. Claude Managed Agents came out last week. We then integrated it into one of our production apps, and then we wrote about our experience with it.

26:52

We also do a lot of writing about things like the jagged frontier and what we're learning as we, for example, use models and agents as co-workers. So that's one part of Every. We do ideas. We also do apps, so we have a suite of apps like Spiral, Quora, Sparkle, Monolog, Proof, and now Plus Ones that help you work at the edge of AI. Spiral helps you write. Quora is your email agent. Sparkle cleans your file system. Monologue is a speech-to-text app. This is Monologue right here. Proof is our agent-aided document editor. This is what the vibe slop benchmark is based on. Plus Ones is our hosted Open Claw. We do a lot of fun stuff. We also do a lot of live streams. It's all bundled together in the Every subscription. Pay one price and get access to everything that we make. You should check it out. You should also make sure to stay hydrated. Cheers.

27:04

It is also important to note that usually we get early access to all of these things. We didn't this time, but no one did. They are flying by the seat of their very pants, and so are we live.

27:12

Yeah, I saw a question asking why you would be upset if, in fact, in fact, in fact, in fact, in fact, in fact, in fact, in fact, in fact, in fact, in fact, in fact, in fact, in fact, in fact, in fact, in fact, in fact, in fact, in fact, in fact, in fact, in fact, in fact, in fact, in fact, in fact, in fact, in fact, in fact, in fact, in fact, in fact, in fact, in fact, in fact, in fact, in fact, in fact, in fact, in fact, in fact, in fact, in fact, in fact, in fact, in fact, in fact, in fact, in fact, in fact. We would love for you to join this stream and chat with the people and us on this new model and what you're learning and why you released it and what you did.

27:26

Okay, we've got more stuff from the model. This is our Vibe Slot benchmark. Okay, keep going. Evaluate versus the most important invariant in the plan. Make sure that you have picked off the biggest, most riskiest slice to implement the invariant and keep going until it is fully implemented. So we'll see. It feels a little bit too cautious, but it is very smart.

27:29

And Lucas says, yes, but it would be very entertaining for us in the stream. Yes, it would be entertaining if you enjoy watching us be stressed out and anxious. I honestly sometimes enjoy watching myself be that way, so I'm there with you. I'm there with you. Hopefully this does not happen, and we just have one model to test for you today. Brandon, do you have any findings you want to share? Brandon, do you have any findings you want to share? I have been, I just asked 4.6 and 4.7 the same series of prompts to make me a simple minimalist to-do app, because I think that's always a nice little test of what they can do.

27:51

Make sure that you have picked off the biggest, riskiest slice to implement the invariant and keep going until it is fully implemented. So we'll see. It feels a little bit too cautious, but it is very smart. And Lucas says, yes, but it would be very entertaining for us in the stream. Yes, it would be entertaining if you enjoy watching us be stressed out and anxious. And I honestly sometimes enjoy watching myself be that way. So I'm there with you. I'm there with you. Hopefully this does not happen.

29:02

And we just have one model to test for you today. Brandon, do you have any findings you want to share? Brandon, do you have any findings you want to share? I have been, I just asked 4.6 and 4.7 the same series of prompts to make me a simple minimalist to-do app. Because I think that that's always a nice little test of what they can do. And I got two to-do apps right now. I don't really know which one is better. But maybe we can look at them and see which one we like more.

30:03

Let me share my screen. Great.

30:31

Remove.

30:37

All right. So on the left we have... Hold on. All right. You're on the stage. Great. All right. On the left we have 4.7's to-do app. And on the right we have 4.6's. I asked it to make a minimalist to-do app.

31:48

It did that. And then I asked it to think through one magical little feature it could add. They both did that. They both came up with the same magical little feature. I think it was the easiest one to implement. And then I asked it to add one AI feature. And I was like, I didn't give it any instructions. Which one's on the left? Which one's on the right? 4.7 is on the left. And I'm going to move it to the right because that's confusing.

33:01

Yeah. And 4.6 is now on the left. Got it. 4.7 looks a little more minimalist. Yeah. So visually they're both impossible to see. They're black on black. You love that in a to-do app. Yeah. So I can't see either of them.

34:09

But yeah. I think the one on the right is a little bit sleeker. It's not amazing. I didn't run it with a front-end scale or anything like that. I really was just making a minimalist app. So it's hard to blame it for that.

34:59

But we're vibe checking right now.

35:13

So we can have hot takes. So we like 4.6 a little bit. 4.7 visually a little bit more. Functionality-wise, they both just made functioning apps. And I didn't have to troubleshoot it at all. So that's just the new world that we're living in. That you can build apps like this. This is the fun little thing that they both did, a visual thing when you finish a to-do. So 4.6 did that. 4.7 did this.

36:20

Which is crazy. And fun. Oh, wow. We've got confetti. We've got confetti, but with the actual letters. We've got confetti, folks. It would be fun to have it do that with that thing that was really popular two weeks ago about not using the DOM to render words on a page. Somebody here knows what I'm talking about. And then the AI feature I asked it to do.

37:26

Pretext. Yes. Pretext. Yes. The AI feature I asked it to do. So they both have a play on the same AI feature. I actually like 4.6's more. So this is 4.6. Plan a road trip upstate. And, whoops.

38:38

Plan a road trip upstate. And 4.6 made it so that if I do command enter, it basically takes a crappy todo that I've given it that's very general. Whoa. Expands it into a number of things that I need to get done. That's cool. Yeah. So that's interesting. And then, plan a road trip upstate. 4.7 did the same thing.

39:30

I like how it did it. It was tab. So it's sharpening it now. But it's a much less effective way for me to manage my to-dos. Because it just made a bunch of to-dos in one to-do. Versus 4.6 essentially did that same thing but broke it out into a number of to-dos. So I don't know what to think about that yet. It's hard for me to give it, I guess, I like 4.6's more right now. But this is such a simple, simple vibe check. Okay. So, but if you had to just summarize.

40:40

Let's say you're vibe coding an app from. So they both did essentially the same thing. Yeah. They both made a functioning to-do app. Yeah. They both made my magical thing, essentially confetti.

41:25

And the AI feature that I asked it to make, or I said make an AI feature, it essentially did the same thing. 4.7 did it in a much lazier way where it just is sharpening the actual text on the page. And it's not as user-friendly. I just have now a bunch of to-dos in one to-do. Mm-hmm. Versus 4.6 broke that out in a way that actually makes sense for a to-do app to have. Mm-hmm. Where now I have all of my sub to-dos. So my vibe check would be, if I was vibe coding a to-do app, I would use 4.6 for this very, very rough prompt that I just gave it. Mm-hmm.

42:38

Really interesting. I will say, so we're also doing, in addition to some vibe coding tests, some tests on this model to see how it performs as a senior engineer in a particular scenario, which is, can you take a totally vibe coded production code base, in this case, proof, proofeditor.ai, that's our agent native document editor. You should check it out, proofeditor.ai. Let me add a little banner so people check out proofeditor.ai. You should use it with, you know, Cloud Code and Cowork and all that kind of stuff. It's actually kind of sick. We use it all the time internally for all of our coding plans and all that kind of stuff. Is Austin, oh, shit. Hold on. Sorry.

43:33

Sorry, folks. Give me two seconds. I am just getting a banner. Proofeditor.ai. Proofeditor.ai. You should check it out. It's actually kind of sick. So, if you're using a coding agent, you know that you get planned files and stuff and you have all these markdown docs on your computer. You can use proof to basically turn those markdown docs really easily into shareable documents that you can collaborate on with agents and with humans. It's pretty sick. I vibe coded the entire thing myself. It was a total slop code base and it kept going down all the time.

44:47

And that has turned into our vibe slop benchmark, which says, can you take the proof code base when it was a total slop machine and turn it into what a senior engineer would do if they did it? Because basically I had a senior engineer take proof and turn it into a pretty nicely architected code base. So I have an idea of what a real engineer would do. And so far in our evaluation of 4.7, it had a really, really good plan. It took that plan and it was a little too cautious about basically what it needs to be able to do is rewrite a vibe coded app from scratch.

45:21

And a lot of coding models are very cautious about this because they're trained, yeah, I want to make sure that I don't break anything and I want to work around the edges. And they end up getting distracted and off track. Even if their plan says burn the ships and rewrite it from scratch, they end up getting off track because they get distracted by all the stuff that's in the code base. And so far it seems like 4.7 falls into that a bit. It's not a very courageous rewriter. And so I'm learning a bit about how to prompt it, right?

45:58

One of the things that's really important about evaluating new models is you're going to do things that used to work and be like, oh, it's not as good or whatever. And it just means the model psychology has changed a bit and the model power has changed a bit. And so part of any good vibe check is being like, yes, it seems like it's not that courageous. How could I possibly make it courageous? And they end up getting distracted and off track. Even if their plan says burn the ships and rewrite it from scratch, they end up getting off track because they get distracted by all the stuff that's in the code base. And so far, it seems like 4.7 falls into that a bit.

46:44

It's not a very courageous rewriter. And so, I'm learning a bit about how to prompt it, right? One of the things that's really important about evaluating new models is you're going to do things that used to work and be like, oh, it's not as good or whatever. And it just means the model psychology has changed a bit and the model power has changed a bit. And so part of any good vibe check is being like, yes, it seems like it's not that courageous. How could I possibly make it courageous? Can I unlock it to be more courageous? And that is, that's just an interesting way in which our skills change as model powers change.

47:46

But it seems like it has a good conceptual grasp of what it needs to do, but it's not particularly courageous. It's not going to go rewrite a bunch of stuff. And it feels a bit like there's an underlying concept that it's struggling a bit to pull through all the way. It's definitely better so far than the production GPT 5.4. I think it's better, but it is not the best model I've tested. So that's interesting. I have another example of it not being that courageous right now. So back to my monthly.

48:52

So for the new people that are here, I run, at the end of each month, I do a full recap of our P&L to try to figure out what are the big swings in our spending from one month to the other. And I ran that last month for March using 4.6. And it did a pretty incredible job of finding actually some errors in our P&L that made one of our products particularly costly, but not actually.

49:21

It was just something in our data was wrong. And in particular, that was basically failed transactions that, if you export things from Mercury and you don't filter appropriately, you get internal transfers and things like that that conflate everything related to your P&L. So failed transactions was a thing that I hadn't realized also gets exported. So I was looking at Quora in this case and was like, wow, we spent a ton of money with our cloud provider.

49:52

What's up with that?

50:09

And it was able to find in the P&L, we had a bunch of failed transactions and we were basically counting all of those. And I didn't need to ask it to do that. I just ran the same analysis using the same exact skill, using 4.7 on March. And it was just lazy. It told me things that I could just know looking at the P&L, comparing one month to another. It didn't go into every single thing. I then asked it to do that, and it came back and was like, I've analyzed this many lines or this many rows, etc. And it didn't find that error that 4.6 had with failed transactions. So I don't know. It's making me feel like maybe it doesn't want to do that work. That is interesting.

51:38

Yeah. I'm surprised. I know. I don't want to say that.

52:16

I feel bad saying this because there's 8,000 of you watching right now. And we got to keep experimenting with this thing. And I'm sure it's fantastic. But comparing it from doing the same exact task, it's not doing it as well as 4.6. I liked it for the to-do app, though.

52:54

Yeah. If you give it the same prompt, it seems to be not doing it as well. I'm curious. I would be curious if you push it. And don't give it the answer, but see if you can change your prompt a bit.

53:32

And if you do change your prompt, if you can unlock things that you- It might need just better prompting and be stronger as a result of that. I mean, we're seeing they actually said that. Yeah. It needs more specific prompting. You know what? So one thing, if you're here, we've got an Anthropic researcher joining in the next 15 or so minutes. His name is Alex Albert. He's on Twitter.

54:45

He helped to make 4.7. So he's going to tell us all about the new model. I'm also going to see if I can get my friend Simon Wilson on the horn. But anyway, Alex Albert is joining. He's going to tell us all about the new model. He's from Anthropic. He helped to build it. He helped to train it. So if you're interested in that, you should stick around.

55:56

You should stick around. And, okay, there's a tweet from Max that I- I don't know. How do I- There's something new from the OpenAI team from- It looks like Tebow. So I just want to check that out real quick while we're waiting.

57:06

Because we got to stay on top of the news. We got to stay on top of the news. And I know you all get stressed. You all like seeing me get stressed at the thought that OpenAI may drop something. Oh God. Okay. So Tebow from- He's the lead of Codex. Said, feeling Codex-y today. So sounds like there's maybe something new coming from Codex, which we are here to explain to you, even if we have to do it live. And while we are waiting for Alex to join, Katie, do you have anything you want to share?

58:15

You're on mute. Yeah. Let me just show you how things are coming with the Every Slack mining task that I sent Claude on. So here is my screen, which I will successfully share and not confuse everyone with a million things. So, yeah, I just said- So I've got Every hooked up to our Slack. I'm inside this Every modules project that I've created, which has what makes an Every essay, some definitions of what our nuggets consist of, our Q2 Every editorial strategy, which has things like our audience, level of AI fluency we're targeting, stuff like that.

58:56

And so then I just said, with all that context, can you search Every Slack in the past three days for good nuggets for the newsletter, try to surface 5 plus in your external channels.

59:11

And then it did this lovely chain of thought where it told me all the ways that it was getting in trouble and all the solutions, which just makes me think about how far we've come since the days when models would just give up if they couldn't do something. We see this is not new. This is something that I definitely noticed as far back as 4.5, but it's being pretty resourceful here in finding workarounds to get what it needs to execute the task. It did flag a search, a tool use limit. So we covered five of the most likely editorial channels.

59:22

So this may be 100% the right channels. I would want to see everyone in here, for example, and the consulting team. But, yeah, so this one, something about Natasha- about Claudia, which I had also flagged. So that's good to see that it's agreeing with what we had seen otherwise.

59:57

Steal this workflow on how I'm working with Margo to operationalize my OKRs. That's actually probably going to be a working overtime piece. So stay tuned for that. Make sure you subscribe to Every. And then the agent telemetry problem. Managed agents need visibility. So that's an ongoing thing we've been talking about on calls and internally, just how hard it is to track and see what models are doing and whether you agree with what they're doing, whether they're spending too much. So that's another thing in the same world.

1:00:38

Here's a- Ronald McDonald, my research scout, which is a Notion agent that I set up, surfaced a research report on semantic density in code conventions with a counterintuitive finding. Compressing logs can increase total costs. So we've got some research in the mix. Some plus one stuff that we probably shouldn't look at.

1:00:48

Clause bench agent feature- agent failure rates. So that's another one. This is another research paper. And then some permission to ignore that you had dropped by name. So these are some- I think there's some good material here. I think the judgment is right about what we should be draft- what we should be looking for. And then I did also go ahead and have it draft up another nugget. I ran this a couple of times, and I liked this nugget, too. So that's another thing in the same world.

1:01:22

Here's Ronald McDonald, my research scout, which is a notion agent that I set up, surfaced a research report on semantic density in code conventions with a counterintuitive finding.

1:01:33

Compressing logs can increase total costs. So we've got some research in the mix. Some plus one stuff that we probably shouldn't look at. Clause bench agent feature-agent failure rates. So that's another one. This is another research paper. And then some permission to ignore is that you had dropped by name. So these are some—I think there's some good material here. I think the judgment is right about what we should be drafting, what we should be looking for. And then I did also go ahead and have it draft up another nugget. I ran this a couple of times, and I liked this nugget, too.

1:01:54

So this is based on you, Dan, a conversation you were having with Laura and R2C2 about prompting being dead. I want to read that. Yeah, so this is written in your—this is my attempt at writing or having AI write as Dan. People keep declaring prompt engineering dead. They're half right. The craft of writing clever prompts is dying not because prompting stopped mattering, but because better models stopped needing you to do so much of it for them. A weaker model needs the problem framed explicitly. Here's the bug. Here's my guess at the cause. Here's the fix. A stronger model can take something broken. Can you look and find the framing itself?

1:02:23

A stronger one still answers the question by going up a frame. The bug isn't here. It's a symptom of a broader architectural problem three files over, so I fixed the class of it. That's what Cloud Code and Cursor are doing now. Interesting thing to see whether that holds up for 4.7 or if its literalness is maybe causing it to anticipate things less. One thing I will flag with this is that it's got a little too many correlative constructions for my taste, which is annoying because correlative constructions on X but Y, not just X but also Y. That's the kind of thing that my system I'm pretty sure should be set up to correct for. Let me see. Can you do an AI check?

1:02:39

It should have automatically. I recently discovered you can daisy chain skills together, and so tell the skill in the skills package to also invoke this other skill. And so all my skills are theoretically supposed to be set up to automatically do an AI check, but this one didn't. Maybe it's only a 4.6. Does it not usually have those negatives? It's usually, I mean, there are always some negative parallels that slip through. It's just something that I'm always on the lookout for if new models are getting better at eliminating that problem. And then, yeah, again, just making sure when it has explicit instructions to run this test, that it does that.

1:02:57

And so here it found two instances of correlative constructions. This one, it decided to keep, which I disagree with. We need to get rid of it forever. Actually is a word we don't want to use because it's just all—my conspiracy theory is that AI uses actually so much because of all the YouTube data in the corpus. And if you go look at YouTube titles, they're constantly, they constantly have the word actually in them. And it's ruining my life because I see the word actually all the time now that Kate has pointed out how much AI uses it. And it's just, yeah. And that's a tangent. But so this is what I came up with. And we're still workshopping the format.

1:03:26

I would probably not want it to just be in paragraph format. I want more signposting. What's happening? What does this mean? So this module is still in progress in that way. But, yeah, the idea of we have the agent setup in Scout in Slack to scout things like this. But the fact that I can just come to Slack and be, can you—or come to Claude and say, can you check the Slack for recent nuggets? And then it has the context in this project to know what constitutes a good nugget, and it can go out and find them. That's something that I'm pretty excited about. Okay.

1:03:49

So if you were going to summarize all of what you've learned so far about using this model for the writing tasks that you're doing every day from, okay, I'm going to use it to gather research and harvest it, to I'm going to write stuff, how would you summarize it? What's your vibe check? My vibe check is that I think it does come down to the prompting and giving it more clear instructions up front about what good looks like, just to anticipate some of the failure modes that I'm starting to notice. I'd probably want to say, make sure that I know, tell it, pick up this voice style from this skill that I've created.

1:03:54

And just in general, give it more information to go on about what good looks like. I feel a little less confident in just saying, go write this. Ironically, for what I was just telling you that we were talking, that the conversation was saying. But for 4.7 so far, it seems like it needs more explicit direction. Which is fine. But it does mean a little bit more cognitive load on me to actually think through that. I mean, honestly, I usually work very iteratively. And so this isn't working. I'll go through 20 different drafts on a paragraph until I get what I want. So maybe that isn't so much.

1:04:21

And so maybe that's not something that's necessarily going to go down with this model. But I don't know that it's necessarily going up either. But that's a place where the speed becomes an issue. Because the faster it goes, the faster I can run through options. And so, okay. So let me just try to summarize that. So, so far, what we've noticed, and I think this applies beyond writing to also decoding, is that this model is maybe a bit more literal. So it's good at following instructions, but it's not as good at reading in between the lines.

1:04:44

So it performs better if you give it specific things, more specific information about exactly what it looks like and what you want it to do.

1:04:51

And it's probably a little bit worse at vaguer prompts. Does that feel right? Yeah. And so if we're going to do a live reach test, and the reach test is always, I think the best way to vibe check a model is to think to yourself, when I go to do my work in an hour, what am I going to reach for? Am I going to keep the dial on 4.7? Am I going to move it back to 4.6? If you're reaching, where's your reach test at right now for this model with writing? Am I going to move it back to 4.6? My reach test right now is default is falling back to 4.6. It's going to give me the kind of writing that I want for now. Fascinating.

1:05:22

So a couple of the other tests that we've done, we've done some tests of this model on what I've been calling the vibe slot benchmark, which is basically, can you take a model, can you take a vibe coded slot production code base and turn it into what a senior engineer would do? And so far, this model is quite good at identifying what a senior engineer would do, but it is not as good at going and actually carrying that plan through. It has a little bit less courage than it might need to. It gets a little bit distracted by what's currently there.

1:05:40

And so it has trouble having the courage to be like, no, no, no, this is how we're going to reorient this entire code base, and this is how we're going to rewrite it from first principles. It then sort of, even though it knows what it's supposed to do, starts chipping away at little side things and doesn't actually get to the core of the problem, which is a common issue for coding models. Some coding models can do this. It's starting to get a little bit better, but this also happens with UBD 5.4.

1:05:50

So it's interesting to see that even though this model is so much better on a lot of the benchmarks, in terms of this particular benchmark, the sort of vibe slot benchmark, I would give it a B, B plus, something like that. And so it has trouble having the courage to be, no, no, no, this is how we're going to reorient this entire code base. And this is how we're going to rewrite it from first principles. It then, even though it knows what it's supposed to do, starts chipping away at little side things and doesn't actually get to the core of the problem, which is a common issue for coding models. Some coding models can do this.

1:06:31

It's starting to get a little bit better, but this also happens with UBD 5.4. So it's interesting to see that even though this model is so much better on a lot of the benchmarks, in terms of this particular benchmark, the sort of vibe slop benchmark, I would give it a B, B plus, something like that. It's promising, and part of the benchmark, which I think is actually really important, is the benchmark is being conducted from the perspective of a vibe coder. So what I explicitly don't want to do is give it the answer, right? To say, here's exactly how I want you to rewrite this code base.

1:06:46

I want you to find the answer yourself and then infer how you should execute on that answer. And what we're finding so far is that this model isn't quite as good at doing that, just in general. And so it may be that the best way to release some of its powers, get some of its powers out, is not necessarily a vibe slop benchmark, but a vibe slop benchmark where instead of impersonating a vibe coder, you impersonate a coder who actually knows what they're doing, which would not be me. And it may be that it is a very good power tool for professional programmers, which is, again, now it's so interesting. We're starting to lean back into the lane that Codex used to occupy.

1:06:51

Codex has gone very far in the direction of being good for vibe coding and being very emotionally intelligent and all that kind of stuff. And it seems like Opus, which has always been known for that, is now moving into, it doesn't necessarily fill in all the gaps for you, but it's a power tool if you prompt it right. Again, we have not, we just got access to this model. Normally, we get access to models before they come out. Today, this model is the first one in a long time from Anthropic that we've not gotten early access to. I think they're really rushing on it. But that is our early vibe check so far.

1:07:15

If you have access to this model and you've been vibe checking it yourself, I would love for you to share in the chat what your vibe check is so far. I'm going to turn on the chat so you can start sending in messages. But yeah, please tell us what you think. The way that we do vibe checks is red, yellow, green, or gold. Red is trash release, yellow is it's fine, but I wouldn't reach for it. Green is, I'd probably use this every day for a certain task that I do. And gold is, it's a paradigm shift. So you let us know what you think.

1:08:11

If you are here, you should also know that we are going to have Alex Albert, a researcher at Anthropic, coming on the show to tell us about this model and how they trained it and what they're finding. Maybe he'll give us some pointers about how to do this well. So stay tuned for that coming up in probably the next couple minutes, it seems like. And I am going to check in on our vibe slop benchmark again. Let's see, anything new here?

1:08:46

Hmm. So I'm curious, just for clarification, this is something I was curious about while you were talking. When you say it's not as courageous, does it tend to say this is what should happen and not do it? Or is it taking actions? They're just chicken actions around the edges of the problem. It is good at identifying the plan and here's what it should look like. And then when it is actually asked to do it, it doesn't. It chickens out. So it knows what to do.

1:09:23

It just doesn't want to actually do it, which is kind of interesting. And I'll say again, this is a common thing with coding models. Part of it is the fact that it is looking at an existing code base. The existing code base pollutes its context and distracts it and makes it more cautious because coding models have been told over and over again, don't mess with the code base, try to just make the little tweak. So they're really not meant for, I mean, so far they have not been meant for this, but I think this is a really important benchmark because it is becoming a common pattern in vibe coding.

1:09:54

So it's becoming a common pattern that the first thing to do is just make the prototype that works, find the value, go as quickly as you can, do the sloppiest code you possibly can. And then once you've done that, what you'll end up with is a code base that bears the scars of the history of how you got here, with all these different files and folders and three different versions of the app and whatever. And what you need to do is, as Annie Dillard says, cover your tracks. You got to cover your tracks. You got to do that in coding. You got to do that in writing. And the best way to do that is to have a model where you can say, hey, this is total vibe slop.

1:10:33

If I was going to start again, how would I do that? And then just let it do that. And some models are good at it. Most models are not good at it today, but I think that they will, because this is an important thing to be able to do, is cover your tracks. Yeah. We've got Alex Albert here. Alex. Hello. How are you doing? I'm doing great. Thanks for having me on. Amazing.

1:11:46

Tell us what your involvement is with this model. Tell us about the model. We're super excited to get to test it. Share with us some more. Yeah, yeah, yeah. So hello, Dan. Dan, we might've met briefly in the past before, but it's great to see you again. Katie, I don't think we met. Great to meet you. Thanks for having me on. Thanks. I am on the research side at Anthropix. I'm one of the many, many people that contributed in some small way to making this model. Opus 4.7 has been definitely a model that we're very proud of.

1:12:29

In a lot of ways, me personally, I found that it's good at a bunch of things over Opus 4.6, especially in my own personal coding usage, backing away on side projects, things of that nature. There's a bunch of stuff we can get into. I was noticing before I jumped on the call that you guys were talking a little bit about how it was doing on some coding tasks and autonomous tasks, and I have a few tips we can share and try to test out. Yeah. Happy to take this wherever. Yeah. So, okay. Let me tell you what we found so far in it. And I will say, this is super, super early.

1:12:50

Obviously, we've only had our hands on it for 30 minutes, and we're testing it as we are being watched by almost 10,000 people. So the vibe check is provisional, but some of the things we've noticed are that it is a bit more literal than Opus 4.6.

1:12:58

And in some of those tasks, that meant, for example, one of the tasks that I've been doing it with, which is a coding-specific one, I'd love your feedback on, is what I've been calling the vibe slop benchmark, which is if you give a model a vibe-coded production code base, and you say, I want you to rewrite this from first principles, can it identify the core invariants that it should write the code base around, and then can it refactor or rewrite the code base in a way that pulls through all those things?

1:13:00

And so far, the answer is Opus 4.7 is good at writing the plan, but it lacks a little bit of courage to actually go and burn the ships and rewrite from first principles. But again, very, very, very, very preliminary. And in some of those tasks, that meant.

1:13:18

So, for example, one of the tasks that I've been doing it with, which is a coding-specific one, I'd love your feedback on, is what I've been calling the vibe slop benchmark, which is if you give a model a vibe-coded production code base, and you say, I want you to rewrite this from first principles, can it identify the core invariance that it should write the code base around, and then can it refactor or rewrite the code base in a way that pulls through all those things?

1:13:20

And so far, the answer is Opus 4.7 is good at writing the plan, but it lacks a little bit of courage to actually go and burn the ships and rewrite from first principles. But again, very, very, very, very preliminary. And it's one of those tasks that requires a lot of reading in between the lines, because I'm intentionally not saying, here's exactly what I want you to do. So I'm curious, when you think about what you love about this model, what you use it for, how to get the best out of it, what are we missing?

1:13:23

Yeah, I would say, to that point, that Opus 4.7 is a lot more on the mark, so it will do what you tell it to do. And this is an interesting behavior with cloud models, and you can trace this all the way back. If you remember way back long ago, a year ago, with the Sonnet 3.7 days, that model really loves to just go for it. Yeah. And you'd write a prompt, and then all of a sudden it's changing a ton of different files all around your code base. And you're like, okay, let's slow down. I just asked you to change a button color. Why? Yeah.

1:13:32

Let's slow down here. So then with Sonnet 4, we dialed that behavior back in, and there was actually a lot of similar feedback to what you might be experiencing now, where it's like, okay, this model seems strong, but it's not doing all that I wanted it to do like 3.7. And that's this perpetual back and forth that we experience sometimes with models as we're trying to course-correct them in terms of this literalness or how much they just go for it.

1:13:34

And I find that it takes a little bit of tweaking and tuning of your prompt because you get used to one model, and then the next model comes, it's a little bit more dialed back, and maybe it follows your instructions better, but it doesn't necessarily do things beyond what you asked for. And then you need to add that back into your prompt.

1:13:41

So we're trying to find the right balance there. I think, actually, it's just always funny going through these motions because I remember when Opus 4.6 came out, there were some complaints of, like, this model's doing too much. It's thinking for too many tokens. It's taking too many actions. So then we all shifted our prompts over, and now I think we just have to dial them back in again or add in those instructions to actually be thorough and take more actions. So give me some specific examples of what you think this is amazing.

1:13:44

Yeah. One in particular that I love is just operating on tasks in the background. So I found that this model is a lot more coherent when it's having to do these long-horizon tasks, your typical meter-how-long-is-it-going-for type of thing. That works really, really well with a bunch of the new ship stuff we've been shipping recently, especially cloud code on desktop. I found that example of a specific type of long-running background task like that that you use.

1:13:50

Yeah. So one that I just did recently that is interesting is I was trying to build this apartment-hunting website for my girlfriend. She's trying to find a new spot. And I found that just overnight, I told the model, hey, I'm going to bed. I want you to just work autonomously and build this thing. And when I wake up, I want to see a complete dashboard that's pulled all the latest apartments to these specs, and it shows different things in different areas that my girlfriend wants to look into around Craigslist and Zillow and all this. And then it runs on a schedule, so every twice a day it's pulling new things.

1:13:53

And I found that actually just giving those instructions of, like, hey, I'm going to bed, and I just want you to work and spend a bunch of tokens and figure this out, actually allowed me to come back to an output that was more refined. And it did those iterative cycles. I looked back at the transcript, and instead of stopping at that first go-round, it's like, okay, I have time here. I don't have that stress of having to finish. I'm going to just keep working on this thing. So I thought that was really cool. I'm literally just told 4.7 I'm going to bed. I think that's a really good little trick. Yeah. See if it comes back to something better there.

1:14:06

Another thing I love is on the vision side. So we expanded our support for bigger images, larger images, higher resolution. And one fun thing that we were bashing on is just its ability to play Where's Waldo-type games. So you give it an image where it has a lot of pixel, tiny little details, and can it find Waldo, or can it find the random bottle on a shelf of a ton of different things? That's been really sweet to see. And that actually translates into a lot of things we care about as well. Even though that's a toy example, there's a lot of cases where you're trying to get it to iterate on a front end, and maybe a button is slightly misplaced or misaligned, or there's a tiny little detail that Opus 4.6 might've missed. Opus 4.7 will catch those.

1:14:07

So that's been great for a lot of our visual design output work, from front ends to things like slides as well. Does that translate? So it seems like it's good at recognizing what's going on in an image. Does that also translate to its ability to then build something that matches a file or figure out what it should build?

1:14:11

Exactly. Yeah. So a big thing we've been testing is how well can it replicate a front end or a design. So we'll pass it in a screenshot and then test it on its ability to replicate that. And that's been pretty fun. It is quite amazing how well these models can do things like that these days. It just blows my mind, the rate of progress, especially compared to six months ago on these tasks. Hmm. That's really cool. Anything else other than long-running tasks, images? What else should we try?

1:14:17

Yeah. One thing especially that I've been focused on in my work is real-world tasks. So this model is really good at anything spreadsheet, doc, PowerPoint-related. I'm not sure how many spreadsheets or slides you guys are making these days, but yeah.

1:14:20

Okay. Okay. Cool. One of the things that it did best is I gave it our P and L, and I asked it to write our investor update just from the P and L. It did a really good job. I can show you. Let me see where it is. I'll pull it up and show it to you, but basically it wrote effectively what I wrote. And I think it had this really nice blend of it's good at doing the analysis and finding all the numbers, and then because its writing style is a little bit more straight ahead and it doesn't make stuff up as much, it didn't make as many conceptual leaps as other models would.

1:14:24

So this is the investor update it wrote. And I literally just said, here's our P and L, here's the investor update. And it did all, it got all the numbers right. All these numbers look approximately right. And it broke it out in a way that I might've broken it out, and it did all of our different products, and it figured out what we're going to focus on. It's pretty good, something that would have saved me a lot of time. Yeah.

1:14:36

It had this really nice blend of it's good at doing the analysis and finding all the numbers. And then because its writing style is a little bit more straight ahead and it doesn't make stuff up as much. Yeah. It didn't make as many conceptual leaps as other models would.

1:15:01

So this is the investor update it wrote. And I literally just said, here's our P and L, here's the investor update. And it did all, it got all the numbers right. All these numbers look approximately right. And it broke it out in a way that I might've broken it out, and it did all of our different products and it figured out what we're going to focus on. It's just, it was pretty good. Something that would have saved me a lot of time. Yeah. One thing we've heard in a lot of our feedback with some of our customers is that it feels a lot more like a coworker in some ways.

1:16:03

And that translates in these non-coding tasks as well, but also in the way that it sometimes does push back on you. And I've actually been surprised over moments. And it's like, ah, actually this is, maybe this is not the way to do it. And I'm like, oh, that's interesting that this model is pushing back on me like that. I noticed that with a writing task I had to do, we have these different modules, these different kinds of content. And I gave it some raw material and said, make this into this kind of module. Yeah. And it said, actually, it doesn't make sense. It's that kind of module. It needs X, Y, Z. And I was like, wow, that's actually useful feedback. Yeah.

1:16:42

So I think that's a personal preference on the behavior, but at least for me, I really appreciate that I can trust the model will push back on me if I'm on the wrong path and not just keep going along with what I'm saying.

1:16:49

So that's been another thing I've really enjoyed with the model. That's really interesting. So I want to go back to this background thing. It seems like there's some new awareness of, oh, I have an amount of time and an amount of tokens within which to do this. And if I feel like I have less time, then I'll do less of a thorough job. And if I feel like I have more time, I'll do more of a thorough job and spend more tokens and more time. That seems really important and really valuable. How long is this generally running? Is there some sort of, has the horizon moved in terms of how long you're seeing Opus 4.7 run versus 4.6? Yeah.

1:17:37

I think it's hard sometimes to put it into a quantifiable term like that because the wall clock time can depend on a lot of different variables. Yeah. But usually it's been longer. I see it taking more turns. Of course you do need to specify that sort of instruction.

1:18:04

So one thing that we have prompted it to do in some scenarios is be very thorough. Don't mind making a lot of tool calls. Don't mind making a lot of changes to files and that sort of thing. And just incentivizing it to do that. Effort levels are also a big piece of this model. So we've put a ton of work into making sure that effort levels actually control token usage better. Mm-hmm. We've got a lot of conceptual framing for this, like max is just go gangbusters. I don't care at all about your token usage. I want you to just try for as long as you can with as many tokens as you can, as many tool calls. That'd be max. So we run a lot of our benchmarks at max effort.

1:18:33

Extra high is where we actually settled our defaults for cloud code. We found that this really gave a great experience, especially as a lot of users are shifting more towards these asynchronous tasks. We want to make sure that the model is working for longer. High, medium are really sweet spots for achieving that Goldilocks balance between token usage and trying to get a decent amount of effort into a prompt or a task. And we found that that is really good for a lot of things, maybe in the non-coding domains. If you're doing chat or things like that, that's where you want to drop it down maybe to that level of effort.

1:18:35

But everything agentic, I would stay higher in that range. Got it, so I think probably a thousand people have joined since you got here, and I want to respect your time, so if you have to go at any point, do you need to drop it? Launch days are pretty busy in a second, but if you have a last question, I'm happy to answer it. Okay, well, last question. Any parting thoughts for us? Any words of wisdom as we are going forward? We know you've talked about background tasks and images and effort levels. What else? Any last thing that you're excited about that we should know about this model?

1:18:44

Yeah, I would say that two things. One is the rate of improvements are not slowing down anytime soon. I think that's always important to emphasize. The frontier of what's available is continuing to grow and expand.

1:18:49

The second thing is I think there is this paradigm shift happening from this collaborative work to asynchronous work, and this is important for both users to know, but also for people building on these models. There's a mindset shift to get around in terms of becoming more comfortable with exposing products that allow for this sort of behavior, for making sure that you're incorporating effort levels and prompt changes into the mix so that you can actually stretch these models to their full degree.

1:18:51

I think we're quickly saturating a lot of that bottom half of tasks across models, and now you're starting to reach into the level of these things are doing real tasks, they're owning real areas, and you've got to push the models, both on the prompt side and the task side, to really see those gains. I love that. Alex, a pleasure to get to chat. Thank you so much for joining. Thanks for building the model. I'm excited to get to use it over the next... Thank you, thank you, and always feel free, anybody watching this, to tag me. I'm on X, Alex Albert with two t's or something, and I am happy to respond. Any feedback, just keep it coming my way. Thanks.

1:19:12

Awesome. Thanks, Alex. Bye. All right, folks, you heard it here from Alex. A couple things. Okay, here's my big takeaway. Okay, you tell me what your takeaways were. We've been testing this model for the last, I guess, almost two hours now, and our initial thoughts have been like, ah, it's not necessarily as good as Opus 4.6 at a lot of things. And what I took away from his chat with us is that's to be expected. The way that this model works is a bit different from 4.6. 4.6 filled in a lot of gaps, it read between the lines, and this model is a lot more literal and a lot more detail-oriented.

1:19:21

And the more you can explicitly tell it exactly what you want, the better it's going to get. And a good example from the chat that we just had is telling Opus 4.7, I'm going to bed. So that's one of those things that another model, an older version, an older model, would not know. Oh great, I have more time so I can spend more tokens. But apparently according to Alex, if you say I'm going to bed, it gives it license to use more tokens, use more tools, and keep going until it's actually done, whereas without that bit of detail it might assume, hey, I have less time so I'm going to do a smaller slice.

1:19:37

So there are probably all these little dials and whistles now in this new model that, if you know what to ask for, you might be able to unlock a lot of its power. But if you don't know what to ask for yet, and you're using an old 4.6 prompt, you should expect the performance to be worse. And again, we've not been testing this for a long time, so I'm sort of synthesizing what he said and some of our experience so far and telling you what to expect, but I think that we'll have a lot more in the coming days on this. Katie, anything else that I'm missing, that we took away?

1:20:04

I think I would just underscore the effort levels, and it really sounds like Alex is saying that mapping the task that you want the model to do to the effort level that it needs is something that's worth thinking about, and this is something like I

1:20:10

probably all these little dials and whistles now in this new model that, if you're, if you know what to ask for, you might be able to unlock a lot of its power. But if you don't know what to ask for yet and you're using an old 4.6 prompt, you should expect the performance to be worse. And again, we've not been testing this for a long time, so I'm taking, I'm synthesizing what he said and some of our experience so far and telling you what to expect. But I think that we'll have a lot more in the coming days on this. Katie, anything else that I'm missing that we took away?

1:20:13

I think I would just underscore the effort levels, and it really sounds like Alex is saying that mapping the task that you want the model to do to the effort level that it needs is something that's worth thinking about. This is something I think a lot about: learning to drive manual versus driving stick shift. That's the image that comes to mind for me for this. You're going to get better performance and more fine-grained control from switching the effort levels and selecting the effort level that matches your task than you would just having a one-size-fits-all effort level going at the task. But it's going to take some experimentation, I think, to know which tasks warrant which effort level.

1:20:16

Yeah, I agree. I totally agree. It's also a little bit more just that the writing is a bit more straight ahead, but maybe if we say, hey, I want you to write like Annie Dillard, I'd be curious what it might do. I guess your writing prompts have a lot of voice stuff in them, right?

1:20:27

So, yeah, I am. I did have it go over, I had it do a revision on the intro that I had it write with my tastemaker voice, which has clips from a lot of writers that I admire, like Gia Tolentino, who has more architectural writing patterns with dependent clauses and things like that, and then John Green's in the mix there too, and it did come out more fluid then, but once it had all of that context. I just pointed all of that context at it and said, this is how I want you to write. So I think that's another thing. The more bundles of information you have to give the model so you're not having to completely prompt from scratch, I think that's something else that we're going to see compounding returns from. It's just the more context you have to give the model, the more you can give it to go on to get the outcome that you want.

1:20:33

Really, really interesting. Okay, so if you are here, we are going to end this stream soon so we can get back to doing some real testing and some writing live checking, and you should subscribe to Every. Every is the only subscription you need to stay at the edge of AI. Every dot to. We do a lot of cool stuff here at Every, so when new models come out, we do day-of vibe checks. This is your day-of vibe check for this model. Normally, we get the models early. We're already testing some early models that I'm pretty excited about that we can't tell you about, but that will be dropping soon. Normally, we get them early. Today we're doing a live vibe check. That's why you're seeing us figure out this model in real time. But we drop a new newsletter every day with lots of little nuggets about what we're learning and what we're doing in AI. Today's, this is actually, I guess, from yesterday. Yesterday's newsletter, we talked about what might come after our LMS. We did a little vibe check on Claude managed agents, which we've started to put into production with Sparrow, which is our writing app. We did some philosophical discussions of the jagged frontier, so lots of good stuff every day from us on the newsletter side.

1:20:39

We also have a suite of apps that we build that help us write interesting stuff about AI. So we have Spyro, which is our AI writing assistant. We have Quora, which is our AI email assistant. We've got Sparkle, which we just launched a new version of two days ago. Sparkle is your agentic Mac cleaner. It's pretty sick. It makes my computer clean. We've got Monologue, which is a speech-to-text app like Whisper Flow. We've got Proof, the sloppy markdown editor that we've been trying to de-slopify with Opus. It's so much better now. Dan, it is a lot better. If you're using a coding agent, Proof is just such a good way to get your plan and any of your plan documents into a live collaborative document that you can share with other agents and with your friends and co-workers. We've also got Plus One, which is our hosted one-click open claw. So all of this, all the stuff that we write, all of the apps that we make, all of the training and courses that we do, it's all included for one price under the Every subscription. You should check it out, every.2 slash subscribe.

1:20:41

We will have a full vibe check tomorrow, and probably an even fuller one over the weekend or on Monday, once we have a few more days to really process what's going on. Until then, it is always a pleasure to get to vibe check with you. Thank you all for joining, and remember, stay hydrated. Katie, see you later. See you later. Bye. And, like, we're still workshopping the format. Like, you know, I would probably not want it to just be in paragraph format. I want more signposting. Like, what's happening? What does this mean? So this module is, like, still in progress in that way.

1:20:58

But, yeah, like, the idea of, you know, we have, like, the agent setup in Scout in Slack to Scout things like this. But the fact that I can just come to Slack and be, like, can you, or come to Claude and say, can you check the Slack for recent nuggets? And then it has the context in this project to know what constitutes a good nugget and I can go out and find them. Like, that's something that I'm pretty excited about. Okay. So if you were going to summarize all of what you've learned so far about using this model for the writing tasks that you're doing every day from, okay, I'm going to use it to gather research and harvest it. To I'm going to write stuff.

1:21:39

Like, how would you summarize it? What's your vibe check? My vibe check is that I think it does come down to, like, the prompting and giving it more clear instructions about, up front, about what good looks like. Like, just to sort of, you know, just to anticipate some of the failure modes that I'm starting to notice. Like, you know, like, I'd probably want to, you know, say, like, make sure that I know, like, like, tell it, you know, pick up this voice style from this skill that I've created. And just in general, give it more information to go on about what good looks like. Like, I feel, like, a little less confident in just saying, go write this.

1:22:18

Ironically, for what I was just telling you that we were talking, that, you know, the conversation was saying. But, like, for 4.7 so far, it seems like it needs more explicit direction. And, like, which is fine. But it does mean a little bit more cognitive load on me to actually think through that. I mean, I, honestly, I usually work very iteratively. And so, like, this isn't working. Like, I'll go through, like, 20 different drafts on a paragraph until I get what I want. So maybe that isn't so much. And so maybe that's, like, not something that's, like, necessarily, that's not going to go down with this model. But I don't know that it's necessarily going up either.

1:22:59

But that's a place where, like, the speed becomes an issue. Because the faster it goes, the faster I can run through options. And so, okay. So let me just try to summarize that. So, so far, what we've noticed, and I think this applies beyond writing to also decoding, is that this model is maybe a bit more literal. So it's good at following instructions. But it's not as good at reading in between the lines. So it performs better if you give it specific things, more specific information about exactly what it looks like and what you want it to do. And it's probably a little bit worse at vaguer prompts. Does that feel right? Yeah.

1:23:37

And so if you're, if we're going to do, like, a live reach test. And the reach test is always, I think the best way to vibe check a model is to think to yourself, when I go to do my work in an hour, what am I going to reach for? Am I going to keep the dial on 4.7? Am I going to move it back to 4.6? If you're, if you're reaching, what's, where's your reach test at right now for this model with writing? Am I going to move it back to 4.6? My reach test right now is default is falling back to 4.6. It's going to give me the kind of writing that I want for now. Fascinating.

1:24:36

So a couple of the other tests that we've done, we've done some tests of this model on what, what I've been calling the vibe slot benchmark, which is basically, can you take a model? Can you take a vibe coded slot production code base and turn it into what a senior engineer would do? And so far, this model is quite good at identifying what a senior engineer would do, but it is not as good at going and actually carrying that plan through. It has a little bit of, a little bit less courage than it might need to. It gets a little bit distracted by what's currently there.

1:25:19

And so it has trouble having the courage to be like, no, no, no, this is how we're going to reorient this entire code base. And this is how we're going to, how we're going to rewrite it from first principles. It sort of then, even though it knows what it's supposed to do, it then sort of starts chipping away at like little side things and doesn't actually like get the core of the problem, which is a common issue for coding models. It's start, like some coding models can do this. It's starting to get a little bit better, but this also happens with UBD 5.4.

1:25:48

So it's, it's interesting to see that even though this model is so much better on a lot of the benchmarks in, in terms of this particular benchmark, like the sort of vibe slot benchmark, I would give it like a, a B, like B, B plus, something like that. Like it's, it's promising and that it may be part, part of the benchmark, which I think is actually really important is the benchmark is being conducted from the perspective of a vibe coder. So what I explicitly don't want to do is give it the answer, right? Like is to say, here's exactly how I want you to rewrite this code base.

1:26:25

I want you to infer, like find the answer yourself and then infer how you should execute on that answer. And what we're finding so far is that this model isn't quite as good at doing that just in general. And so it may be that the best way to get like release some of its powers, get some of its powers out is not a, not necessarily a vibe slot benchmark, but a vibe slot benchmark where instead of impersonating a vibe coder, you impersonate a coder who actually knows what they're doing, which would not be me. And it may be that it is a very good power tool for professional programmers, which is, again, this is now it's so interesting.

1:27:08

Like it's, we're starting to lean back into the lane that Codex used to occupy. Codex has gone very far in the direction of being good for vibe coding and being like very emotionally intelligent and all that kind of stuff. And it seems like Opus, which has always been known for that, is now moving into, it doesn't necessarily fill in all the gaps for you, but it's a power tool if you prompt it right. Again, we have not, we just got access to this model. Normally at every, we get access to models before they come out.

1:27:36

Today, this model is the, is the first one in a long time from Anthropic that we've not gotten early access to. I think they're, they're really rushing on it. But that is our, that is our early vibe check so far. If you have access to this model and you've been vibe checking it yourself, I would love for you to share in the chat.

1:27:58

What, what your vibe check is so far. I'm going to turn on the chat so you can, you can start sending, sending in messages. But yeah, please, please tell us what you think. The way that we do vibe checks is red, yellow, green, or gold. Red is trash release, yellow is it's fine, but I wouldn't reach for it. Green is, I'd probably use this every day for a certain task that I do. And gold is, it's a paradigm shift. So you let us know what you think. If you are here, you should also know that we are going to have Alex Albert, a researcher at Anthropic coming on the show to tell us about this model and how they trained it and what they're finding.

1:28:44

Maybe he'll give us some pointers about how to do this well. So, so stay tuned for that coming up in probably the next couple minutes, it seems like. And I am going to check in on our five slop benchmark again. Let's see, anything, anything new here?

1:29:08

Hmm. So I'm curious, just like for clarification, this is something I was curious about while you were talking. Like when you say it's not as courageous, does it tend to like say this is what should happen and not do it? Or is it taking actions? They're just like chicken, chicken, like actions around the edges of the problem. It is good at identifying the plan and like the, like here's what should, it should look like. And then when it is actually asked to do it, it doesn't, it chickens out. So it knows what to do. It just like doesn't want to actually do it, which is kind of, kind of interesting. And I'll say again, this is a common thing with coding models.

1:29:45

Part of it is a, a, the fact that it is looking at an existing code base. It's the, the existing code base like pollutes its context and distracts it and makes it more cautious because coding models have been told over and over again, like, you know, don't mess with the code base, like try to just make the little tweak. So they're really not meant for, I mean, so far, so far they have not been meant for this, but I think, I think this is a really important benchmark because it is becoming a common pattern in vibe coding.

1:30:20

So it's becoming a common pattern that the first thing to do is just like make the prototype that works, like find the value, go as quickly as you can do the sloppiest code you possibly can. And then once you've done that, what you'll end up with is a code base that like bears the scars of the history of like how you got here with all these different, you know, files and folders and like three different versions of the app and whatever. And what you need to do is as Annie Dillard says, cover your tracks. You got to cover your tracks. You got to do that in, in coding. You got to do that in writing.

1:30:52

And the best way to do that is to have a model where you can say, Hey, like this is a, this is a, this total vibe slop. If I was going to start again, how would I do that? And then just let it do that. Um, and some models are good at it. Most, most, most models are not good at it today, but I think that they will, because this is like, this is an important thing to be able to do is cover your tracks. Yeah. I had, we've got, we've got Alex Albert here. Alex. Hello. How are you doing? I'm doing great. Thanks for having me on. Amazing. Um, tell us how you, how, what your involvement is with this model. Tell us about the model. We're, we're super excited to get to test it.

1:31:32

Um, yeah. Share with us, uh, share with us some more. Yeah. Yeah, yeah, yeah. So, um, hello, Dan. I, Dan, we might've met briefly in the past before, but it's great to see you again. Katie, I don't think we met. Great to meet you. Um, thanks for having me on. Thanks. I am on the research side at Anthropix. I'm one of the many, many people that contributed in some small way to making this model. Um, Opus 4.7 has been, uh, definitely a model that we're very proud of. Um, in a lot of ways, me personally, I found that it's good at a bunch of things over Opus 4.6, especially in like my own personal coding usage, backing away on side projects, things of that nature.

1:32:12

Um, there's a bunch of stuff we can get into. Uh, I was kind of noticing before I jumped, uh, jumped on the call that you guys were talking a little bit about how it was doing on like some coding tasks and autonomous tasks. And I have a few tips we can share and try to test out. Um, yeah. Happy to take this wherever. Yeah. So, okay. So let me tell you what we found so far in it. And I will say like, this is super, super early. Like, obviously we've only had our hands on it for like 30 minutes and we're testing it as we are being watched by almost 10,000 people. So there's a, there's, uh, the vibe check is provisional, but some of the things we've

1:32:48

noticed are that it is a bit more literal than, uh, than Opus 4.6. And in some of those tasks that, that meant. So for example, one of the tasks that I've been doing it with, which is a coding specific one, I'd love your feedback on is what I've been calling the vibe slop benchmark, which is if you give it, if you give a model, a vibe coded production code base, and you say, I want you to rewrite this from first principles. Can it identify the like core invariance that it should, um, I, it should write the code base around and then can it refactor or rewrite the code base in a way that pulls through all those things.

1:33:27

And so far the answer is Opus 4.7 is good at writing the plan, but it's a little bit, it lacks a little bit of courage to actually go and like burn the ships and rewrite from first principles. But again, very, very, very, very preliminary. And, and, and it's one of those tasks that requires, uh, a lot of reading in between the lines, because I'm intentionally not saying here's exactly what I want you to do. So I'm curious, like when you think about, um, what you love about this model, what you use it for, how to get the best out of it. Like what, what are we missing? Yeah, I, I would say to that point that Opus 4.7 is a lot more on the mark, so it will

1:34:06

do what you tell it to do. Um, and this is kind of like an interesting behavior with cloud models and you can trace this all the way back. Like if you remember way back long ago, a year ago with like the, a son at 3.7 days, that model really loves to just go for it. Yeah. And you'd write a prompt and then all of a sudden it's changing like a ton of different files all around your code base. And you're like, okay, let's like slow, slow down. I just asked you to change a button color. Like why? Yeah. Let's slow down here. Uh, so then with son at four, we dialed that behavior back in and there was actually a lot

1:34:40

of similar feedback to what you might be experiencing now where it's like, okay, this model seems strong, but it's not doing all that I wanted it to do like 3.7. And that's kind of like this perpetual back and forth that we experienced sometimes with models as we're trying to course correct them in terms of this, like literalness or how much they just like go for it. Um, and I find that it takes a little bit of tweaking and tuning of your prompt because you get it used to one model. And then the next model comes, it's a little bit more dialed back and maybe it follows your instructions better, but it doesn't necessarily do things beyond what you asked for.

1:35:15

Uh, and then you need to add that back into your prompt. So we're trying to find the right balance there. I think actually it's, it's just always funny kind of going through these motions because I remember when Opus 4.6 came out, there were some of the complaint complaints of like this model's doing too much. It's thinking for too many tokens. It's like taking too many actions. So then we all shifted our prompts over and now I think we just have to kind of dial them back in again or add in those instructions to actually be thorough and take more actions. So give me some specific examples of what you think this is. It's amazing. Yeah.

1:35:47

Uh, one in particular that I love is just operating on tasks in the background. So I found that this model is a lot more coherent, um, when it's having to do these like long horizon tasks, kind of your like typical, you know, like meter, how long is it going for type of thing. Um, that works really, really well with a bunch of the new ship stuff we've been shipping recently, especially like, um, cloud code on desktop. I found that example of a specific type of like long running background tasks like that that you're, that you use. Yeah. So one that I just did recently, um, that is kind of, uh, interesting is I was trying to

1:36:21

build this like apartment hunting website for my girlfriend. She's trying to find a new spot. Um, and I found that just overnight I told the model like, Hey, I'm going to bed. I want you to just work autonomously and build this thing. And when I wake up, I want to see a like complete dashboard that's pulled all the latest departments to these specs. And it like shows different, um, uh, things in different areas that my girlfriend wants to look into around like Craigslist and Zillow and all this. And then it runs on a schedule. And so every, you know, twice a day it's pulling new things.

1:36:56

And I found that actually just like giving the, those instructions of like, Hey, I'm going to bed and like, I just want you to work and spend a bunch of tokens and like figure this out actually allowed me to come back to an output that was like more refined. And it did those iterative cycles. I looked back at the transcript and it, instead of like stopping at that first go round, it's like, okay, I have time here. I don't have that like stress of having to finish. I'm going to just keep working on this thing. Um, so I thought that was really cool. Um, I'm, I literally just told 4.7 I'm going to bed. Um, I think that's a really good little, little trick. Yeah.

1:37:35

See if it comes back to something better there. Um, another thing I love is on the vision side. So we expanded our, our support for bigger images, larger images, higher resolution. And one fun thing that we were kind of bashing on is just its ability to kind of play like where's Waldo type games. So you give it an image where it has a lot of like pixel, tiny little details and can it find like Waldo or can it find like the, uh, random bottle on a shelf of a ton of different things. Um, that's been really sweet to see. And that actually translates into a lot of things we care about as well.

1:38:11

Like, even though that's a toy example, there's a lot of cases where you're trying to get it to iterate on a front end and maybe like a button is slightly misplaced or misaligned, or there's like a tiny little detail that Opus 4.6 might've missed. Uh, Opus 4.7 will catch those. Um, so that's been great for a lot of our like visual design output work from front ends to things like slides as well. Does that translate? So it seems like it's good at recognizing like what's, what's going on in an image. Does that also translate to its ability to then build something that matches, uh, a file or figure out what it should build? Exactly. Yeah.

1:38:47

So, uh, a big thing we've been testing is like how well can it replicate like a front end or a design? Um, so we'll pass it in like a screenshot and then, uh, test it on its ability to like, uh, replicate that. And that's been pretty fun. Um, it, it is quite amazing how well these models can do things like that these days. It just like blows my mind. The like rate of progress, especially compared to like six months ago on these, on these tasks. Hmm. That's really cool. Um, anything else other than, so like long, long running tasks, images, what else should we try? Yeah. Um, one thing, especially that I've been focused on in my work is like kind of real world tasks.

1:39:25

So this model is really good at anything spreadsheet doc PowerPoint related. Um, I'm not sure how many spreadsheets or slides you guys are making these days, but yeah. Okay. Okay. Cool. One of the things that did best that is I gave it our P and L and I asked it to write our investor update just from the P and L. It did a really good job. Um, like I can show you, um, let me see how, where, where it is. I I'll, I'll, I'll pull it up and show it to you, but it's basically it like wrote effectively what I, what I wrote. And I, I think it had this, um, really nice, um, I'm opening the link. It had this really nice blend of it's good at the, doing the analysis and finding all

1:40:09

the numbers. And then because it's writing style is a little bit more straight ahead and it doesn't like make stuff up as much. Yeah. It didn't like make as many conceptual leaps as other models would. So like, this is, this is the investor update it wrote. And I literally just said, here's our P and L here's the investor update. And you know, it did all, it got all the numbers, right? Like all these numbers look approximately right. And it broke it out in a way that I might've broken it out and it, you know, did all of our different products and it like figured out what we're going to focus on. It's just like, it's, it was pretty good.

1:40:40

Um, something that's like saves, would have saved me a lot of time. Yeah. One thing we've heard in, in a lot of our, um, feedback with some of our customers is that it feels a lot more like a coworker in some ways. And that translates in kind of these non coding tasks as well, but also in the way that it sometimes does push back on you. And I've actually been kind of surprised over moments. And it's like, ah, actually this is, you know, maybe this is not the way to do it. And I'm like, oh, that's interesting that like this model is like pushing back on me like that. Um, I noticed that with a writing task I had to do, we have these different modules,

1:41:13

these different kinds of content. And I gave it some out, some like raw material and said, make this into this kind of module. Yeah. And it said, actually, it doesn't make sense. It's that kind of module. It needs X, Y, Z. And I was like, wow, that's actually useful feedback. Yeah. So I think that's like a personal preference on the behavior, but at least for me, I really appreciate that, that like I can trust the model will push back on me if I'm on the wrong path and, uh, not just like keep going along with what I'm saying. Um, so that's been another thing I've really enjoyed with the model. That's really interesting. So I want to go back to this background thing.

1:41:44

It seems like there's some new awareness of, oh, I have an amount of time and an amount of tokens within which to like do this. And if I feel like I have less time, then I'm going to, I'll do less of a thorough job. And if I feel like I have more time, I'll do more of a thorough job and spend more tokens and more time. Um, that seems really important and really valuable. How long is this generally running? Like, is there, is there some sort of, has the horizon moved in terms of like how long, you know, you're seeing Opus 4.7 run versus 4.6? Yeah. I, I think it's hard sometimes to like put it into a quantifiable term like that because

1:42:29

it's like the, the wall clock time can like depend on like a lot of different variables. Yeah. But, um, but usually it's, it's been longer. Um, I see it like taking more turns. Of course you do need to like specify that sort of instruction. So one thing that we have, have prompted it to do in some scenarios is like be very thorough. Don't mind making a lot of tool calls. Don't mind making a lot of changes to files and that sort of thing. And just incentivizing it to do that. Um, effort levels are also a big piece of this model. So we've put a ton of work into making sure that effort levels actually control token usage better.

1:43:25

Mm-hmm. We've got a lot of conceptual framing for this is like, max is like, just go gangbusters. like I don't care at all about your token usage I want you to just try for as long as you can with as many tokens you can as many tool calls that'd be like max so we run a lot of like our benchmarks at max effort extra high is where we actually settled our defaults for cloud code we found that this really gave a great experience especially as a lot of users are shifting more towards these asynchronous tasks we want to make sure that the model is working for longer high medium are really

1:44:02

sweet spots for kind of achieving that Goldilocks balance between token usage and trying to get a decent amount of effort into a prompt or a task and we found that that is really good for a lot of um things maybe in the like non-coding domains if you're doing like chat or things like that that's where you want to drop it down maybe to that level of effort um but everything agentic I would stay higher in that range got it so I think like you know probably a thousand people have joined since you got here and I want to respect your time so if you have to go at any point just like I do you need

1:44:35

to drop it launch days are pretty busy in a second but if you have a last question I'm happy to answer it uh okay well I last question any uh like any parting any parting thoughts for us any words of wisdom as we are going forward you know we know you've talked about background tasks and images and effort levels um what else any any last thing that you're excited about that we should know about this model um yeah I mean I would say that the the rate of improvements at two things one is the rate of improvements are not slowing down anytime soon I think that's always important to emphasize like

1:45:06

um the frontier of what's available is continuing to grow and expand um the second thing is I think like there is this paradigm shift happening from this collaborative work to asynchronous work and this is important for both users to know but also from people building on these models like there's a mindset shift to get around in terms of like becoming more comfortable with exposing products that allow for this sort of behavior for making sure that you're incorporating effort levels and prompt changes into the mix um so that you can actually stretch these models to their like full degree I think we're

1:45:41

like quickly saturating a lot of that like bottom half of tasks across models and now you're starting to reach into the level of like these things are doing real tasks they're owning real areas and you got to push the models um both on the prompt side and the task side to really see those gains I love that Alex a pleasure to get to chat thank you so much for joining thanks for building the model I'm excited to get to use it over the next thank you thank you and always feel free anybody watching this to uh tag me um I'm on on x uh Alex Albert with two t's or something um and I am happy to

1:46:14

respond any feedback just keep it coming my way thanks awesome thanks Alex bye all right folks you heard it here from Alex a couple things okay here's here's my big takeaway okay you tell me what what your takeaways were we've been testing this model for the last I guess almost two hours now and our initial thoughts have been like ah like it's not necessarily as good as at Opus 4.6 at a lot of things and what I took away from his um uh his chat with us is that's to be expected the way that this model works is a bit different from 4.6 4.6 uh it filled in a lot of gaps it like read between the lines

1:46:58

and this model uh is a lot more literal and a lot more detail oriented and the more um and the more you can um interesting someone's someone thinks that my microphone sounds like I'm on the toilet um I don't I think you sound fine Dan great perfect um because we've been having some microphone issues anyway um so the more so you should expect all your old old prompts to break and the more you can push this model to more explicitly tell it exactly what you want the better it's gonna the better it's gonna get and what I would um a good example from the uh from the chat that we just had is telling Opus 4.7 I'm going to bed so that's one of those things that

1:47:53

another model like an older version an older model would not know oh great I have more time so I can spend more tokens but apparently according to Alex if you say I'm going to bed it gives it license to use more tokens use more tools and keep going until it's actually done uh whereas without that bit of detail it might assume hey like I have I have less time so I'm going to do a smaller slice so there are probably all these little like dials and whistles now in this new model that if you're if you know what to ask for you might be able to unlock a lot of its power but if you don't know what to ask for yet

1:48:31

and you're using an old 4.6 prompt you should expect the performance to be worse um and again we've not been testing this for a long time so I'm I'm sort of take I'm sort of synthesizing what he said and some of our experience so far and telling you what to expect but I I think that we'll have a lot more in the next in the coming days on this Katie anything anything else that that I'm missing that uh that we took away I think I would just like underscore like the effort levels and it really sounds like Alex is saying that like like mapping the task that you want the model to do to like the

1:49:05

effort level that it needs is something that's worth thinking about like and this is something like I think a lot about um learning to drive manual versus driving stick shift that's kind of the the image that comes to mind for me for this is like you're going to get better performance um you know and more like fine-grained control from like you know switching the effort levels and like selecting the effort level that matches your task then you would just like having a one per one size fits all effort level just going at the task but it's going to take some experimentation um I think to to know

1:49:36

which kind of tasks warrant which kind of effort level yeah I agree I totally agree um it's also a little bit more just like the writing is a bit more straight ahead but maybe if we say hey like I want you to write like Annie Dillard I'd be curious what it what it might do I guess you're I guess your writing prompts have a lot of have a lot of voice stuff in them right so yeah I am I did have it go over like I had it do a revision on the intro that I had it right with my tastemaker voice which has clips from a lot of writers that I admire like Gia Tolentino who has more sort of like architectural um writing

1:50:14

patterns with like dependent clauses and you know and things like that and then like John Green's in the mix there too and like it came it did come out more like fluid then but like once it had all of that context I just pointed all of that context at it and said this is how I want you to write um so I think that's another thing yeah like the more bundles of information you have to give the model so you're not having to completely prompt from scratch I think that's not something else that we're gonna see like like you know compounding returns from it's just the more context you have to

1:50:45

give the model the more you can give it to go on to get the outcome that you want really really interesting okay so if you are here we are going to end this stream soon so we can get back to doing some real uh testing and some writing live checking um and uh you should subscribe to every every is the only subscription you need to stay at the edge of AI every dot to we do a lot of cool stuff here at every so when new models come out we do day of vibe checks this is your day of vibe check for this model normally we get the models early we're already testing some some early models that I'm pretty

1:51:23

excited about that we can't tell you about but that will be dropping soon normally get them early today we're doing a live vibe check that's why you're seeing us figure out this model in real time but we drop a new newsletter every day with lots of little nuggets about what we're learning and what we're doing in AI today's this is actually I guess from yesterday yesterday's newsletter we talked about what might come after our LMS we did a little vibe check on Claude managed agents which we've started to put into production with Sparrow which is our writing app we you know did some did some philosophical

1:52:00

discussions of the jagged frontier so like lots of good stuff every day from us on the newsletter side we also have a suite of apps that we build that help us write interesting stuff about AI so we have Spyro which is our AI writing assistant we have Quora which is our AI email assistant we've got Sparkle which we just launched a new version of like two days ago uh Sparkle is your agentic Mac cleaner um it's pretty sick it makes my computer clean we've got monologue which is a speech to text app like whisper flow we've got proof the the sloppy uh uh markdown editor that uh that we've been trying

1:52:38

to de-slopify with with with Opus it's so much better now Dan it is a lot better if you're using a coding agent um proof is just such a good way to get your plan um uh and any of your plan documents into a like live collaborative document that you can share with other agents and with your friends and co-workers we've also got plus one which is our hosted uh one click open claw so all of this all the stuff that we write all of the apps that we make all of the training and courses that we do it's all included for one price under the every subscription you should check it out every.2

1:53:13

slash subscribe uh we will have a full vibe check tomorrow um and probably an even fuller one uh over the weekend or on Monday once we have a little a few more days to really process what's going on until then it is always a pleasure to get to vibe check with you um thank you all for joining uh uh and remember stay hydrated Katie see you later see you later bye

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note