Why Opus 4.8 Pulled Me Back to Claude
Description
Opus 4.8 dropped today, and it's so good Anthropic could have called it Opus 5. Every CEO Dan Shipper has been testing Opus 4.8 with the team for the past week. In this day-zero vibe check, he breaks down why Anthropic is so back—and the one thing keeping him from going all-in on Claude. Dan covers: 1. How Opus 4.8 jumped 30 points past Opus 4.7 on our Senior Engineer benchmark—and edges out GPT-5.5 2. Why it's the best writing model we've tested, with fewer AI tells on high effort 3. The slide deck that made it the first model to nail one-shot knowledge work 4. Why Kieran Klaassen calls it the most human model he's used 5. The two catches: it's heavily reasoning-sensitive, and the Claude app is still messy compared to Codex Read the full vibe check on Every: https://every.to/vibe-check/opus-4-8-vibecheck See the slide deck Opus 4.8 made https://docs.google.com/presentation/d/1jGL0OBNeTh-k0rp4I-gQJxQcEWPB4Y6U/edit?slide=id.p3#slide=id.p3 Every is the only subscription you need to stay at the edge of AI. Start your free trial today: https://every.to/subscribe
Summary
Generated by claude-sonnet-4-530-second take
Opus 4.8 is legitimately paradigm-shifting for serious Claude users—scoring 63 on Every's senior engineer benchmark (vs 62 for o3-mini, 30 points above Opus 4.7), 79.6/100 on writing benchmarks (vs 73 for o3-mini), and reversing Claude's recent momentum loss to OpenAI. The speaker and his team tested it for a week and found it's the best writing model they've seen, emotionally intelligent for interpersonal work, and excellent at knowledge tasks like slide decks. Critical caveat: performance is extremely reasoning-level dependent (extra high is vastly better than medium), and the Claude desktop app's messy UI remains inferior to Codex's harness, limiting daily-driver adoption despite the model's quality. This is Anthropic clawing back mindshare after months of losing ground.
Key takes
- Opus 4.8 matches or beats o3-mini on real-world tasks: Scores 63 vs 62 on senior engineer benchmark (refactoring vibe-coded codebases), and human engineers score 80s-90s, meaning these models are approaching senior-level capability on structured rewrites.
- Writing quality jumped significantly: 79.6/100 vs 73/100 for o3-mini, with notably better voice matching, fewer AI tells, and more expressive output—first time a model consistently produces "depth" in auto-generated slide decks instead of thin slop.
- Reasoning levels dramatically affect output: Extra high reasoning performs substantially better than high/medium for both coding and writing; the model is "very sensitive" to this setting, meaning casual users on default settings will miss its best performance.
- Harness quality now rivals model quality: Despite Opus 4.8's superiority, the speaker still uses Codex daily because Claude's desktop app has fragmented tabs (chat/code/cowork) that feel like "shipping their org chart," while Codex is fast, simple, with a working in-app browser. The app friction prevents full adoption of a better model.
- Emotional intelligence and frame-challenging stand out: Best model tested for interpersonal/psychology work; thinking traces show it explores multiple permutations of situations and actively pushes back on user framing rather than just agreeing—useful for management, friendship, internal decision-making.
- Claude regained "die-hard Claude stan" loyalty: Even team members who'd switched to o3-mini for writing/coding are flipping back; one tester (Kieran, GM of Quora, runs 50 agents daily) rated it "paradigm shift"—a rare grade—and called it "the most human model" he's used.
Useful details
- Senior engineer benchmark methodology: Gives model vibe-coded slop, asks for first-principles rewrite, compares to two human senior engineers' rewrites. Opus 4.8: 63, o3-mini: 62, Opus 4.7: ~33, human engineers: 80s-90s.
- LFG bench example (3D cozy island scene): Opus 4.8 output was "lush and vibrant" with rich detail; o3-mini had more diversity of elements but felt "less lush," more straightforward task-completion.
- Reach test ratings from three users: Speaker (CEO doing coding/writing/decisions): gold/green (harness drags it down), Kieran (coding all day, 50 agents): straight gold paradigm shift, Katie (senior staff writer): green. All historically Claude fans.
- Writing benchmark tasks: Introduction writing, promo emails, mid-article paragraphs, voice continuation from single paragraph sample.
- Slide deck test: Asked to make beginner deck on "compound engineering" philosophy—first auto-generated deck with actual depth, well-styled, not thin.
- Claude app problems: Three tabs (chat/code/cowork) run by different teams, confusing navigation, lacks Codex's in-app browser quality.
- Every team: ~30 people, applied AI lab for future of work, publishes day-zero vibe checks and detailed written versions at every.to.
Caveats / counterpoints
- Not the speaker's daily driver: Despite model superiority, he still uses Codex most of the time due to app quality—"harness matters as much as the model."
- Performance craters on lower reasoning settings: Medium reasoning is noticeably worse; users must manually select high or extra high for best results, which adds friction and isn't obvious to new users.
- Limited testing window: Only one week of internal testing; longer-term edge cases and failure modes may emerge.
- Writing benchmark is internal/proprietary: No public benchmark validation; scores are Every's own evaluation framework.
- No mention of cost, speed, or API availability: Comparison focuses on quality but ignores practical deployment considerations.
- The transcript provides no counterarguments from Anthropic, OpenAI fans, or skeptics about whether these improvements matter at scale or in production.
Ken relevance
High relevance for AI ops and content workflows. If Ken's systems involve writing (newsletters, content generation), this is the new quality bar—79.6/100 suggests meaningfully better output that might reduce editing overhead. The emotional intelligence angle is interesting for agent systems handling customer support or interpersonal coordination tasks. The harness/app quality issue reinforces Ken's likely belief that UI/UX around models matters enormously for adoption—even a better model loses if its interface sucks. For investing, this signals Anthropic is competitive again after losing ground; the "paradigm shift" rating from experienced users suggests sticky retention if they fix the app. The reasoning-level sensitivity is a product design flaw that limits mass adoption but might not matter for power users in Ken's context. If Ken's building on Claude APIs, this is a meaningful quality upgrade; if building UI layers, the Claude app's fragmentation is a cautionary tale about org structure leaking into product.
Watch verdict
Skim. The takes are valuable (reasoning sensitivity, harness quality gap, writing benchmark scores), but the video is repetitive (second half largely restates the first) and spends significant time on Every's business pitch. Read the written vibe check at every.to instead for the same insights in less time, unless you want to see the 3D island visual comparison or hear the speaker's enthusiasm firsthand.
Transcript
It's model release day! Opus 4.8 drops today. But honestly, they could have called it Opus 5 because this is a really great model. Anthropic, I know you're trying to under-promise, but you are over-delivering. We have been testing it internally for about a week here at Every, and here is your day zero vibe check. Before we get into it, what is Every? Every is the only subscription you need to stay at the edge of AI. You can think of us like an applied AI lab for the future of work. We're about 30 people. We're all early adopters of these tools. We write about all the new models, all the ways that we use them for coding, writing, design, company building, and more. We have a suite of products that we build for ourselves to help us work better with AI, and we also do a lot of training and courses. It's all available for one subscription on the Every website, every.to. And if you want to read the in-depth written version of this video, we also publish a written vibe check on Every as soon as the model drops. So make sure you go there to subscribe and read it. Okay, let's get into it. So headline is to me, Anthropic is back. Opus 4.7 wasn't that great of a model. Yes, benchmark improvements, but not that usable, pretty slow, hard to love. And what I found for myself is I was using only Claude and GPT 5.5 for almost everything for the last month or two. And even internally at Every, we have a lot of die hard Claude stands. We've been going hard with Claude for the last year or so. And you could feel even the die hard Claude stands were thinking Claude is pretty good, but I'm actually starting to use GPT 5.5 for some of my writing in ways that I'd never had before or some of my coding. I think especially with the Claude desktop app being just so clean and fast and feeling like the future versus Opus 4.7 itself being a slow, hard to use model, and then its harness, the Claude desktop app. It's just messy. Anyway, vibes were honestly bad for Anthropic for a little while. And this model is just a legitimately great model. It is the top of the pack for us in terms of our benchmarks. So we have a senior engineer benchmark, which measures these models on how well they do at senior engineer tasks. Opus 4.8 scores a 63 on the benchmark, which is about 30 points higher than Opus 4.7 and is just a hair. It's one point higher than GPT 5.5. So it's very similar to GPT 5.5, maybe a little bit more depending on how you score the benchmark on senior engineer tasks. It's really good for writing. It's expressive. It doesn't have a lot of AI tells, especially on the higher reasoning settings. And we'll get into that in a bit. And it's really good at knowledge work. It did this slide deck. One of the things we test always is how well does it do on knowledge work tasks? And one of those tasks is how well does it make a presentation? It made a slide deck. We'll link to it in the YouTube and maybe I'll be able to put a little screen share up here, but it made a slide deck explaining a topic. Our engineering philosophy is compound engineering. So it just made a beginner slide deck for compound engineering. And it's really good. A lot of these decks, when they're automatically generated, feel thin, and this had depth. Everything was pretty well styled. It was just a great first pass at a deck, which is the first time I've really seen that. And it's actually really hard to make a model that improves that much on all these different dimensions at once. I think we usually see the labs with a pendulum swinging back and forth. It's like one release is too cautious, and the next release it's way too, it just goes off and does tons of stuff without you asking. This model they just seem to have gotten something really right. It just feels good. Kieran Klassen, who's the GM of Quora and was one of the internal Every testers on this, said to him it's the most human model that he's worked with. That's why I think they could have called this Opus 5 and we would have been happy. There are some catches, though. There are some things this model doesn't do well. The first one is it's very sensitive to reasoning. We got really great performance on extra high reasoning, both for writing and for coding and less good performance on high and medium. So as you're testing it, especially on your hardest programming challenges and for really important writing, I highly recommend the high and extra high settings. It makes a big difference. The second thing is it is still not really my daily driver. And that's only because the Claude app is so much better than the Claude app. We're entering this world where the harness matters as much as the model does. And the Claude app has the scars of the history of how Anthropic got here. It's got the chat tab and the code tab and the cowork tab and each tab is shipped by a different team and you can feel it. Whenever I get in there, I don't know which tab to go into. And Claude is so fast and simple. And it just works really well. And it has a couple other bells and whistles like the in-app browser working really well, which changes the game for knowledge work. So I'm still in Claude all day. However, I'm now flipping back and forth between the Claude app and the Claude app in a way that I was not just because this model is so good. So let's get into some of the details. Okay, so first thing that we always do is a reach test and the reach test is our simplest measure of how good a model is, which is do you reach for it? And if you do reach for it, in what situations do you reach for it? We have three reach test participants today. We have three reach test ratings on this model from our team. One's from me, one's from Kieran Klassen, who I mentioned earlier, the GM of Quora, one's from Katie Parrott, who's a senior staff writer. On the reach test, this is clearly an S tier paradigm shifting model. So I'm gold, but I'm also gold slash green. And that's just because the harness isn't that good. Kieran is a straight gold paradigm shift, which is very rare. I gotta say it's very rare to give a paradigm shift grade to a model. So pay attention to this. Kieran's a gold and Katie is a green. So if you're doing the kind of work that we're doing, so I'm doing a lot of CEO work, which is work across all sorts of different things like coding and writing and decision-making, it's a paradigm shift model wrapped in a pretty good, okay-ish to pretty good harness. So that makes it paradigm shift to green. For someone like Kieran, who's coding all day and running 50 agents at once, paradigm shift. Kieran is also historically the biggest Claude stan on the team. So if you are a Claude person, you're going to love this model. And Katie is a green. Katie is using it mostly for writing and knowledge work, also historically a big Claude fan, especially for writing. I think it's currently going back and forth between this model and Claude for most of her work. Now let's get into some of the more detailed benchmarks. Coding, it is a powerhouse at extra high reasoning. It got a 63 on our senior engineer benchmark. Opus 4.7 is lower, and it's just a nose higher than GPT 5.5. The senior engineer benchmark, what it does is it gives the model a vibe coded code base. It says, this is vibe coded slop. Can you please rewrite it from first principles? And I actually have two human engineers who have done the rewrite themselves. So I can compare what the models do to what a human senior engineer would do. And these human senior engineers usually score in the eighties or nineties. Opus 4.8 is at a 63. GPT 5.5 is at a 62. It's really good. It's very close on a task like this. We also tested it on LFG bench, which Kieran has, which tested it on some real world style coding tasks, building a SaaS, building an e-commerce website or building a 3D game landscape. And it's also very good at that. Kieran found that it writes very readable code. The output bridges the gap between being a very good coder and being creative. And there's something that comes up a little bit more and you can see. So one of the tasks on the LFG bench is what Kieran calls a cozy island benchmark. It just says, can you make a 3D cozy island scene, a 3D game? Here's Opus 4.8. It's pretty good. It's rich. It's detailed. There's a lot of stuff going on. And here's 5.5. There is still a lot of detail in 5.5. There's, I feel a lot more diversity of things happening, but it feels less lush and vibrant. So Opus 4.8 just has a little bit of that. I find it has depth and character in a way that 5.5 feels a little bit more straightforward, just knocking out tasks for you. We also tested it on writing. According to our internal writing benchmarks, it is the best writing model that we've tested. We tested it on things like writing an introduction to an article, writing a promo email, writing a paragraph in the middle of a piece, all that stuff. Opus 4.8 scores a 79.6 out of a hundred on the writing benchmark. GPT 5.5 is 73. It's fast. It's expressive, especially on high. It doesn't have as many of the AI tells. It gets a little bit worse on medium. So be careful of that. But it's very good at figuring out a writer's voice from the context, which is very impressive. If you give it a paragraph of your own voice and say continue, it'll keep going in a way that feels very much like you. Adjacent to writing is using it for personal stuff. I used it for some internal psychology, interpersonal stuff. It's actually very good at that. It's very emotionally intelligent. I find that of the models I've tested, it's the best at pushing my own frame. I wouldn't call it disagreeable exactly, but when you look at the model's thinking traces, it's very thoughtful. It's going through all the different permutations of a particular situation and then helping to expand how you might think about it, which I think is so interesting to watch. And it's so useful for me as someone who wants to use it for management stuff or friendship stuff or any kind of interpersonal stuff. It's really good at that. Last thing is knowledge work. And again, we found it has unmistakable daily driver smell. It's great at building slide decks. It's versatile. You can move between coding and writing in the same thread and it all just works really well. Again, the problem is just in the Claude desktop app, it's not quite getting the most out of it as you could. I really hope they make that better. But in general, this is a banger model. If you are a Claude stan, you're going to love it. If you've been converted to Claude, I highly recommend you at least add it as part of your arsenal. I think it'll be an eye-opening experience for some of the things that you can do. And if you want more of this, we have a lot more on Every. We have an in-depth vibe check from the team on every.to. We will also have a bunch of updates rolling as we get more and more testing in over the next couple of weeks. I'm super excited about this model. Go check it out on the Claude desktop app, go check it out in Claude code and remember to stay hydrated. Opus 5 because this is a really great model. Anthropic, I know you're trying to under-promise, but you are over-delivering. We have been testing it internally for about a week here at Every, and here is your day zero vibe check. Before we get into it, what is Every? Every is the only subscription you need to stay at the edge of AI? You can kind of think of us like an applied AI lab for the future of work. We're about 30 people. We're all early adopters of these tools. We write about all the new models, all the ways that we use them for coding, writing, design, company building, and more. We have a suite of products that we build for ourselves to help us work better with AI, and we also do a lot of training and courses. It's all available for one subscription on the Every website, every.to. And if you want to read the in-depth written version of this video, we also publish a written vibe check on Every as soon as the model drops. So make sure you go there to subscribe and read it. Okay, let's get into it. So headline is to me, Anthropic is back. Opus 4.7 wasn't that great of a model. Yes, benchmark improvements, but not that usable, pretty slow, hard to love. And what I found for myself is I was pretty much using only codecs in GPT 5.5 for almost everything for the last month or two. And even internally at Every, we have a lot of like die hard Claude stands. We've been really going hard with Claude for like the last year or so. And you could kind of feel even the die hard Claude stands were like, you know, codecs is pretty good. I'm actually starting to use GPT 5.5 for some of my writing in ways that I'd never had before or some of my coding and stuff like that. I think especially with the codecs desktop app being just so clean and fast and feeling like the future versus Opus 4.7 itself being a slow, hard to use model, and then its harness, like the Claude desktop app. It's just, it's kind of messy. Anyway, fives were honestly bad for Anthropic for a little while. And this model is just like a legitimately great model. It is the top of the pack for us in terms of our benchmarks. So we have a senior engineer benchmark, which measures these models on how well they do at senior engineer like tasks. Opus 4.8 scores a 63 on the benchmark, which is about 30 points higher than Opus 4.7 and is just a hair. It's one point higher than GPT 5.5. So it's, it's about, it's, it's very similar to GPT 5.5, maybe a little bit more depending on how you score the benchmark on senior engineer like tasks. It's really good for writing. It's expressive. It doesn't have a lot of AI tells, especially on the higher reasoning settings. And we'll get into that in a bit. And it's really good at knowledge work. Like it did this slide deck. One of the, one of the things we test always is how well does it do on knowledge work tasks? And one of those tasks is how did, how well does it make a presentation? Like it made a slide deck. We'll link to it in the YouTube and maybe we'll, I'll be able to put a little screen share up here, but it made a slide deck explaining a topic. Our engineering philosophy is compound engineering. So it just made a beginner slide deck for compound engineering. And it's really good. Like I think a lot of these decks, when they're automatically generated, they feel kind of thin and this just had, it had depth. It had, it was, everything was like pretty well styled. It was just like a great first pass at a deck, which is the first time I've really seen that. And, and it's just, it's actually just really hard to make a model that improves that much on all these different dimensions at once. I think like we usually see the labs, like pendulum swinging back and forth. It's like one release, it's too cautious. And the next reason it's like way too, it just goes off and does tons of stuff without you asking this model. They just seem to have gotten something really right. It just feels good. Um, Kieran Klassen, who's the GM of Quora and was one of the internal every testers on this said it to him. It's like, it's like the most human model that he's worked with. That's why I think they could have called this Opus five and, and we would have been happy. There are some catches though. There are some things this model doesn't, doesn't do well. The first one is it's very sensitive to reasoning. We got really great performance on extra high reasoning, both for writing and for coding and less good performance on high and medium. So as you're testing it, especially on your hardest programming challenges and for really important writing, I highly recommend the high and extra high settings. It makes it, it makes a big difference. The second thing is it is still not really my daily driver. And that's only because the codex app is just so much better than the cloud app. We're entering this world where the harness matters as much as the model does. And the cloud app has the scars of the history of how claw of how anthropic got here. You know, it's got the chat tab and the code tab and the cowork tab and each tab is like they're kind of shipping their org chart. Like each tab is run by a different team and you can just kind of feel it. Whenever I get in there, I'm like, I don't know which tab to go into. And codex is so just fast and simple. And it just works really well. And it has a couple other bells and whistles like the in-app browser working really well. That is just changes the game for knowledge work. So I'm still in codex all day. However, I'm now flipping back and forth between the cloud app and the codex app in a way that I was not just because this model is so good. So let's get into some of the details. Okay. So first thing that we always do is a reach test and the reach test is, is, is our, is our simplest measure of how good a model is, which is, do you reach for it? And if you do reach for it, what do you, in what situations do you reach for it? We have three reach test participants today. We have three reach test ratings on this model from our team. One's from me, one's from Kieran Klassen, who I mentioned earlier, the GM of Cora, one's from Katie Parrott, who's, who's a, who's a senior staff writer. On the reach test, this is clearly like an S tier paradigm shifting model. So I'm gold, but I'm also like gold slash green. And that's just because the harness isn't that good. Kieran is a straight gold paradigm shift, which is very rare. It's, I gotta say, it's very rare to give a paradigm shift grade to a model. So I would pay attention to this. Kieran's a gold and Katie is a green. So if you're doing the kind of work that we're doing, so I'm doing a lot of CEO work, which is work across all sorts of different things like coding and writing and decision-making. It's a paradigm shift model, like wrapped in a kind of like pretty good, okay-ish to pretty good harness. So that makes it paradigm shift to green. For someone like Kieran, who's coding all day and just like running 50 agents at once paradigm shift. Kieran is also historically the biggest Claude stan on the team. So if you're, if you are a Claude person, you're going to love this model. And Katie is a green, Katie is using it mostly for writing and knowledge work is also historically a big Claude fan, especially for writing. I think it's currently kind of going back and forth between this model and codex for most of her work. Now let's get into some of the more detailed benchmarks. Coding, it is a powerhouse at extra high reasoning. It got a 63 on our senior engineer benchmark, opus 4.7. And it's just a nose higher than GPT 5.5. The senior engineer benchmark, what it does is it gives the model a vibe coded code base. It says, this is vibe coded slop. Can you please rewrite it from first principles? And I actually have two human engineers who have done the rewrite themselves. So I can compare what the models do to what a human senior engineer would do. And these human senior engineers usually score in the eighties or nineties. Opus 4.8 is at a 63. GPT 5.5 is at a 62. 62. It's really good. It's like, it's very close on a task like this. We also tested it on Kieran has a benchmark called LFG bench, which tested it on some real world style coding tasks, like building a SaaS, building a e-commerce website or building a sort of 3d game landscape. And it's also very good at that. Kieran found that it writes very readable code. The output is kind of bridges the gap between being a very good coder and being kind of creative. And there's, that's something that comes up a little bit more and you can see. So one of the tasks on the LFG bench is this, what Kieran calls a cozy island benchmark. It just says, Hey, like, can you make a 3d cozy island scene? Like sort of like a 3d game. Here's Opus 4.8. It's pretty good. It's rich. It's detailed. There's, there's a lot of stuff going on. And here's 5.5. There is still in 5.5. There is still like a lot of detail. There's a, I feel like there's more diversity of things happening, but it feels less lush and vibrant. So Opus 4.8 just has a little bit of that. I find that it has depth and character in a way that 5.5 feels a little bit more straightforward. I'm just going to knock out tasks for you. We also tested it on writing. It is according to our internal writing benchmarks. It is the best writing model that we've tested. We tested it on things like writing an introduction to an article, writing a promo email, writing a paragraph in the middle of a piece, all that kind of stuff. Opus 4.8 scores a 79.6 out of a hundred on the writing benchmark. GPT 5.5 is 73. It's just fast. It's expressive, especially on high. It doesn't have as many of the AI tells. It gets a little bit worse on medium. So be careful of that. But it's, it's very good at figuring out a writer's voice from the context, which is it's very, it's really impressive. If you give it, if you give it a paragraph of your own voice and say, continue, it'll keep going in a way that feels very much like you. Sort of adjacent to writing is using it for like personal stuff. I use it for, I used it for like some internal psychology, interpersonal stuff. It's actually very good at that. It's very emotionally intelligent. I find that of the models I've tested, it's the best at pushing my own frame. I don't, I wouldn't call it disagreeable exactly, but when you look at the models thinking traces, it just, it's very thoughtful. It's really going through all the different permutations of a particular situation and then helping to expand how you might think about it, which I think is just so interesting to watch. And it's so useful for me as someone who wants to use it for that kind of stuff, like for management stuff or friendship stuff or any kind of interpersonal stuff. It's really, really good at that. Last thing is knowledge work. And again, we found it, it has unmistakable daily driver smell. It's great at building slide decks. It's like versatile. You can move between coding and writing in the same thread and it all just sort of works really well. Again, the problem is just in the cloud desktop app, it's just not quite, you're not quite getting the most out of it as you could. I really hope that they make that better. But in general, this is a banger model. If you are a cloud stand, you're going to love it. If you've been converted to codex, I highly recommend you at least add it as part of your arsenal. I think it'll, it'll be an eyeopening experience for some of the things that you can do. And if you want more of this, we have a lot more on every, we have an in-depth vibe check from the team on every.to. We will also have a bunch of updates rolling as we get more and more testing in over the next couple of weeks. I'm super excited about this model. Go check it out on the cloud desktop app, go check it out in cloud code and remember stay hydrated.