We Tested GPT-5.5 for 3 Weeks. It's a Beast.
Description
OpenAI just dropped GPT-5.5—and after three weeks of hands-on testing at Every, the headline is its coding ability. On Every's Senior Engineer Benchmark, GPT-5.5 scored 62.5 out of 100. That’s about a 30-point leap over Claude Opus 4.7. Human senior engineers score in the 80s and 90s, so it's not all the way there yet, but it's the first model that's really closing the gap. Our team has been running GPT-5.5 on everything: coding, writing, knowledge work, and Open Claw. A few things stood out: 1. GPT-5.5 rewrites code from first principles instead of patching existing messes 2. It's an absolute beast when paired with a contract-style plan written by Claude Opus 4.7 4. It's a surprisingly restrained, effective business writer and voice mimic 5. Opus 4.7 still gets more trust for sharp insight and evaluation, and still wins on vibe coding from underspecified prompts Every is the only subscription you need to stay at the edge of AI. Subscribe today: https://every.to/subscribe Read the full vibe check: https://every.to/p/gpt-5-5 00:00 It's Model Release Day! 02:04 Vibe Check Roadmap 02:45 The Senior Engineer Benchmark, Explained 04:04 Why 5.5 Executes Better 05:48 Opus 4.7 Plan Secret Sauce 07:00 Real World Build Tests 07:53 Where 5.5 Falls Short 08:50 Language Strengths and Weaknesses 09:08 Writing Quality and Voice 09:50 Knowledge Work and the Codex Desktop App 11:00 Daily Driver Verdict
Summary
Generated by claude-haiku-4-5-20251001GPT-5.5 Review: 3 Weeks of Testing Summary
Main Topics
- Coding Performance - Senior Engineer Benchmark testing and model comparison
- Writing Capabilities - Business writing and voice replication assessment
- Knowledge Work - Desktop agent experience and agentic tasks
- Model Workflow Optimization - How to get the best results from GPT-5.5
Key Points
Coding Performance
- Senior Engineer Benchmark Score: GPT-5.5 scored 62.5/100 (human senior engineers: 80-90/100)
- Comparison Gap: 30-point improvement over Opus 4.7 (which scored ~33)
- Critical Finding: Performance dramatically improves when using a plan written by Opus 4.7 (mid-40s to 62.5)
- Key Strength: Ability to identify core principles and execute comprehensively without getting distracted by existing code
- Language Performance: Excels in TypeScript and Swift; weaker at Ruby
Planning Synergy
- Opus 4.7 produces cleaner, more specification-focused plans that GPT-5.5 executes exceptionally well
- GPT-5.5 has "boldness" and "agency" to delete files and completely rewrite rather than patching
- Detailed, contract-style prompts draw out the best performance
Writing
- Strong at business writing (investor updates completed in near-final form on first attempt)
- Good at voice replication and style mimicking without overdoing it
- Less personality than Opus models but more restrained and professional
- First GPT model in a long time that competes with Claude for writing tasks
Knowledge Work & Desktop Experience
- Best-in-class desktop agent experience through Codex desktop app
- Significantly faster than Opus 4.7 (hardware advantage evident)
- Excellent at web browsing, dashboard writing, and complex data analysis
- One weakness: reduced eye for detail compared to Opus 4.7 for tasks requiring sharp insights
Notable Quotes
> "Write a plan with Opus 4.7 and GPT 5.5 is an absolute beast."
> "If you give it a plan like that, it just actually does it. It has the confidence to do it."
> "It's the most usable way right now to get the power of a frontier model in a user-friendly, human-friendly, collaborative package."
> "I feel like it's Christmas every couple of weeks around here because we keep getting access to new stuff and it keeps being awesome."
> "All of our powers are just increasing day by day."
Takeaways
When to Use GPT-5.5
✅ Best For:
- Large-scale code refactoring with detailed specifications
- TypeScript/Swift projects
- Business writing and professional communication
- Desktop agent tasks and knowledge work
- Daily driver model for general purpose use
⚠️ Consider Alternatives For:
- Ruby/Rails projects
- Tasks requiring extremely sharp analytical insights
- Design-forward product engineering
- Open-ended vibe coding without detailed specifications
Workflow Recommendations
- Use Opus 4.7 to write detailed plans → Then execute with GPT-5.5
- Provide specific, contract-style prompts to activate GPT-5.5's execution capabilities
- Use for agentic desktop tasks - significant performance and speed advantages
- Leverage for business writing - excellent voice and style matching
Bottom Line
GPT-5.5 represents a genuine step change in AI capabilities, particularly when used as part of a synergistic workflow. Its combination of speed, execution confidence, and user-friendliness makes it a compelling daily driver model, though it works best when paired with detailed planning and specification.
Transcript
It's model release day. Let's go. Today is the release of Spud GPT 5.5 from OpenAI. We've been testing it internally for about three weeks. It's a sick model. We've tested on everything from coding to writing to knowledge work tasks. And it's actually a real step change in a lot of abilities. The big headline for me though is its coding ability. On our senior engineer benchmark, it scored a 62.5 out of 100. And by comparison, human senior engineers score consistently about 80 or 90 out of 100 on this benchmark. It's not in the senior human engineer range yet, but it's getting up there. Opus 4.7 by comparison scored in the low 30s pretty consistently. So there's a 30 point gap between the work that GPT 5.5 does and the work that Opus 4.7 does. But there's a catch. The best performance that we got out of GPT 5.5 came when we ran it on our senior engineer benchmark using a plan written by Opus 4.7. Write a plan with Opus 4.7 and GPT 5.5 is an absolute beast. We're going to go into a ton of detail on all the tests we ran and what we found, what this new model is useful for, what its drawbacks are. You can get it into your work and into your life today. I love model release days because those are the days where we publish our vibe checks. This is a video version of the vibe check, but if you go to the link down below, you're going to see our written version. It's a couple thousand words with all of our detailed thoughts on exactly what we did with this model and what we thought released on the day of. We write new things every day to keep you at the edge of AI. We also build things ourselves as part of the Every Subscription. We have a fleet of six apps that we've built with Codex and CloudCode with these models to help us in our work. Things like Quora, which is an agent for your email, or Sparkle, which is an antique file organizer, or Monologue, which is a smart application app. All these are things we build ourselves and release to our audience as part of our subscription. We also do a lot of training, so we run camps and courses to teach you everything you need to know to code, to write, to do design with AI. And we even do this for big companies and executive leadership teams. So if this is interesting to you, you should go to every.to and subscribe. Okay, so let's get into the vibe check. I've got four main categories for you. We've got coding, we've got writing, we've got knowledge work, and we've got open claw. How does GPT 5.5 stack up versus 5.4 and versus Opus 4.7 and other models? We're going to go into that in detail. So let's start with coding. The headline is the Senior Engineer Benchmark. As I said, GPT 5.5 scores about 30 points higher than Opus 4.7 on the Senior Engineer Benchmark. But it only does that when it is used with a plan written by Opus 4.7. I think this tells us a lot about the psychology of this model, what it's good at, what it's not good at, and when to use it in your workflow. To understand this a little more, I want to talk to you about what this Senior Engineer Benchmark is, the SE Bench. It's a benchmark I invented, and the benchmark basically gives the model a vibe-coded slap codebase. It's a real codebase from an app that I built called Proof, and it says to the model, hey, this is vibe-coded slap. How would you rewrite this from first principles? Basically, is it able to rewrite the codebase? Do it in a clean first principle way where there's conceptual clarity, and it looks like a senior engineer did it. And the gold standard for this benchmark is actual code written by senior engineers. I had two different engineers do their own rewrites of this codebase. So we get to compare what these models do versus real human engineers. And first of all, this benchmark is not saturated. There's still a lot of room on the frontier. The best human engineers get around 80 to 90 points on this benchmark out of 100. The best score we've ever seen is GPT 5.5. It got about a 62.5 using a cloud Opus 4.7 plan, and it got in the low to mid-40s without. That's still much better than Opus 4.7, which has got about a 33. If you do some prompting, you can get 5.5 to make its own plan good enough that its performance on this benchmark starts to approach its performance when it uses an Opus 4.7 plan. But it still only gets to about the low to mid-50s. So for my money, Opus 4.7 is just a better planner than 5.5. What this model can do that other models don't seem to be able to do is it's able to identify the underlying core principles or invariants that need to be true in the codebase, in the plan that it writes. And then when it starts writing code against that plan, it doesn't get as distracted by the existing codebase. It doesn't go into patch mode and start patching little holes. It has the assertiveness, the boldness, the agency to actually go and delete a bunch of files and really start from scratch and then carry through the idea it has from start to finish over several hours. It's not perfect yet. It's far from perfect. As I said, it's still about 30 points away from senior engineers on this benchmark, but it is a step change in terms of how good it is versus other models. What happens with Opus 4.7 is Opus 4.7 actually produces a really good plan. It's very good at writing something that's conceptually clear and clean, something that feels contract driven and gives the model a sense for if this rewrite is good enough, this big file will only be 100 lines, that kind of thing, that level of detail. That's what drives GPT 5.5's performance. That's what makes it so good is if you give it a plan like that, it just actually does it. It has the confidence to do it. But Opus, if you give it its own plan, it's a really beautiful, well-written plan. It says, oh, this is too much effort. I'm just going to pick off a small little slice and then it just starts to patch around the issue rather than actually do the rewrite as asked for. GPT 5.4, the old OpenAI model that's being replaced today, does the same thing. It does it better than Opus 4.7, but it still does the same thing. That's what makes it so good is if you give it a plan like that, it just actually does it. It has the confidence to do it. But Opus, if you give it its own plan, it's really beautiful, well-written plan. It says, oh, this is too much effort. I'm just going to pick off a small little slice and then it just starts to patch around the issue rather than actually do the rewrite as asked for. GPT 5.4, the old OpenAI model that's being replaced today, does the same thing. It does it better than Opus 4.7, but it still does the same thing. But GPT 5.5 on extra high reasoning just has that little bit of extra oomph that allows it to actually go in and execute. Okay, well, what is it about the Opus 4.7 plan that makes GPT 5.5 better? And I think that's really interesting because it says something about what makes a good use of 5.5. If you look at the plan that Opus 4.7 wrote for this benchmark versus the plan that GPT 5.5 wrote, it has all the right concepts, but it's pretty long. And it doesn't have a lot of the really specific programmary contract style. Okay, here's what good looks like. Here's what you should delete. Here's how many files should be left. All that kind of stuff, it's really written well for humans. And I think they've done a lot of tuning on 5.5 to make it good for humans. And if you look at 4.7, part of the reason people don't like it is because it's terse and it feels a little bit more robotic than you're used to with anthropic models. But it turns out that terseness, that exactness, that contract style plan is really good for 5.5. So if you want to get the most out of 5.5, using Opus 4.7 as the planner or making sure you prompt 5.5 with a lot of exact detail is going to be the thing that draws out this boldness, this agency, this ability to carry a big plan through from start to finish. It's really impressive. Naveen, who's the GM at Every who runs Monologue, tested it on building a to-do app called Dayline. And he was really impressed by how it was able to build a native iOS and Mac app that was beautiful, works really well, and it can just churn through a bunch of features in a plan until it's done. It's really good at that, especially if the plan is well-specified. Naveen also used it to build a release of Monologue, the app that he runs. And he used about 900 million tokens with GPT 5.5 in pre-release as part of his testing. He thinks it's his favorite model for everything. He said he was only able to hit the deadline that he needed to in order to get a new feature out for Monologue by using this model. Naveen's an incredible senior engineer, and if he likes this model, it's worth paying attention to. There are some things that it's not as good at. Kieran Klass and the general manager at Quora, which is our AI email agent at Every, tested the model on his LFG bench, which essentially replicates his real development process to see how a model does on more product-forward engineering tasks, where you're not necessarily refactoring a whole code base, but you're building a feature and something that includes a lot of front-end and design and product thinking. And he found that for these types of tasks in the LFG bench, 5.5 did pretty well, but Opus 4.7 had a higher ceiling, especially on design-forward tasks. It just has a better aesthetic sense than 5.5. And Mike Taylor, who runs technology consulting at Every, had a similar feeling about it in his benchmarks, where getting into VibeCode, relatively complex app from scratch with an underspecified plan, it just didn't do as well as Opus 4.7, which ripped through the entire task. So whether you're VibeCoding or you're trying to do some senior engineer type stuff, a better specified plan is going to be the way to get the most out of this model. One other thing that we noticed is that GPT-5.5 is really good at writing TypeScript and writing Swift, but Kieran, who is a huge Ruby stan, thinks it doesn't write good Ruby. So if you're doing a Rails project, you may not be happy with the quality of the Ruby that gets written. But if you're in TypeScript or you're in Swift, you're going to be pretty happy. Okay, now on to writing. It's got a little bit less personality than Opus, especially the older Opus models like 4.6. But it's actually really good for doing business writing. So I used it to write our investor update and it basically one-shotted an update that was close to being ready to send. Katie Parrott, who's a staff writer at Every, has been using Claude models for probably a year or two now, almost exclusively for writing. This is the first GPT model in a long time that she has started to use for writing tasks over Opus or Sonnet. In particular, Katie and Mike both really liked it for its voice replication. It was pretty good at mimicking a style without overdoing it. It's a little bit more restrained, which I think is why it's good at business writing. And that makes it a little bit more subtle as a writer. Okay, now on to knowledge work. This is one of those things that OpenAI was so behind on even three months ago or six months ago. And they have just rapidly iterated to the point where the Codex desktop app is great for any kind of knowledge work. And using it with GPT 5.5 is best in class agent experience you can have on your desktop. For one, it's really fast. It's really powerful. It can use any of the apps on your computer. It's good at browsing the web. It's good at doing things like writing dashboards or doing complex data analysis. And again, it's really fast. And in all of these tests, I haven't talked about this that much, but in all of these tests, we were pretty shocked at how slow Opus 4.7 feels in comparison. You can really tell the hardware advantage that OpenAI is working with right now. You can feel it. One thing about this model that I ding it for versus Opus is some of the training It's really powerful. It can use any of the apps on your computer. It's good at browsing the web. It's good at doing things like writing dashboards or doing complex data analysis. And again, it's really fast. And in all of these tests, I haven't talked about this that much, but in all of these tests, we were pretty shocked at how slow Opus 4.7 feels in comparison. You can really tell the hardware advantage that OpenAI is working with right now. You can feel it. One thing about this model that I ding it for versus Opus is some of the training that's being done to make it feel digestible has come at the cost of some of its eye for detail and knowledge work. If you're using this model for tasks that require really sharp insights, I would definitely consider using Opus 4.7 over 5.5 for that. For example, even the senior engineer benchmark, like grading the model trajectories, I just trusted 4.7 more than I trusted 5.5. Even if I want to use 5.5 more as my daily driver. It's a great daily driver model. So that's the gist of 5.5. I feel like it's Christmas every couple of weeks around here because we keep getting access to new stuff and it keeps being awesome. And I feel like my powers and everybody else on the team, like all of our powers are just increasing day by day. You should definitely give this model a try. It's the most usable way right now to get the power of a frontier model in a user-friendly, human-friendly, collaborative package. And that's a big achievement. Until recently, I pretty much used Claude for everything. Over the last couple of months, I've really switched my usage where I use Claude a decent amount on mobile. I like the mobile app. But for anything on my computer, anything agentic, I'm pretty much in ChatGPT for everything. And I got to say, I like it in OpenAI. If you're an Opus stan, the harness is really still built to work better in Opus. But they're releasing stuff every day that makes it pretty good in OpenAI. And it's been one of those stable experiences for me over the last couple of weeks where it's not forgetting stuff as much. It's not doing dumb stuff as much. It still does it a little bit. But because it's free under my ChatGPT subscription, I'm okay with a little bit of the trade-off in power. And I think that they'll fix that hopefully soon. So if you didn't like 5.4 in OpenAI, I recommend giving it a shot in 5.5. So that's it, folks. That's our report from the frontier. If you liked this, you should like and subscribe. But more importantly, you should go to every.to. We have a couple thousand words on this model with everything that we tested and everything that we found. We also have a suite of apps for subscribers. Stuff that we build ourselves that help you work better with AI. Stuff that helps you write your emails. That helps you organize your files. That helps you write better. We've got a ton of good stuff in the Every subscription. If you try this model, we'd love to know what you think. So please leave a comment. And remember, stay hydrated. See you next time. On the moon! underspecified plan, it just didn't do as well as OVS 4.7, which is ripped through the entire task. So whether you're VibeCoding or you're trying to do some senior engineer type stuff, a better specified plan is going to be the way to get the most out of this model. One other thing that we noticed is that GPT-5.5 is really good at writing TypeScript and writing Swift, but Kieran, who is a huge Ruby stan, thinks it doesn't write good Ruby. So if you're doing a Rails project, you may not be happy with the quality of the Ruby that gets written. But if you're in TypeScript or you're in Swift, you're going to be pretty happy. Okay, now on to writing. It's got a little bit less personality than Opus, especially the older Opus models like 4.6. But it's actually really good for doing business writing. So I used it to write our investor update and it basically one-shotted an update that was like close to being ready to send. Katie Parrott, who's a staff writer at Every, has been using Claude models for probably a year or two now, almost exclusively for writing. This is the first GPT model in a long time that she has started to use for writing tasks over Opus or Sonnet. In particular, Katie and Mike both really liked it for its voice replication. It was pretty good at mimicking a style without overdoing it. It's a little bit more restrained, which I think is why it's good at business writing. And that makes it a little bit more subtle as a writer. Okay, now on to knowledge work. This is one of those things that OpenAI was so behind on even three months ago or six months ago. And they have just rapidly iterated to the point where the Codex desktop app is great for any kind of knowledge work. And using it with GPT 5.5 is best in class agent experience you can have on your desktop. For one, it's really fast. It's really powerful. It can use any of the apps on your computer. It's good at browsing the web. It's good at doing things like writing dashboards or doing complex data analysis. And again, it's really fast. And in all of these tests, I haven't talked about this that much, but in all of these tests, we were pretty shocked at how slow Opus 4.7 feels in comparison. You can really tell the hardware advantage that OpenAI is working with right now. You can feel it. One thing about this model that I kind of ding it for versus Opus is some of the training that's being done to make it feel digestible has come at the cost of some of its eye for detail and knowledge work. If you're using this model for tasks that require really sharp insights, I would definitely consider using Opus 4.7 over 5.5 for that. For example, even the senior engineer benchmark, like grading the model trajectories, I just trusted 4.7 more than I trusted 5.5. Even if I basically want to use 5.5 more as my daily driver. It's a great daily driver model. So that's basically the gist of 5.5. I feel like it's Christmas every couple of weeks around here because we keep getting access to new stuff and it keeps being fucking awesome. And I feel like my powers and everybody else on the team, like all of our powers are just increasing day by day. You should definitely give this model a try. It's the most usable way right now to get the power of a frontier model in a user-friendly, human-friendly, collaborative package. And that's a big achievement. Until recently, like I pretty much used Cloud for everything. Over the last couple of months, I've really switched my usage where I use Cloud a decent amount on mobile. I like the mobile app. But for anything on my computer, anything agentic, I'm pretty much in codex for everything. And I got to say, I like it in OpenClaw. If you're an Opus stan, the harness is really still built to work better in Opus. But they're releasing stuff every day that makes it pretty good in OpenClaw. And it's been one of those stable experiences for me over the last couple of weeks where it's not forgetting stuff as much. It's not doing dumb stuff as much. It still does it a little bit. But because it's free under my ChatGBT subscription, I'm okay with a little bit of the trade-off in power. And I think that they'll fix that hopefully soon. So if you didn't like 5.4 in OpenClaw, I recommend giving it a shot in 5.5. So that's it, folks. That's our report from the frontier. If you liked this, you should like and subscribe. But more importantly, you should go to every.to. We have a couple thousand words on this model with everything that we tested and everything that we found. We also have a suite of apps for subscribers. Stuff that we build ourselves that help you work better with AI. Stuff that helps you write your emails. That helps you organize your files. That helps you write better. We've got a ton of good stuff in the EverySubscription. If you try this model, we'd love to know what you think. So please leave a comment. And remember, stay hydrated. See you next time. On the moon!