LIVE VIBE CHECK: GPT-5.5 Has it all
Description
The new OpenAI model is faster and easier to use than Opus 4.7, with surprising strength across coding, dashboards, writing, and enterprise workflows. Come join the Every team for a live vibe check to see how it compares across the board. Every is the only subscription you need to stay at the edge of AI. Start your trial today: https://every.to/subscribe Ready our full GPT-5.5 vibe check: https://every.to/p/gpt-5-5
Summary
Generated by claude-haiku-4-5-20251001LIVE VIBE CHECK: GPT-5.5 Has it all
Main Topics
- GPT-5.5 Release & Testing: A comprehensive three-week testing period by the Every.to team across coding, writing, design, and knowledge work
- Model Comparison: Detailed comparisons with Opus 4.7, GPT-4, and Claude models
- Senior Engineer Benchmark: Performance metrics showing GPT-5.5's strengths in code refactoring
- Codex Integration: New capabilities combining GPT-5.5 with OpenAI's Codex development environment
- Workflow Applications: Real-world use cases from developers, product builders, writers, and growth professionals
- Image Generation Integration: New workflow using GPT Image 2.0 to design and implement web applications
Key Points
Model Performance
- Senior Engineer Benchmark Score: GPT-5.5 scored 62/100 vs. Opus 4.7's 33/100 on code refactoring tasks (30-point improvement)
- Best when combined: GPT-5.5 performs optimally when using plans created by Opus 4.7
- Speed advantage: Significantly faster than Opus 4.7 for knowledge work and execution
- Execution capability: Unique ability to execute large-scale code rewrites without getting distracted by existing code
Strengths
- Reliability over creativity: Described as a "workhorse" model—dependable, thorough, and consistent
- Personable & collaborative: Better tone than previous OpenAI models; feels less condescending
- Context management: Excellent at maintaining context across multiple code bases and long conversations
- Attention to detail: Can implement complex requirements from minimal specifications
- Multi-modal capabilities: Works well with image inputs and can follow visual design specifications
Weaknesses
- Design inconsistency: Some visual design elements less polished than previous versions (though typography improved)
- Language-specific limitation: Not as effective with Ruby or non-TypeScript languages
- Design creativity: Less creative in visual design compared to Opus 4.7; better suited for structured tasks
- API not yet available: Only in ChatGPT and Codex; API access coming within days
Use Case Applications
- Senior engineers: Code refactoring, complex bug fixes, large codebase analysis
- Product builders: Full-stack development, one-shot app creation, iterative feature building
- Knowledge workers: Campaign planning, research synthesis, meeting notes analysis
- Writers: Content creation, editorial work, faster iteration with direction-following
- Growth professionals: Dashboard creation, analysis, marketing plan generation
Notable Quotes
> "It's a very good model. When I ran it in 4.7, I was like, 'Oh, this is the best coding model out there.' And 5.5, GPT 5.5 scored the same on that specific benchmark."
— Kieran Klassen (CEO/Founder, Every/Quora)
> "For tasks where I know I'm not going to be able to pay that much attention, and I just need to get it done, I want to make sure it's safe. I can delegate it without concern. I think this model has it."
— Mike Taylor (Head of AI Tech Consulting, Every)
> "This is a workhorse. It just gets the job done without fuss. For most writing tasks, I don't need witticisms. I need the job done."
— Mike Taylor on the model's writing capability
> "The bottleneck is not one-shooting things because that's working. For me, the bottleneck is how do you collaborate as a human with the model?"
— Kieran Klassen on the future of AI development
> "I am not even specifying much of the task anymore. I just trust the model much more to figure things out."
— Roman Hewitt (Head of Developer Experience, OpenAI) on internal usage
> "The level of attention to detail that the model puts in is incredible."
— Dominic Kunder (Codex DX Lead, OpenAI)
Takeaways
For Developers
- Adopt GPT-5.5 for execution: Best choice for implementing well-specified tasks, large refactors, and multi-file projects
- Combine with Opus 4.7: Use Opus for planning, then GPT-5.5 for implementation for optimal results
- Use Codex harness: The desktop app and plugin ecosystem dramatically enhance model capabilities
- Leverage computer use: New screen control features enable autonomous task completion across native apps
For Product Teams
- One-shot capability: Minimal specs needed; the model can figure out complex requirements
- Rapid iteration: Speed enables fast feedback cycles within meetings or during other activities
- Image-to-implementation workflow: Use GPT Image 2.0 for design, then GPT-5.5 for implementation
- Context integration: Connect Slack, Notion, and other tools for informed decision-making
For Writers & Content Creators
- Faster iteration: Superior at taking direction and refining work based on feedback
- Reliable output: More dependable than previous models; less prone to unexpected creativity
- Integration benefits: Slack integration enables real-time feedback incorporation during writing
- Breaking from Claude: Worth reconsidering for those who switched away from OpenAI
For Growth & Business Professionals
- Knowledge work: Viable alternative to Claude for dashboard creation, analysis, and planning
- Speed matters: The performance difference makes it comfortable to use interactively during meetings
- Automation recommendations: Codex now suggests and implements automations automatically
- Context awareness: Can synthesize information from multiple sources for comprehensive planning
Strategic Observations
- Model trajectory: OpenAI has made GPT models viable for non-engineering knowledge work—a significant shift
- Ecosystem matters: The model's power is multiplied by Codex, plugins (100+), and integration features
- Implementation challenges remain: Design-focused work still benefits from Claude; language support varies
- Next frontier: Focus shifting from "can we build this?" to "how do we collaborate effectively with AI?"
For Every.to Subscribers
- Access comprehensive vibe checks on all model releases with hands-on testing across multiple domains
- Get detailed benchmarks and real-world examples before making adoption decisions
- Leverage consulting services for enterprise AI transformation and adoption
- Use integrated app suite (Spiral, Quora, Monologue, etc.) optimized for AI-native workflows
Transcript
Hello, everybody. Welcome to Vibe Check Day. We've got a new model coming out. It's actually, it is out. GPT 5.5 is out. We've been testing it for the last three weeks or, and it's really fucking great. It's really great. It's a very exciting model. I've loved it. It has become my daily driver. So what we're going to do on this stream for you is we're going to go through our Vibe Check, which we just published on Every, Every.to. Every is the only subscription you need to stay at the Edge of AI. I'm going to go through all my findings. We're going to have other people from the team who are going to come on. Mike Taylor, come on over. This is Mike Taylor. This is good. This is my best idea. Why don't you take that chair? So we've got me. We've got Mike Taylor, who's our head of AI tech consulting at Every. Mike, you're not quite in the shot. All right. Great. And we are going to go through this model for you. So what I'm going to do is I'm going to share my screen. And here we go. Here we go. Yes. So GPT 5.5 is out. It is live right now. They are not releasing it on the API quite yet, but it should be out in your codecs, in your chat GPT. And I assume that will start rolling out over the next couple of hours. They're holding it in the API for just a bit because it's just a very powerful model and they're doing a lot of testing on it, which I think is probably good. It seems to be the new standard for these models. we seem to be reaching a level where people are, oh, this is worrying, but okay, here we go. Here's our vibe check. We just published it on every.to slash P slash GPT dash five dash five. And our headline is GPT 5.5 is OpenAI's Workhorse model. What's really it. So it is out. It is out. There's a, there's a blog post on their site. You, you should be able to see it. We, what we, what we found in testing this model is it's actually quite rare for a model to be really good at both senior engineering type tasks and a really good workhorse. And this model is that it is super collaborative. It's super fast. It's actually like pretty personable. And it did the best on our senior engineer benchmark of any model that we've tested. So on our senior engineer benchmark, which tests how good a model is at rewriting a, an existing code base in the way that a senior engineer would benchmark against real on a real code base that two real senior engineers rewrote separately., GPT 5.5 scored a 62 as its best score. Opus 4.7, 62 out of a hundred Opus 4.7,,, by comparison scored, I think its best score was a 33. So there's like almost a 30 point swing between Opus 4.7 and GPT 5.5 on this task. There's an interesting caveat though. Do not throw out Opus 4.7 yet because our,, best, the best performance of this model of GPT 5.5 comes from,, using a Opus 4.7 plan. So if you use them together, they get super powerful., codexes are GPT 5.5 is the model you want to be coding., but you actually at Opus 4.7 plans, I think are actually still better than, than, than 5.5. But,, when you actually put it on this benchmark, it is like miles better at this task., so we have a full, we have a full vibe check here., it's,, it's on every, every.to slash,, you should be able to see it. Every.to slash p slash vibe slash p slash gpt dash 5 dash 5. Full vibe check is here,, published already., it's a totally new pre-trained model. So this isn't a fine tune of a previous, a previous GPT 5 version. It's the spud model that you've been hearing a lot about. It's super powerful., but there, there's like a mix of reactions here. So it has absolutely become my favorite, my daily driver model., it's, I use it for everything from my day, my day-to-day work to like real engineering tasks, but other people on the team have different reactions. Mike Taylor, who's our,, head of,, technology consulting at every has, has us had a slightly different,, experience with it. Mike, do you want to,, just talk a little bit about what you found in testing this model? Yeah, sure.,, I was also very impressed. I think it's the most reliable model that tested., I felt like comfortable and safe, like getting into a Waymo,, it's,, it like feels pretty good, but I still think like Opus is the Tesla,, it's like a little dangerous., and,, when I'm micromanaging tasks, I, I still use Opus as my daily driver, like if I'm heavily involved,, but I would say for,, tasks where I know I'm not going to be able to pay that much attention,, and I just needs to get done. I want to make sure it's safe. I can delegate it without,, with peace of mind. I think like this model has it., and give me, give me an example of like the tasks that you might want to,, use this model for that,, and we just got Kieran Klassen, Kieran GM of every, of Cora at every, and,, the creator of compound engineering. Kieran, welcome. Hello everyone.,, so Mike, just give us, give us an example of a task where you're, this is actually a really big difference. I might, I'm, I really want to try this today using this model that you might not realize. Yeah. So one thing I've been using it for this week is,, creating the curriculum for training material,, because this is,, quite an onerous task., you have,,, tons of call notes from different people across the organization., we're, we're working on AI adoption,, with a few companies trying to get,, the teams on blocks so they can use more AI. And,, it just requires a lot of diligent work of, we need to go through all the call notes, make sure that,, all the different themes and issues that we found in those notes have represented in the curriculum. And,, it like hasn't failed on that once., I've been,, and, and, and also I would say that with, with Opus,, I was using Opus for that previously. And with Opus, I always felt I had to like go back through it line by line., and it came up with some really sharp stuff and like some of the titles were, cool. I'm, oh, I'd love to teach that. But, but actually, you don't want cool titles all the time when you're doing corporate training, They actually want something dependable, reliable that,, regular people will find accessible. So I would say like if I was writing marketing copy, it would be Opus probably. But,, if it was,, something where I need it to be,,,, unoffensive, I need it to be,, reliable, I needed to capture all the main notes and not be like a little bit wild. And,, yeah, this is, this is the model. Great., so we've got a couple more people to, to add into this vibe check. So we've got, as I said before, we've got Kieran Klassen, the GM of Cora and the creator of compound engineering, and we've got Naveen, the GM of Monologue. Hello, Naveen. How's it going? Hi, hi. Good. I'm really excited., okay. So I, we, I, we've got, if you look at our vibe, if you look at our vibe check,, what we do on these vibe checks is every time we publish one,, this comes from,, three weeks of internal testing. We always do the reach test, which is, do you reach for it every day for? And if you do, what do you reach for it for? And what do you reach for it over? So GBT 5.5 is absolutely my, my daily driver model. I'm a green. Kieran though is a yellow. So Kieran, you have some, I think you have some, some nuanced thoughts on the strengths and weaknesses of this model. Can you talk to us about what you found? Yeah, absolutely. So for, for me to use a model and use a daily,, it has to help me build Cora. I'm building a new version of Cora., and it's a lot of product work. It's a lot of coding. It's a lot of testing., it's very wide work, lots of work. So front and back end, everything. And what I noticed is that while 5.5 is a very good model, like in my benchmarks that I ran,, like when, when I ran it in 4.7, I was, oh, this is the best coding model out there. And 5.5, GBT 5.5 scored the same on that specific benchmark. So it's a very, very good coding model., but why, why don't I use it every day is that it's, it's also, it's also, it feels more like a specialist and less like a generalist. That's how I look at it. Like Claude is the generalist that is a very good coding model, but it's also very good at product work. It's good at going into details. It's,, it's going, it's good at like looking at the big picture and for,, 25.5, it's very good in execution and going into details, but sometimes it like breaks down if you look at it from far away and you just see things not being coherent. And I need that. I need a generalist. I need some, something that is working for everything. Although I'm using GPT 5.5 for review for execution. I use it in my flow, but it's not the daily driver that I go for first. Like the daily driver is still 4.7 because it's better of a generalist. And since I do product work and more general work,, it's, it's,, yeah, it's, I think it's still the same philosophy as we ever always seen. Open AI just like have a different perspective on what engineering work is than anthropic. And you see that in the model. It's not good or bad. It's just, some people are more aligned with one than the other. And I think Naveen, for example, is more aligned with the open AI take on what engineering work is than my take. I'm more of like a product engineer generalist. Naveen is maybe more of like an engineer engineer. And, it's just for different people. It's not better or worse, but for me personally,, while it's a very strong model and I've one shot at like amazing things with,, 5.5, all the models are very good, which is a very good luxury we have. Like there's nothing bad about any of these models. They're amazing. Yet like these little things make you go for one or the other., so on that note,, Naveen, GM of monologue,, you're, you're usually a big GBT fan., so I'm curious what you think of this release and, and how you compare and contrast Kieran's feeling about this model. Cause I, I'm having this experience where I feel I totally see some of the things that Kieran says, and I have a slightly different feeling. And I think that you have some different feelings. So there's, I think there's a lot of different viewpoints and different types of people doing different types of work that make this model feel, Holy fucking shit. I've never, I I've never seen a model do this versus,, maybe it's not as, as good of a generalist for product work,, like Kieran's finding. So Naveen, tell us what you found in your testing. So I, I used to agree with Kieran, with the older models of GPT, when he says this, yeah, I am with you because I also used to reach out to cloud code,,,, opus models,, but coming to this 5.5, it just feels really well rounded model here in my case, I am,, work,, in Python code base or shift code base, like native Mac app,, when it coming to web app codec, like codecs is not really good. So when you're building front end or anything, I always reach out to cloud code or some other model, but with 5.5, I didn't reach out to any other model. I know, Opus 4.7 released, I tried it a couple of times, but I just went back to 5.5 because it's really great for like the everything before,, I, I also do support, right. From on log, I do ton of,, replies,, obviously we use AI to do that before I used to use cloud code because it's like really good at writing, but with 5.5, I just completely moved to move that as well., and main drawback with the older models is vibe coding where I know for sure I can't pick codecs as my vibe coding tool., if I get come across like any other,,, any idea, I just always go to cloud code, but in this case as well, I vibe coded three different apps in last couple of years. Can you show us some of the apps? Cause, cause one of the thing, one of the cool things is that,, Naveen used like 900 million tokens on GPT 5.5 over the last couple of weeks., so he's really, really,, making those GPS go burr. So Naveen, can you show us, I know you did an app called Bayline. So I'd love for you to show us an app that you are able to vibe code with GPT 5.5 on the side, while you're also shipping a bunch of features for monologue while you were also, you had pink eye. So you were pretty, you're out of commission for a little bit there. I think I really helped me wipe code this because I'm just,, resting, but I have monologue dictating into,,, codecs. Let me share my screen. I'm sharing my screen. Hey,, so first one second, one second. Okay., I removed myself and then I'm gonna add you. Good. Okay. Now we can see your screen. Can you zoom in a little bit? Okay. Nevermind. Oh, okay., let me open my downloads. I just like tried to download recent things. So one, okay. Vipe coding day line. Yeah. I want to show you guys,, this thing where it's a Raycast alternative. I personally use a Raycast notes like this notes where you can just bring this whenever you want to on top, but the thing is, this is like a plain text, I don't want this plain text. So I thought, okay, let me make the to-do list, which is day to day. I can just have this like to-dos on a daily basis. So you can see how I'm already using it., and yeah, this whole thing, I didn't look at single code. It just completely, I just gave a screenshot,, of Raycast., okay. I really like Raycast. Like just go take,, do, do it. And then it wipe coded. What's really cool thing about it is though, the Mac app that you see here, you can see, I, I hit enter, you can type it, I hit enter. You can go back all these like interactions, minor interactions that I'm using. It's like difficult to implement,, in a Mac app. And it did that all in one single thread. This is one single, I think 200 million token thread that you're seeing right now. That's 200 million tokens. Oh my God. Yeah. It's all in one single thread., here I used,, build iOS apps,, plugin. That's the plugin that,,,, codex comes with. This is my long prompt. I want to implement it today. This is what I was doing when I got pinged. So yeah. And then it implemented the first thing you can see 49 messages., wow. It just went at it. It implemented the initial version. I was like started, Oh, can you actually take a screenshots on the Mac app and then,, do it for iOS as well. Now there is iOS app as well that syncs automatically. I'm just like blown away with,, you can see, I just keep on having this to and fro and I'm just now regularly using it. I don't think I will be releasing it for anyone else, but it's just fun to test it, test it out. And today I experienced one bug,, and I just said that it fixed it and I'm just running it again. That's it. So that's really, that's really interesting. Yeah. I think there, there's something really interesting here. I want to pull out., because,, because I've done a lot, I've done a lot of testing with this and I think some of your results are similar to mine and the, the, some of the things that I've found are, so I have this senior engineer benchmark. And I said, if you test GPT 5.5 on the senior engineer benchmark versus Opus 4.7,, 5.5 scores about a 62 and four, seven scores about a 33. So there's a 30 point difference, but the really interesting thing is that 5.5 scores the best. It gets that 62 when it uses an Opus level plan, when it uses a plan literally written by Opus 4.7. So that's really interesting. Right. And I tried to tease apart what, what was going on.,, the thing that 5.5 can do that four, seven cannot do is if you give it a prompt that says, Hey, I want you to like go and rewrite a major part of this code base. I don't, I don't know what it is, but this is just a vibe. This is a vibe coded slop code base, like figure out how to,, how you would rewrite it from first principles and then do it both four, seven and five, five can figure out, okay, here are the core principles. Here are the core invariants that I would rewrite and write a plan for it. But what happens is when you then say, okay, go execute it for Opus 4.7 says to itself, says to you, Hey, actually like this is way too big of a project. I'm just going to pick off like a little piece. And even if you push it a little bit to, no, do the whole thing, it will still only do a little bit of a patch over the problem. It doesn't want to go and like actually rewrite the whole thing and actually delete a bunch of code. So it gets distracted and a little bit almost intimidated by a really big code base and a big rewrite. Whereas GPT 5.5 is really interesting for its ability to take that plan and actually execute it over many turns, over many, many, many hours, over many, many tokens. And for having this like ability to,,, to like have the almost like the courage or the assertiveness to be, okay, I'm going to go and like delete a bunch of code and I'm really going to think about this from first principles without getting as distracted by the existing code. And that's just something that that's new. I just have not seen that in, in an Opus model. I've not seen that in GPT 5.4. It's even less present in GPT 5.5 on high reasoning. This is only really on extra high., and, and some of the things when you think about when you, when I looked at, okay, it's, it happens on extra high and it happens on extra high with an Opus 4 7 plan. What's different about the Opus 4 7 plan? The Opus 4 7 plan is, is very, it's very terse and very spec like and very contract heavy. It says things like a good rewrite will take this gigantic file and get it down to 500 lines. it has almost like contracts that it's giving to 5.5. And even though Opus won't itself execute those contracts, 5.5 is totally happy to go do that. And I think that's a really, that's a show, something really interesting about the character of this model. And I think it's part of what you're seeing to be in, in your vibe coding results because you're giving it, you're not doing a detailed engineering spec or whatever, but you are,, if you go look at the original prompt, you are giving it a lot of detail that a, a like baby vibe coder maybe wouldn't. And it, what it's able to do is take that detail and keep it in its head and actually like turn it into like a full thing from start to finish. Does that make sense? I see that. Do you want to talk a little bit about that? I would love that yeah. Yeah so I have a, one of my benchmarks is, is like purposely very under specified. I don't use plan mode and it's just create a version, like a clone of type form, but call it talk form. And the goal is like the backend is exactly the same as type form, structured data but the front end is you're just talking to the person you're interviewing them to get the answers for that structured data because it's i don't think this exists but i think it would be cool if it did exist and i've vibra-coded that a few times in workshops and and and the prompt is like four lines on purpose because i what i really like to see is like from scratch with no repo with a very underspecified prompt like as a beginner like is this model accessible and so i we we talked about this because i was like oh it didn't do very good on vibe coding and then you guys are what are you talking about this is the best vibe coding model ever and i think that was the difference i think opus is like very ambitious to start and then it has i don't know if you've noticed it says like okay are you guys ready to wrap this up like if i'm having a long conversation with it it almost like feels like panics that it's using too many tokens yeah i hate to anthropomorphize but that's how it feels it's like okay do you want to wrap this up and i'm like no no not yet and it's like okay but now do you want to wrap it up and and so i wonder if that's the difference between what we're doing was like codex seems pretty chill it's like dogged and determined it's just gonna like keep pounding it out time after time after time so i think it really depends on what shape of job that you're doing when you're using these models 100 percent naveen any any reactions there yeah i i i don't completely i'm not seeing that so maybe it says that i do a lot of big brain dump prompts and then recently i'm with dan what you observed which is i recently had to add mcp support remote mcp support to monologue i initially implemented really bad mcp server i think i just did it like so what happened is last week we are like doing the launch i just asked gpt 5.5 to come up with a plan this is where i use the plan and it added a plan across front end back end client side both ios mac os right and then i just said i'm like too lazy so i said implement this plan and then it did that it continuously kept on working across all multiple code bases and that's when i know okay this model is really good at managing that multiple code like across context i think that's why i'm like really curious to test this single thread white coding app as well because it i'm not sure what they're doing with compaction i think the context between compaction or the context between each thread that models are like using it's really good so it's not losing that context i think the product shape that i think that's what makes this model really good yeah i actually have like a side project where i have a one of the carpathy style and knowledge bases that's but i'm doing it away from keyboard it's just running in the background and i i initially set it up as a ralph wiggum loop right like every every time claude stops it it checks to see like what's the next task we have to do and then it takes that task down and then it commits and then it stops and then it starts up again and and i switched to codex on on that a while back and and and i've just just been running and and i realized that i'm it doesn't need that that like ralph wiggum loop anymore so i just switched it to just let it compact itself and and it's been going a lot faster i need i needed less harness essentially which was pretty nice do you i'm curious mike do you have examples to show us of of more knowledge-type tasks that you worked on it with yeah yeah i can pull some stuff up let me hear my okay that'll be great and then while he's doing that kieran what else what else do you have to say about this model what else do you think you you've experienced that we haven't talked about yet yeah so one observation was that design looks a little bit less good than the previous model so it just looked a little bit chaotic but design in sense of structure and typography look better so if you if you if you like typography it's better and so so this is the app where i try to just one shot from one prompt an entire app with like a rubber ducky store that where you can customize ducks and all of that it nailed it everything worked which is like this is the first time with 4.7 this ever worked so 4.7 and 5.5 are beasts this is a web app this is react like a next app it's very very very good at this so if you do web you need stuff to work in full detail this is great i've seen here so i run the lfg bench which is using compound engineering to just run all the steps by itself and it took three times more tokens in the planning and the review state similar to 4.7 so it takes the tokens this was extra high but as you can see here like yeah it looks cool the typography side and the alignment and the padding but the the right side like is weird like what is that even i don't even understand what it is so there is some some degradation like the previous models were better at this in design and yeah i'm sure they're going back and forth they're like okay now we need to structure more make it less chaotic and more structured this one is the island test and i said just make a cozy island and it's pretty good it did the the ground was blue instead of brown and animations were not working completely correctly but it's very close like 4.7 opus did better there like the but yeah it's it's still it's getting better every time and it's cool that you can see where they push the model towards and yeah so design wise this one is actually good only the colors of here like the the nice thing like what i've seen before with gpt models is it adds stuff it doesn't really need to add and it does a little bit too much in the ui that that's not there anymore so like that extra it's it's more like opus where it just shows the stuff it needs to show which is which is nice it doesn't do extra things and yeah so design but like one-shotting i did a i did a library and i just here i can share that i had an idea one morning where i was like oh it's like actually one of the more important things now with ai are like doing a call with someone and going through a product and sharing what you like about it and not like about it and so i created this library call it's i shared my screen but in the morning i was having coffee and i just pulled up this and and i was like hey what what if we have like a react package if you can share my screen then you can see the package got it yeah so i created this like riff rack as like let's go riff together and record it and what i did was i just record screen sharing and then feed it to a model to see what was told but i was like well what if we record like the dom interactions as well in addition to all the other stuff so a model can very precisely see when you say oh this is dumb or this should move somewhere else you click on something it has that element and then that can be extracted very well into like an improvement well as you can see it's three commits yesterday and it's initial commit one one review commit and one more review commit i didn't do anything this is just lfg in one go like one prompt in monologue copy into like compound engineering lfg flow into codex five point or in gpt 5.5 and it did it and it's super detailed it had oh yeah there was one iteration where i said hey like maybe we should have a consent screen it actually proposed that it was like hey like maybe we should do a consent screen that's probably something you want that's missing here came up in the reviews of it yeah sounds good but it has all these things and it has these things hey you can enable recording for urls by doing this and you can even create your own param where you can do your own so like there are these details that are super nice that it adds that i haven't thought of but it's not too much it's not sloppy it feels neat and organized and so it's very good at these things this is typescript i use ruby in cora it's just not good at ruby so it's just very hard for me to use because it's not good at ruby it just does typescript stuff but then in ruby and that's not how you write ruby so for me that is honestly the biggest blocker like if you do typescript or anything open ai trained on it's probably amazing but i built cora and use ruby and ruby is not liking the model as much so you can see like this is like an amazing experience and and i i i rolled this out in cora i will try it out it will just record your screen record all your events and you get a zip file stored somewhere you can share it with me and then feed it and just say lfg ideate come up with the the things we need to build here and yeah that's just a morning one shot and that's the world we live in and i'm not even surprised like some people say oh wow wow amazing guy for me that's the norm one-shotting things for me the next thing is like how can we one shot 10 things and release them automatically so i'm more like for me the bottleneck is not one-shotting things because that's working for me the bottleneck is like how do you collaborate as a human with the model because that's the unblock and i think like the ideation the polish at the beginning and the end like that's where you need most help as a human still i love it thank you for sharing that karen i'm gonna now bring in our head of growth austin austin welcome what's up and austin has been testing this model for three weeks alongside us and the codex app and he is in an interesting place because he's using it for head of growth type stuff so for more knowledge work type pass and he's coming from cloud code so awesome give us the headline on i think you started this experiment using cloud code and co-work mostly and you ended using mostly codex tell us about your experience yeah over the last few weeks the codex app in particular and think it's really because of like the the power and speed of the codex app relative to the cloud desktop app has become my daily driver for everything i do for someone like me without like a technical or engineering background i feel way more comfortable doing any engineering work with the latest open ai models that can be everything from like like running code to build like underlying dashboards across multiple tools shipping landing pages that look at like our figma design or whatever doing analysis creating plans and when i you first like were nudging me to at least try codex a few months ago dan and i i did try it and i was like either i'm too stupid to use this it thinks i'm too stupid to use this or it's talking to me like i can't hang at its level and i was i i wish i could get as much out of this as as you did or as i knew like naveen was a big fan of it and i just kept going back to cloud code in the cli because i was like okay it's executing it at like a level i like and it's it's like nice to me and and engaging and as someone who is still relatively new to this stuff that was important to me yeah you're like codex makes me feel i'm dumb makes me feel i'm dumb it makes me feel bad about it and this came up a lot in the previous open ai models in codex because i would make a plan and in planning mode it would say i have four questions for you and for each question for the options all the options i was i have no idea what the hell you're saying please explain it to me and the tone of the open ai model was like why why would i have to explain this to you why don't you just do what's recommended and especially in knowledge work when i'm trying to run like a campaign or planning it's a weird feeling to have and all of that is gone for me in in the new model that we've been playing with for the past three weeks i both understand what it's saying i also really trust what it's saying the the best example i have here is we are about to launch our plus one product in a few weeks and we're sprinting to do it i'm sprinting to get a big like campaign together and i pointed in the codex app i pointed the model at our notion our slack where like all of our meeting notes and conversations have and i just i did a monologue brain dump of like the high level version of what i wanted the goals of the campaign to be and a previous campaign plan we had and i said go make a campaign plan looking at this and it's informed by like my taste and the team's taste and thoughts because it looks at our meeting notes where we talked about what we want to do and the plan that we're going to go to market with is like 90 of what it came up with i was like so impressed and it was so easy because it grabbed all of this context it asked me questions that made sense to clarify and the most important thing to me is that inside of the codex app it's so easy to use and it's just so fast i find the claw desktop app impossible to use due to the like clunkiness and slowness after having experienced how fast for knowledge work codex is with with these models and then the the other thing that i really really like that i think is driven a bit more by how the codex app works with what i feel rather than the models itself is the way that it recommends automations to me for knowledge work and then just builds them and starts starts using them if i if i ask for them is so powerful and impressive and they're automations that i rely on every day so to me like the only way i would switch back to claude and opus whatever new model is my daily driver is if the if the cloud desktop app got like dramatically faster and better so as like a an easy way to work it it was competitive with the codex app in in that way or if like something about the the model for knowledge work like far surpassed where the open ai models are and to me they're they feel about even for most stuff besides design i agree with karen that we use ai for a lot of like video clipping image generation stuff and over and over again i find myself still reaching for claude for that stuff yeah that makes sense i i do i think it's really important just to like talk about how remarkable that shift has been like even three months ago definitely six months ago but even three months ago it was it was one of those things where you just couldn't use codex for any knowledge work it was too slow and it treated you like and now it's super fast and it's just good for the kinds of things that if you're ahead of growth or you're like me you're doing like analysis of like some of our metrics or you're writing investor updates or whatever it's just it's the thing that i want to reach for i think obviously we've been talking about opus 4.7 has some unique powers so for example i think it's planning is much better we think it's design is much better but it's also even in the places where it's roughly equivalent which i think there are a decent number of places like that it is actually a lot slower and 5.5 is just a fast model and that that makes a huge difference i think speed is in some ways it's it is a type of intelligence and it gives gives you a lot more power to have it be so fast yeah it passes this i don't have a benchmark version of me yet i want to make one of like my like do my knowledge work while i'm in meetings test where i'm like that or i did this on sunday where i was watching the nba playoffs and i'm mostly watching the game and i'm just i'm like nudging it along and my brain is only maybe like 10 working and when i'm in between meetings i'll check the outcome of whatever it is and this is a lot of like planning execution stuff and it moves so fast that i like can't keep watching the basketball game or i can't i like during the meeting i'm like oh i like want to keep go looking at it and then it has something good that i can actually go do the work and i don't feel i have that experience with opus that's really interesting okay so we've got i have a question is it okay then yeah like how is this up with remotion i know you've been playing with remotion a lot so how is gpt 5.5 with pre-motion i have so a thing that i do a lot as someone who like needs to make social assets for us but has no real design ability or sense is i like to point these models at designs our designers have already made primarily through the figma mcp or to screen record how our gm's products already work we launched a new version of sparkle with these models you can do a screen recording give it to give it to one of the models and say like hey make me a video with this storyboard and over and over again with the exact same prompts the open ai models just like take the recording and like zoom in a little bit it's like very very odd it's like the laziest possible approach and with the exact same prompts i find that opus and like all of the recent opus models have been this way for a few months they take this level of creativity and they they engage with you back and forth and questions for how to make it more creative that i haven't found it close at all i'm intrigued to see how much open ai catches up here because dan and i were messing around with the new image generation tool for youtube thumbnails and it was really good and it's really fast it's so good it's crazy i was hesitant to try it because i i'm like in my head i'm like this is the one thing the open a models can't do and right before i jumped on i knocked one out for the the the thumbnail for this live stream was generated on the new open ai image generation tool in like three minutes and i was i don't have to give it any notes see yeah like gpt image too is really cool can you show us austin yeah i think so let me because i i haven't seen it yet so i want i want to see it but yeah let me see let me see if i can find the the chats so you can see like the conversation and just like all it took cool all right i've got it here let me show my screen i would say all right cool can you all see this yep we can all right cool so what i gave it was two thumbnails actually that that dan had made and with gpt image with gpt image too so it has it has this baseline this is for the video we recorded before the can you just click on one of them i just want to show people how good oh yeah for sure yeah so this is this is one that dan made i think very quickly right i made this really quickly this is like essentially one shot with the new gpt image model giving it a thought giving it a picture of me and this is a little bit more massaging with gpt 5.5 and the new image model and it's like that's that's kind that's just crazy that it does that and it's very it's very good at like oh i want you to change this one little thing but keep the rest the same and it's also good at it looks like me which a lot of these models have not been that's been harder for them yeah it looks you this is i would say like a 95 approximation of our logo which sometimes these things are very bad at so i gave it that and then i gave it a few more just like screenshots of the video and said hey this is the text for the youtube video about to do give me three rips on this and it made this one which is i believe what we're using it also made this one which is like a little busy but really good like in the busy state it's like quite quite good typically these kinds of icons these things make would look absolutely horrible and they look pretty clean and it has the right number of fingers and that's not from a picture my fingers are not in there that's crazy and this was the this was the third one just off of that prompt i gave it and the interesting thing for me with this is that my my process in in cloud code with opus has previously been that i made a skill for like a youtube kit skill and i had to preload it with all this context where it has the youtube api it has the figmcp connection it has all this stuff so it can go look at the internal taste that we've created and what's worked for us previously and the analytics behind our youtube thumbnails and all this was was like hey we're making this video and then it creates this thumbnail and when we do these live streams and we do these vibe checks and reactions we have to move so fast we're doing so much stuff that the only way we're able to react quickly with real insights and have stuff that performs well is tools like this that get us from zero to one in seconds i love it i love it okay thank you austin that's amazing so we've got some really cool stuff coming up we've got we're gonna have mike who's our head of consulting in our tech practice talk to us more about how he uses gpd 5.5 for knowledge work and his vibe check on that the good the bad the ugly all that stuff we're gonna have some special guests joining so there's more special guests non-every guests i will not say exactly who it is but you'll want to stick around for that before we get there what you should know is you got to stay hydrated it's important to here mike you got to stay hydrated it's important to keep yourself and your agent liquid cooled you should also know that every is the only subscription you need to stay at the edge of ai and i want to show you a little bit about every in case you don't know who we are and what we're about we are every we do these vibe checks every time a new model drops we've got one this is our vibe check for today i want to show it to you look at this it's beautiful we've got a full team across coding writing design we we build tons of apps we're we're we're building when using open claw and what happens is we get access to these models so we got it we got access to 5.5 about three weeks ago and we get to put it through its paces and so you get the benefit of if you're in every subscriber of getting our take on day one from from our hands-on testing a couple if you're just getting here a couple things that are important to know about this model we found it's it's really spikes on on some of its senior engineering ability especially in this in the back-end department and especially if you give it a good plan so on our senior engineer benchmark which measures how good a model is at taking a vibe coded slop code base written by yours truly and rewriting that code base in the way a senior engineer would comparing it to real senior engineer code and on that benchmark gpt 5.5 scored a 62 out of 100 by comparison gpt opus 4.7 scored a 33 so there's about a 30 point swing there and so this model has some special characteristics that we have not found in any other model that we've tested a couple other things that are that are really interesting about this model it's it has that spiky senior engineer ability and it's also super fast and and and pretty good to talk to which is a it's rare to see a model improve on both dimensions at once i think you can you can see that a little bit with with opus 4.7 it got a lot better at programming it's it and it but it's also much more terse and a little bit more it feels it lost some of its empathy and i think open ai did a really fantastic job of making this much smarter but also keeping it pretty pretty personable and pretty fast so it's usable for the like day-to-day knowledge knowledge work tasks that that were that we're used to when we at every when we do these these vibe checks we always do a reach test which is are you using this every day are you reaching for it over other things and if so what are you reaching for it i'm a i'm a green on this model this is my daily driver for pretty much everything the only thing that i still use opus for is every once in a while when i'm not talking to my claw i'll i'll talk to claude on mobile i think it's i i really like claude on mobile mobile but other than that it's it's pretty much my go-to for kieran kieran is a yellow i think there's some things kieran you like some things you don't like but it's it seems like opus 4.7 is still really your your go-to model your your daily driver with using this this model for certain other things naveen you are a green you're usually an open ai stan so so know that but i think that you you just you love this you love this model mike is a yellow and katie parrot who's our lead staff writer is a green so it's not all greens but there's there's a lot of people psyched about this model and even if you're a yellow i think i think pretty much everyone agrees there's a lot of power here a couple things that we've so that we saw so again it's really good at doing that senior engineer benchmark it's really good at if you give it an idea for a big thing it's going to build like pushing that idea through over many rounds and many many tokens it's a little bit weak it's it's really good at some visual design stuff but it's a little bit weaker at other things so for example kieran did this benchmark test where he asked it to to design this karen do you want to give us just a little bit of a a refresher on exactly what we're seeing here and what you think it says about the character of gp 5.5 as a model yeah and it's also good to understand what it is the benchmark because there i also test the harness here so it is testing the the the codex cli harness and it's testing the plugin framework and it's testing compound engineering so there's a lot of things going on here it's not just a prompt in result out so it's also like how does gpt 5.5 work within the codex harness like what features does the codex harness have and that is also represented in here but what i've seen change from previous things for design is it's better at typography it's less fiddly like weird details that were in in the previous version but it's not as good at like visual so you can see that the the left side looks good with the typography but the right side looks it doesn't doesn't make any sense like there's like there does not really make a lot of sense what it even is or what it's trying to do there and that was better in other models but also i do see it improve in structure and maybe you can argue actually structure is more important and finishing things and not having any bugs and like staying on course is more important because these visuals are probably done by a designer or by an image model which they have which is very good so you could even argue that i'm not saying this is bad or good i'm just observing this changed and normally in these changes if you see them in different places you can see what they try to do and this is a great model but also i test the model within the harness so i when i say yellow it means yellow for also the harness in which i use it and the harness is is very good but there are other harness that also very good so also what what i'm what i'm what i like to do is i also recently like to try models against each other in the same harness in like cursor for example and and just see them head to head i will do more of that i remember last time i did a vibe check for a cursor i say yellow and now i'm using cursor every day so like me saying yellow on 5.5 means maybe that i need to adjust how i work or learn the powers or it just means i need to change something to see it work and i'm not saying i won't try it i'm actually running right now i'm running like three codex things in the background doing stuff so i'm using this model for sure and but it's a big thing to say daily driver and go super green so i'm very very enthusiastic about this model it's super good and i'm very excited that i can do things in one shot because that was never really the case with like older versions where it just stopped working or got trapped at some point i think 5.3 was the first time where i was like okay open ai you got my attention this is actually like very very very usable and 5.5 is just that but then way better it's like if they keep going this is great this is the best thing any engineer or creator can have we're just they're so good now that you yeah it's hard to compete and they're so good and you can all use all the models so really the vibe check here is for me like what is different and what are the differences between all the models it doesn't mean it's good or bad i love it okay so we've got about three minutes before our next guest mike taylor our head of technology consulting at every do you want to tell us you've been using this for some of some knowledge work use cases do you want to give us some some thoughts on on what you found in doing that i'm just going to open open up some of the things that you probably going to want to be looking at yeah do you want to open the dashboards yeah oh yeah let's do let's we'll do the dashboards okay yeah so this is a opus dashboard and you'll probably recognize that well maybe like introduce the benchmark first yeah yeah so the benchmark here is this is actually something that comes up again and again in every workshop i do can we make a custom dashboard from and this is something that used to take a lot of time a lot of it tickets to accomplish and now people are just vibing it and this looks vibed right this is opus and you can see you probably recognize the color scheme if you've done a bunch of vibe coded dashboards like if you just scroll down you've got like yeah it's like very very bright like dark dark background and you've got the purple at the bottom as well the vibe coded purple oh yeah the purple always yeah now if you switch over to the other dashboard and so this this was five five this is gpt five five right so it's like way better it looks more professional it looks more professional like if you scroll down it starts to get really nice the color scheme was like more tasteful this the spacing was better i actually like wanted to read it so i would say like this is the one this is the first benchmark i ran i was like oh my god this is like agi and and it's it's a little bit more nuanced after after that but just to point that one out because that's a like that comes up in every workshop and i would say like this is like a very very clear win and and it surprised me because i i just saw that like with opus it did the best powerpoint deck i've ever seen and the reason was that it was better at design and it had better vision than previous models so i wasn't expecting this weirdly gpt 5.5 failed the powerpoint test versus opus but it beat opus on the dashboard so it's not a clear picture that it's like worse at design better at design i think depends on what you're designing i love it thank you for thank you for sharing that we have now i'm very excited to welcome to the stage a very special guest romaine are you ready are you ready to be on the stage no he's already at mine because you are i'll take you off if you're not romaine is not ready okay so just just a little heads up romaine hewitt is the head of developer experience at open ai and he's joining us to talk to us about this model that's coming up soon and while we wait i'm going to share my screen again and i'm going to just say every is the only subscription every is the only subscription that you need to stay at the edge of ai we get these models way before they're released and on release day you can expect a vibe check like this from us where we use it for everything from coding to writing to design to marketing to strategy and this page is not loading but this page does load we've got a really really sweet long vibe check here it's several thousand words a lot of it's free a lot of it some of it's for paid subscribers but if you're in every subscriber you get a lot of other cool stuff so as part of every we have a bundle of apps that we build some of the apps are built by some of the the amazing people on this call on this live stream we've got quora which is an ai agent for your email and soon an inbox kieran i would love for you to show us some some inbox type stuff if you feel ready to do that we also have monologue built by naveen it's a smart dictation app and it we recently launched monologue notes so monologue started as you it's like a speech to text app like whisper flow or super whisper but we just added notes so you can transcribe a note on on your walk you can you can record a note in a meeting you can brain dump into it it's really great i love it and it's very available to all your ai agents so it's really easy to get your notes where you want them so we've got quora we've got monologue we've got sparkle which organizes your desktop we've got spiral which is a agentic ghostwriter we've got proof which is a markdown editor for your agent and we've got plus ones which is our hosted open claw so this is all available for every subscriber as you pay one price you get access to everything that we write everything that we make and all of the we do a lot of we do a lot of live streams and camps and stuff to teach you how we how we use stuff at the edge and we also have a pretty cool consulting practice where we work with big companies work with executive leadership teams to get their hands on get their hands on stuff like cloud code or codex and get them using agents and get agents into the org so if you have not checked out every you should go to every.to and you should subscribe and i think we're still waiting i think we're still waiting on romaine while we're doing that mike anything you want to add on some of the benchmarks you did or or maybe anything i haven't talked about on the consulting side yeah so maybe if you want to bring up those two files called persona one for each model so i'll tell you a little bit about this benchmark so last year i worked on the startup failed unfortunately but which is why i'm working for a living but but your loss is my game it's still going as a side project yeah we have we have some customers and stuff but but but it's now largely away from keyboard i have claude and codex working on it while i'm here but yeah it's it was a virtual focus group startup so we used ai to roleplay as your customers so i got very very deep last year into cloning people i we built a whole panel of 300 people that we cloned based on user interviews and then this is my clone just for ethical concerns don't share anyone else's clone but this is this is a clone of me just a snapshot taken just before i moved here to new york and yeah if you're seeing that on the screen here it's just a great benchmark because obviously we have spiral right where it's it learns to write you but the way that i write isn't the way that i talk and especially once it's gone through editorial it's a little bit different and i think it's maybe an easier benchmark to to to like write like an editorialized piece of content because that's more consistent it's more standard the way i talk is is really odd it has like inflections it has a lot more personal information so what this is is i i provide it with a transcript of this long half an hour interview i did and that's the prompt and then i have some extra stuff on top to say impersonate me and actually the first time i tried it it refused to impersonate me i think some alignment checks alignment is good and well but so so the first version failed on this bench but that that was quickly fixed and and it worked beautifully so i don't know if you have the this is the opus one is it yeah the opus one if you if you look especially if you scroll down to like maybe the second or third question it's it's it's trying to too hard like there's a lot of like ums and ahs oh mate a lot actually i think exactly maybe i do talk i like there's some some of that but not not all of that yeah yeah and i i think it's just like feels so eager to be me that it's almost like a super version of me it's really emphasizing my ticks whereas if you go to gpt 5.5 it's just much cleaner it's just like a much smoother experience and it picked out a bunch of things that i actually really do think or or like very very plausibly like could think which which is like as good as you can get really for this type of benchmark great so i i agree i think we found the writing to just be a little bit smoother good good at voice but not overdoing it maybe a little bit less ai ish and now i am very excited to welcome romaine hewitt and dominic dominic how do you pronounce your last name kunda but you can just say dom it's fine okay i'll say i'll say dom welcome both of you can you introduce yourselves yes thank you so much for having us this is roman i lead developer experience and i found dominic join me who leads a lot of the dx efforts for codex and so we're both very excited to talk to you today i'm excited we're psyched we've been we've been using this model for the last three weeks we love it it does some special things that we've never seen a model view before what like tell us about like what what you've seen internally and what you're excited about for this model it's pretty amazing how much now like the work has changed at the point i like relying on this model i think we we've seen like across the board how much the model is smarter at any professional task right we talk a lot about writing software and of course we hear from engineers that the the model is writing better code it's like more thorough when it comes to like analyze analyzing a repository of complex code bases but also takes on like more and more of the other tasks that we do every single day and and i think what's very exciting to me is not just the model itself like gpt 5.5 but it's also the combo with codex because when you have such a great model you also have to have the right software and environment to to make it shine and and we're seeing like so much of that with codex right because codex can help the model not just like figure out the files and the context but also access all of the tools that it needs to really carry the work check-in output and and really like deliver things that you can just ship right yeah i i i feel like 5.5 has especially in a combination with like the features we shipped last week in the codex app and some of the features we shipped today has drastically changed how i work like compared to even just a few weeks ago i know that i can have codex handle everything from like the documentation to i i think i've built like 50 demo apps in the last week just like trying to push codex to its limits and like like there's surely no way that it can do it and it just like just hammers through it like last night i think it was like 11 pm or something i was like let let me see if i can like one shot like a full 3d modeling software and it it fully did it i i remember one of our colleagues was like well can you add like download like support for like the file type for like 3d files and i was it already had it it was already there so like the level of like attention to detail that like the model puts in is is incredible dom do you have any of those demos that you can show i don't have the 3d modeling one but i do have some let me share my screen here sweet very impromptu demo by the way so we just like got out of the launch room flipping all the switches publishing the blog post and we just showed up so amazing i we're doing it live we're doing it live that's what we like that's how we do we want it we want it like this okay don you're live we've got your screen this was one of the one of the moments where i truly had like a mind-blowing moment and it like the final result is actually in the blog post as well if you want to take a look at it later but my colleague sent me this picture right like this is just a screenshot of some software that he had been building and it's mostly like for like illustrative purposes so i thought like okay let's take this and i gave it to five five and i simply told it i had a couple of other folders in the project that i didn't want it to deal with so i totally ignored those but then use webgl to and like the real data from the artemis 2 mission to actually implement this and it went ahead and like fully built this as a working app and it went to the it cited its sources it actually went to nasa downloaded like the flight data of the artemis 2 mission and then rendered all of this and some of the attention to detail that originally i was thrown off by was it had if i open this in our built-in browser here it had this like slightly weird curve where i was like why is this what like this doesn't look fully like how the data like how the curves looked when i saw like animations from other people and the result was actually that it saw in my image that i had given it that these two planets were like scaled like this so try to like replicate the scale but that's not real scale in space so it had to like take the real data trajectory and then like fix it for those for those moments i had i asked it later on to like add a true scale moment so like here you can see the actual like slingshot and then how it goes over to the moon and if we go towards the end is also the flyby which is we have like the full flyby and like how it like sends the moon around it and like all of the data it tells me here like what were like the mission moments and stuff like this was like one of those moments where i'm i i didn't give it a lot of instructions on what i wanted this to look i just wanted it to look like the picture and it was like cool i got you and so i've been experimenting a lot with this and so one of the things that we silently launched because i definitely want you to try it but i didn't we just like shipped this is inside our like build if i find it build web apps like build web apps plugin if you have that already installed or if you install it now we added a new skill here called front-end app builder and this one actually uses now the the design capabilities of image gen that we launched this week together with like five fives attention to detail to build apps for you so i told it earlier like right before this like gpt5 is out here's the blog post but like a minimalist well-designed page and it went in and like this is one shot like the other things is me just asking to start the app but this is the app that it built wow with like the data from the from the website and the design if you look at it is like pretty spot on with minor details of like image and generating the individual assets later on in slightly different positions but it's pretty spot on for like a one shot that's really interesting so is this like is this a new workflow for doing front-end design is like have the image model like render a website and then have 5.5 turn that into html and css yeah it's a great one we've heard like from the community that one of the things that we still have to improve and push the frontier on is really like the front end of the taste capabilities but now and we're still pushing on that very hard but at the same time like with imagine 2.0 the quality of the designs that you can create is pretty outstanding but also gpt 5.5 is so thorough at taking an image input and implementing it that you can now iterate on your design first and when you're happy with it just get it done with gpt 5.5 to implement it like this one which is pretty magical yeah and i think i think one of the interesting things there is it really brings together a lot of the capabilities that the model has gotten better at like one like the attention to detail but then also it went through the entire like like loop here so it used one of the things we added today as well is browser use inside this in-app browser so it actually controlled the in-app browser to like take screenshots try things out and then it would like do checks here on like does it work on mobile does it work on desktop take screenshots and like verify that work and really put in so that attention to detail before it felt comfortable with like all right this is good enough now that's really the magic of bringing like the codex harness the codex surface and all of these tools right i think if the model does not get everything right from the first try you don't even have to interrupt it because it checks it won't work like now with computer use outside of the codex app but even inside the codex app you get the model to just like keep on iterating until it's satisfied with the answer that's really interesting what i what i'm really curious about is i feel you guys are on this crazy arc this crazy run right now where if i think about codex three months ago or six months certainly six months ago it was like very there was no desktop app and it was very like senior engineering both in terms of the the user base and what it was like to talk to it was like super slow it would get very very technical answers austin who's our head of growth was on the user base and what i was saying was saying it just made me feel stupid and i feel i feel like over the last especially the last three months kieran was saying since 5.3 which i agree with you've been on this run where like every release it just it's gotten smarter it's gotten more personable and and really good for senior engineering work it's it it's tops our senior engineer benchmark for what it's able to achieve but it's also just really good for just general agentic knowledge work so tell us about that shift and and the trajectory you're on yeah it's it's almost something that has been like very surprising to us in a great way it's the like even at openia internally the adoption of the codex app right it started with us being engineers and building with it and of course as we started with the cloud and then the cli and then the id extension so we had like different manifestations of the coding agents but we always believed into this like agentic delegation vision and we truly needed a surface with the codex app to like achieve that that that vision of like being able to send complex tasks and have the model handle them but what's happened in the past like few weeks even at open ai and outside of open ai is like not just the best engineers in the world like moving to codex and taking on this like agentic delegation workflow it's also more and more of the like adjacent teams adjacent roles like actually using the codex app for any parts of their work like dom and i have i've completely changed the way we work at open ai like most of us here like not just for building software but also to connect to like your slack and your notion with you can show the plug-in screen but we have now like like 100 plugins where you can connect like all of your tools and all of the context to bring into this codex app and as of today also with the launches we also introduced like the artifacts so the ability to like hey if you're working with spreadsheets or with if you're working with like any documents and files or if you produce like assets you can now visualize those in the codex app so it's really turning into a very advanced like productivity app for much more than just the developer yeah go ahead i was just gonna say i think one of the interesting things is that like as i mentioned like over the last couple of weeks like the way i work has drastically changed like with these launches like especially we had a big launch last week we had a have a big launch this week and we had smaller ones along the way right like image gen was certainly a big launch but then we had things like chronicle launch and research preview which inside codex like can watch your screen so you can learn more about how you're using other tools and all of these different capabilities have really helped me like stay on top of all of these things where like the amount of times that i would kick off a task because of some feedback and i wouldn't even actually go in and describe in like long words what i need to get fixed out like often a prompt would look much more like like go like go look at like the last feedback that like roman sent me on slack and like go fix that and send him a send him a message when the pr is up and like that's the that's the full prompt and five five is intelligent enough to like figure out with like my memories and my context window like which and my like chronicle history like which roman am i talking about it knows to use slack to like pull that message and then like knows possibly like what are the docs that are related to the to the launch and then actually like what uses an automation to like keep a track to keep track of the pr until the deploy preview is live and slack it back to him and i i just send off one message and it's wild for me to think that it can just do that in like a message yeah and i think what's really magical too with the launch of last week to put things in perspective is that now you have the codex app but the codex coding agent historically could only manipulate like your files in the repo for instance and then connect to some skills or plugins to bring that context but there's a lot of work you do on your computer that is not always just an api or a plugin and now if you combine like chronicle as a research preview we like which understands a little bit how you work on your computer but also computer use i don't know if you all watching this i've tried it but i really recommend trying it because you can't really describe until you experience it like how good it is because to my knowledge there is no other computer use implementation out there where you can simply ask a task to codex and then it would kick off its own like app in the background and then its own cursor and you can still use your computer and so you can still like send like multiple of these tasks in the background for those that don't have let's say an api or a plugin the model can actually figure everything out including your your native app on a mac let's say like the reminders app or the notes app like anything that you use on your computer codex can also drive so now if you bring all of these things together you you have this like combo of codex and gpt 5.5 to really drive any task that you do in your days which is pretty pretty outstanding really cool kieran ask a question yeah i i would love to it's it's really cool these examples and this inspires me to like actually like really kick the tires and clearly the model better than me since you used it even more than we did and one question is with these examples which are very cool like can you also if you are allowed to share more concrete examples like what you do in like actual day-to-day work and how you use 5.5 because yes this is a cool example but like do you ship codex or do you like use harness engineering or like maybe maybe harness engineering changed with 5.5 i would love to hear some of the like the inside stuff that you can share around these yeah i think one of the i'm like trying trying to think back of like the last couple of weeks of like really leveraging 5.5 whether like some of the standout examples but i think one of the one of the parts is right now i've spent a lot of time improving like the codex documentation developers at openair.com was like one of one of those examples and we had this on what like sunday night before we launched chronicle the next morning where i thought it would be helpful for people to understand the impact that chronicle has win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win pull it into the developer's website, which is an entirely different code base, implement it, and then like read through the docs that I had drafted in a Google doc to figure out how it actually has to like document these examples and like create them. And all of those things were just like running in the background. I think like one of the interesting things is that I can contribute way more features now, even like smaller delightful features to the app or like try out an idea in the app. Not all of them necessarily ship, but like inspire people inside the team without necessarily doing a lot of context switching, if that makes sense. And I was chatting like on top of that, I was chatting with some early developers you, Dan, who have had access to the model for like a couple of weeks. And like the things that came coming back in those chats were, I am not even specifying like much of the task anymore. I just trust the model much more to figure things out, like navigate large, complex code bases and like extract what's needed, solve complex bugs that were still hanging around, but now like five, five could solve them. And to your question, Kieran, on harness engineering, I was chatting with Will at an engineer at RAMP and they're using their own harness. So I was curious how that worked for them. And he told me actually like complete plug and play, like from GPT 5.4 to 5.5, but better yet, 5.5 started to discover new tools that other models could not discover. It's like all of the sudden it realized how to access the database, how to fetch data. It was complete plug and play, but also like new magical powers they did not have with any other model before. I love it. Thank you guys for sharing. I know you only have like a minute left, but a big question that we have not gotten to yet is GPT 5.5 is out right now in codex and in chat GPT, but it is not available in the API. Can you talk about why and when we can expect to be able to use it in the apps we're building? Yeah. We are taking like a very,, safe approach for this launch and we are very, very eager to give like 5.5 to everyone. And so we expect that to come extremely soon, hopefully in days maximum., we're just like making sure that like the rollout goes well and that we have all of the safeguards in place for everyone to benefit from this model. But again, back to our OpenAI mission, we want to make sure that everyone can put their hands on this model. Awesome. Thank you guys so much for joining. It was a pleasure to chat with you and thank you for the model. It's great. Thank you for having us. Thank you for having us. Enjoy GPT 5.5. Of course. Thanks. See you guys. All right. So that was the OpenAI team. That was pretty cool. Kieran, Naveen, Mike, any thoughts, anything you want to share? And Laura, welcome. Laura is a staff writer at Every. Yeah. Thank you. Thank you. Laura is a staff writer at Every. Thank you. I'm not ready at all, but thank you. You're not ready. I will take you off then. I'll take you off. So Mike, Kieran, Naveen, any reactions to the OpenAI team? This is why I come in the room, Dan, so you can't surprise me. Yeah. I really want to see more computer use stuff. I think that the Chrome extension for Claude is one of the main reasons I use the model still as well. So I haven't had a chance to really kick the tires on that. But it's also the thing that all of my enterprise clients are most scared of. Yeah. Yeah. So it'll be really interesting to see how that rolls out. Which is an interesting sign. Like if people are scared of it, it might actually be the most useful thing. I feel like one of the things that held OpenAI back for a while is that they were super scared of unhoveling the model and just like letting it do stuff on your computer. And now it's just, oh yeah, it just does stuff. So that's really interesting. I love, I love this. Okay, you're going to do an image generation with this sick image model and then we're going to turn it into a front end thing because it totally short circuits the issue that that we have with, okay, the taste of this model isn't as good. And so they don't even necessarily have to improve along the dimension of can it come up with a good design on its own? They just like let the image model team do that. And I think that's amazing. Yeah. Yeah. On that one, it's interesting because I have a benchmark where I have an image and I say generate the website for this image, but it was so bad. It just put that image as a background. And that's what 5.5 did. So I'm just, I'm, oh, me too. But that's what most more. Are you sure it is thinking on? Because I've had that happen. Yes. Yes. Extra high. Yeah. It's on, but it means that it's not always the model and it's also the harness and how you prompt. So for me, this is again, hey, can I rewrite that in a way or like improve the harness to see, to get everything out of it. And that's why it's so great to talk with OpenAI because they have the best knowledge. we played with it so we can share our opinions and findings, but also they, if they can show you, hey, this is possible, then, then we can now fine tune how we do it and learn from that, which is, I would love to see this evolve into a daily driver by just learning how to use it. That's my take on it. I really want to try my PowerPoint again now. Yeah. Yeah. Yeah. Yeah. Yeah. Yeah. Could smash it. I, I, I, I think the image, following the image is something that I observed, it has improved a lot because in the past it didn't do it. I have a couple of examples that I did recently, let me show my screen. I, I was quite shocked how well it's able to follow the images so i've had an opposite experience so daniel who is our good designer for monologue and every he gave me this design and this is the design oh sorry i think let me show you the old design that we have this is the old design right this is completely wipe coded it didn't fall and then this is the final design can you see any difference between the daniel's design and this one wait show us show it to us side by side if you can so that's a screenshot of the actual developed the actual app yeah yeah you can see it's a local host one right yeah yeah yeah but show us the the original design that i was working from yeah this is the original design oh how can i let's see if i can actually open this yeah sorry maybe let me just put it there yeah i'm now i think so it just did it was just a one shot thing or or what oh so there's a good harness here so i actually used it asked it to take screenshots so it started the web server it used playwright browser and then actually it said browser binary is missing because i never actually used codecs for building web apps because i always thought it's bad at it i reach out to cloud code for this task but i tested it it actually downloaded the playwright i think somehow images are not loading here but it just took ton of screenshots you can see the whole thing is changed now and then i started giving some feedback it pushed it and then i gave some feedback feedback and then finally i think it's the assets so i gave a lot of like here it's not there i really love the comments thing codex has this in-app comments on the ui directly so i think i did couple of iterations to and fro but finally it's like all done in off an hour or something so that's very cool yeah all right it's like following the images really that's really cool so if you just got here we just had the open ai team on they shared a lot about a lot about gbt 5.5 and and we have found that what they taught us is you can use the new image model to make front-end designs and have gbt 5.5 implement them i think that's really awesome now we have representing writers from all around the world we've got katie parrot staff writer at every katie wrote the vibe check so if if you saw the vibe check that we published today which is around here on one of my tabs anyway this is the vibe check look at that look at that see there's katie parrot is that's her name that's because she wrote it so katie thank you so much for writing this for doing a lot of the testing i would love for you to walk us through what you found when you use this product for writing and i'm going to move over move our screen over to some of the stuff that some of the stuff you found yeah totally so just to set the scene a little bit i have been i haven't touched open ai models for writing in almost a year around the time that i think it would have been sonnet 3.5 i've dropped i just found the writing quality on claude models to be so much better that i just didn't feel the need to use chat gpt and in fact i actually did what i thought was breaking up with chat with chat gpt i canceled my personal subscription about three months ago and then and honestly when i first started doing the tasks for this model i was like okay nothing new here i don't need to change anything it really wasn't until i got my hands a little bit more dirty and like went back and forth with it and saw how it took direction that i was like wow this is really good it's so some of the things that we found and mike can speak to his experience too is just that it's it's it's smoother so it like the logical progression from idea to idea flows like flows more naturally it it's less clever i think the that's something that i found with opus and if you scroll down to the third bright pink you can't read it on the screen but yeah if you go in there and zoom really close you can see that this is the gpt551 and at the end of q1 i sat down to review my okrs and discovered that i feel i have roughly half of them i first read that i was like well there's not really much there so that's the thing that i might want to like make more clever and more neurotic but you see and like and i've just told that like capture more of my neurotic energy please and and it does that i can't show you that because it might reveal some details about secret information about the model but yeah and like and and for me like a really big thing that is that's important to highlight is the speed because when you're writing with ai you're you're not going to get it's not going to one shot anything you're going to go back and forth you're going to be like this i like the direction of this but it's not quite right like this hook isn't the right hook and so you want that speed of that and that ability to take feedback and go in the direction that you want with it and i've just found that gpt55 is faster and better at picking up on what when i give a direction even when that direction is like extremely extremely poorly phrased and very like minimal in terms of like direction which is something that we found with opus 4 7 is that it needs a lot of specificity in how you prompt it in order to follow your directions so i find gpt55 to be a more intuitive model as well so are you breaking up with claude i am probably i am thinking about breaking up with claude the beautiful thing is because i work at every i i have subscriptions to both thank you every but for my personal use i will probably bring my subscription level down just because i don't see myself using it as much mike what do you think yeah for writing in particular i thought this is a very dependable model like there's a lot of things that i write where i just need to get the job done and and i i just want to rely on it to not say something stupid and and it does a great job of that i still prefer opus writing just because i like the neuroticism element it is it's very witty right opus is very witty and and but the problem with opus 4 7 i found is that it would say something clever and then just stop and it's like very i'm i asked for like a section and you've given me a paragraph almost like again anthropomorphizing the models but but it feels it's like great i've said something clever now i can like move on to the next thing whereas this is a workhorse is is the good is a good term for it i think you guys picked a really good title because it was probably katie deserves the credit there but but it is a workhorse it just like gets the job done without fuss it just churns through the work and for most writing tasks i don't need like witticisms i need just like job done katie any any reactions to what mike just said or any anything that you found in your testing that we have not talked about yet i think i think like let me see i don't know that's okay if not i can yeah no i think that that's more or less caught it's so hard to talk about writing it's like up short of just reading out the prose that it produced it's doesn't make for good streaming but what i will say is that i did write the the vibe check itself with with codex with with gpt55 and that was an awesome experience not just because of the model although well this may actually be about the model's capabilities the slack integration being able to pull information straight out of slack and when dan did what dan does and zoomed in at like 5 p.m the night before the review had to go out to flag an issue and talk it out i was able to pull the additional messages talk to to codex about what needed to change in the in the draft and get like new sections for dan to be happy with hopefully it was great yeah they came out decently most of it made it into the into the doc or into the published piece and so that the workflow like having that really powerful model sitting at the middle of all these integrations is something else speaking more to the general knowledge worker angle which writing nests inside of i just feel so much more resourceful sitting inside of this this ecosystem which now also includes g gpt agents which i am also extremely excited about have you tried i have not tried yet i so i built i've been working on a project to get ai to manage my okrs for me because i failed at half of them last time and i was like can open like and i have an extreme weakness at project management like can't do timelines can't do dependencies can't do milestones can't do any of that so i just gave my okrs to the model and was like can you help me do this and so i i i talked to the i had the agent i had an agent build me a project manager and yeah it connects to my notion task list it connects to my calendar so i can just say hey what should i do today and it will tell me what needs to ship which is great for my brain and hopefully really cool right well vibe check of the agents coming soon but we are we are approaching time here folks we've been streaming for about an hour and a half if you just got here you should know every is the only subscription you need to stay at the edge of ai we have published a detailed several thousand word vibe check of our three weeks of testing written by katie parrot the overall take that we have is this model's a beast it's really good it it is both a step change in senior engineering capabilities and it is quite good at just being a fast personable like working collaborator which is a pretty rare thing to like increase your efficacy along both of those dimensions at once in a model it scored a 62 on our senior engineer benchmark which measures 60 62 out of 100 which measures how well a model how good a model is at cleaning up a vibe coded slop code base in a way that a human senior engineer would it scored a 62 opus 4 7 scored a 30 32 33 ish so it's like a 30 point swing which is pretty it's pretty crazy one caveat the on the senior engineer bench the 5.5 scored the best when it used an opus 4 7 plan so so katie may be breaking up with opus or or with claude but if you really want the best get the best out of your coding capabilities i actually still think you want both of these models in the mix but anyway this is a really good model if you're into it you should definitely check out our vibe check every every dot to slash vibe check slash gpt55 or you can just go to every.to and it's linked in the home page every i said every is the only subscription you need to stay at the edge of ai you should subscribe at every.to slash subscribe we do ideas apps and training on the idea side we have articles like this every single day we have a daily newsletter where we keep track for everything that's going on in ai you just read one thing and you are absolutely caught up you get all of the stuff that we're learning as we as we build companies with this as we design as we write as we code we also have a suite of apps that you get as part of the every subscription it's apps that we build for ourselves to help us do better work in ai that we include for you in the subscription so we have spiral which is your ai writing assistant with taste we've got quora which is an ai email assistant it's an it's an agent for your email and when it's coming out with an inbox soon so you can i will be doing all my email through quora very soon which is great sparkle which is a file organizer you can see my desktop is completely empty that is technically my second desktop but my actual desktop is also empty and that's because sparkle organizes it for me we've got monologue which is a speech text app it's like super whisper whisper flow it's built by naveen it's it's really really good it's super fast we also just release notes for it so you can you can record any note you want a voice note on a walk an idea storm a brainstorm your meetings they're all transcribed and all available for your agents we've also got proof this is the app that i vibe coded a couple a month or two ago that we had to rebuild because it was this vibe code was slop which became the senior engineer vibe code benchmark proof is sick it is a markdown editor for your agent it's all web-based so it's really easy to share a markdown document between different agents or between you and your colleagues so like coding plans and stuff like that so you should definitely check out proof proof editor.ai also plus ones it's our hosted open claw one click it's in your slack all of this is available as part of your every subscription every.to subscribe all these app plus all the articles we write plus we do a bunch of trainings we do live streams we do camps we do courses we do all that stuff ideas apps and training all part of one subscription we also have an enterprise offering so the the the base every subscription is really intended for individuals and teams but if you're part of a part of a larger company a big startup trying to trying to go ai native or agent native we've got you covered mike do you want to talk a little bit about the consulting offering yeah sure yeah so so i'm a recent addition to the consulting team just joined in february but actually i've been writing for every for the past couple of years and before i joined every i like to joke that i invent reinvented every from first principles i was playing around with ai i was writing about it and then i was consulting so now i'm just doing it as part of a bigger team and more professionally hopefully natalia will tell you it's been to change i'm wearing a shirt now i've had a haircut no it's it's not that corporate here no but yeah in the in the consulting team what we do primarily is ai transformation adoption so quite often a company will come to us and say we just bought cloud code for this whole team how do we get them to use it and how do we get them to be effective with it in a way that is like safe and matches like what we what we can do internally any restrictions limitations so it's hands-on hands-on tactical stuff we're usually doing at least 50 of those workshops are like like actual actions you're actually building stuff in the workshop so that's that's a big focus for us i think you can only really learn this stuff by playing around with it and we obviously have a little bit of theory to to go alongside that too one of the things we've been working on more recently is trying to take the bundle concept the membership concept from the the rest of the business and then thinking about what would subscription look like for consulting because time of materials doesn't really like make any sense anymore when my agent can be working overnight on something for you so that's something we if you're interested in that we're trying to explore that at a minute we're talking to four or five companies about what that looks like so yeah happy to have that conversation so if you are if you want us to come train your executive leadership team if you want us to come set up agents inside your company where to find us otherwise you should subscribe to every and and if you don't that's okay we'll be back the next time there's a model release so you'll definitely see us anytime there's a new model check out every we'll tell you what we think of it it's it's always a pleasure to do these i i love model release days honestly super fun i never eat lunch because i'm too excited and so it's like 3 45 here so i'm gonna probably eat some lunch but what you should know is you should always make sure to stay hydrated cheers keep your agent liquid cooled not hydrating i saw that coke zero katie that doesn't i was looking for my water bottle and i couldn't find it i got it i'll be hydrated next time all right thank you all for joining check out every and we'll see you next time