Every

LIVE VIBE CHECK: GPT-5.5 Has it all

914 summary words 4 min summary Watch video

Start with the signal

4 min read

Summary

LIVE VIBE CHECK: GPT-5.5 Has it all

Main Topics

  • GPT-5.5 Release & Testing: A comprehensive three-week testing period by the Every.to team across coding, writing, design, and knowledge work
  • Model Comparison: Detailed comparisons with Opus 4.7, GPT-4, and Claude models
  • Senior Engineer Benchmark: Performance metrics showing GPT-5.5's strengths in code refactoring
  • Codex Integration: New capabilities combining GPT-5.5 with OpenAI's Codex development environment
  • Workflow Applications: Real-world use cases from developers, product builders, writers, and growth professionals
  • Image Generation Integration: New workflow using GPT Image 2.0 to design and implement web applications

Key Points

Model Performance

  • Senior Engineer Benchmark Score: GPT-5.5 scored 62/100 vs. Opus 4.7's 33/100 on code refactoring tasks (30-point improvement)
  • Best when combined: GPT-5.5 performs optimally when using plans created by Opus 4.7
  • Speed advantage: Significantly faster than Opus 4.7 for knowledge work and execution
  • Execution capability: Unique ability to execute large-scale code rewrites without getting distracted by existing code

Strengths

  • Reliability over creativity: Described as a "workhorse" model—dependable, thorough, and consistent
  • Personable & collaborative: Better tone than previous OpenAI models; feels less condescending
  • Context management: Excellent at maintaining context across multiple code bases and long conversations
  • Attention to detail: Can implement complex requirements from minimal specifications
  • Multi-modal capabilities: Works well with image inputs and can follow visual design specifications

Weaknesses

  • Design inconsistency: Some visual design elements less polished than previous versions (though typography improved)
  • Language-specific limitation: Not as effective with Ruby or non-TypeScript languages
  • Design creativity: Less creative in visual design compared to Opus 4.7; better suited for structured tasks
  • API not yet available: Only in ChatGPT and Codex; API access coming within days

Use Case Applications

  • Senior engineers: Code refactoring, complex bug fixes, large codebase analysis
  • Product builders: Full-stack development, one-shot app creation, iterative feature building
  • Knowledge workers: Campaign planning, research synthesis, meeting notes analysis
  • Writers: Content creation, editorial work, faster iteration with direction-following
  • Growth professionals: Dashboard creation, analysis, marketing plan generation

Notable Quotes

> "It's a very good model. When I ran it in 4.7, I was like, 'Oh, this is the best coding model out there.' And 5.5, GPT 5.5 scored the same on that specific benchmark."

— Kieran Klassen (CEO/Founder, Every/Quora)

> "For tasks where I know I'm not going to be able to pay that much attention, and I just need to get it done, I want to make sure it's safe. I can delegate it without concern. I think this model has it."

— Mike Taylor (Head of AI Tech Consulting, Every)

> "This is a workhorse. It just gets the job done without fuss. For most writing tasks, I don't need witticisms. I need the job done."

— Mike Taylor on the model's writing capability

> "The bottleneck is not one-shooting things because that's working. For me, the bottleneck is how do you collaborate as a human with the model?"

— Kieran Klassen on the future of AI development

> "I am not even specifying much of the task anymore. I just trust the model much more to figure things out."

— Roman Hewitt (Head of Developer Experience, OpenAI) on internal usage

> "The level of attention to detail that the model puts in is incredible."

— Dominic Kunder (Codex DX Lead, OpenAI)

Takeaways

For Developers

  • Adopt GPT-5.5 for execution: Best choice for implementing well-specified tasks, large refactors, and multi-file projects
  • Combine with Opus 4.7: Use Opus for planning, then GPT-5.5 for implementation for optimal results
  • Use Codex harness: The desktop app and plugin ecosystem dramatically enhance model capabilities
  • Leverage computer use: New screen control features enable autonomous task completion across native apps

For Product Teams

  • One-shot capability: Minimal specs needed; the model can figure out complex requirements
  • Rapid iteration: Speed enables fast feedback cycles within meetings or during other activities
  • Image-to-implementation workflow: Use GPT Image 2.0 for design, then GPT-5.5 for implementation
  • Context integration: Connect Slack, Notion, and other tools for informed decision-making

For Writers & Content Creators

  • Faster iteration: Superior at taking direction and refining work based on feedback
  • Reliable output: More dependable than previous models; less prone to unexpected creativity
  • Integration benefits: Slack integration enables real-time feedback incorporation during writing
  • Breaking from Claude: Worth reconsidering for those who switched away from OpenAI

For Growth & Business Professionals

  • Knowledge work: Viable alternative to Claude for dashboard creation, analysis, and planning
  • Speed matters: The performance difference makes it comfortable to use interactively during meetings
  • Automation recommendations: Codex now suggests and implements automations automatically
  • Context awareness: Can synthesize information from multiple sources for comprehensive planning

Strategic Observations

  • Model trajectory: OpenAI has made GPT models viable for non-engineering knowledge work—a significant shift
  • Ecosystem matters: The model's power is multiplied by Codex, plugins (100+), and integration features
  • Implementation challenges remain: Design-focused work still benefits from Claude; language support varies
  • Next frontier: Focus shifting from "can we build this?" to "how do we collaborate effectively with AI?"

For Every.to Subscribers

  • Access comprehensive vibe checks on all model releases with hands-on testing across multiple domains
  • Get detailed benchmarks and real-world examples before making adoption decisions
  • Leverage consulting services for enterprise AI transformation and adoption
  • Use integrated app suite (Spiral, Quora, Monologue, etc.) optimized for AI-native workflows
Full transcript 18177 words · 86 min read
0:01

SPEAKER_03

Hello, everybody. Welcome to Vibe Check Day. We've got a new model coming out. It's actually, it is out. GPT 5.5 is out. We've been testing it for the last three weeks or so,

0:21

SPEAKER_01

and it's really fucking great. It's really great. It's a very exciting model. I've loved it. It has become my daily driver. So what we're going to do on this stream for you is we're going to go through our Vibe Check, which we just published on Every, Every.to. Every is the only subscription you need to stay at the Edge of AI. I'm going to go through all my findings. We're going to have other people from the team who are going to come on. Mike Taylor, come on over. This is Mike Taylor. This is good. This is my best idea. Why don't you take that chair? So we've got me. We've got Mike Taylor, who's our head of AI tech consulting at Every.

1:07

SPEAKER_01

Mike, you're not quite in the shot. All right. Great. And we are going to go through this model for you. So what I'm going to do is I'm going to share my screen. And here we go. Here we go. Yes. So GPT 5.5 is out. It is live right now. They are not releasing it on the API quite yet, but it should be out in your codecs, in your chat GPT. And I assume that will start rolling out over the next couple of hours. They're holding it in the API for just a bit because it's just a very powerful model and they're doing a lot of testing on it, which I think is probably good. It seems to be the new standard for these models. Like we seem to be reaching a level where people

2:01

SPEAKER_01

are like, oh, this is kind of worrying, but okay, here we go. Here's our vibe check. We just published it on every.to slash P slash GPT dash five dash five. And our headline is GPT 5.5 is OpenAI's Workhorse model. What's really it. So it is out. It is out. There's a, there's a blog post on their site. You, you should be able to see it. We, what we, what we found in testing this model is it's actually quite rare for a model to be really good at both senior engineering type tasks and a really good workhorse. And this model is that it is super collaborative. It's super fast. It's actually like

2:58

SPEAKER_01

pretty personable. And it did the best on our senior engineer benchmark of any model that we've tested.

3:05

SPEAKER_04

So on our senior engineer benchmark, which tests how good a model is at rewriting a, an existing code base in the way that a senior engineer would benchmark against real on a real code base that

3:18

SPEAKER_01

two real senior engineers rewrote separately. Um, GPT 5.5 scored a 62 as its best score. Opus 4.7,

3:27

SPEAKER_04

62 out of a hundred Opus 4.7, um, uh, by comparison scored, I think its best score was a 33.

3:34

SPEAKER_01

So there's like almost a 30 point swing between Opus 4.7 and GPT 5.5 on this kind of task. There's an interesting caveat though. Do not throw out Opus 4.7 yet because our, uh, best, the best performance of this model of GPT 5.5 comes from, uh, using a Opus 4.7 plan. So if you use them together, they get super powerful. Um, codexes are GPT 5.5 is the model you want to be coding. Uh, but you actually at Opus 4.7 plans, I think are actually still better than, than, than 5.5. But, uh, when you actually put it on this benchmark, it is like miles better at this kind of task. Um, so we have a full, we have a full vibe check here. Uh, it's, uh, it's on every, every.to slash,

4:28

SPEAKER_01

uh, you should be able to see it. Every.to slash p slash vibe slash p slash gpt dash 5 dash 5.

4:38

SPEAKER_03

Full vibe check is here, um, published already. Um, it's a totally new pre-trained model. So this isn't a fine tune of a previous, a previous GPT 5 version. It's the spud model that you've been hearing a lot about. It's super powerful. Um, but there, there's sort of like

4:57

SPEAKER_01

a mix of reactions here. So it has absolutely become my favorite, my daily driver model. Um, it's, I use it for everything from my day, my day-to-day work to like real engineering tasks, but other people on the team have different reactions. Mike Taylor, who's our, um, head of, uh, technology consulting at every has, has us had a slightly different, um, experience with it. Mike, do you want to, um, just talk a little bit about what you found in testing this model? Yeah, sure. So, um, I was also very impressed. I think it's the most reliable model that tested. Uh, I felt like comfortable and safe, like getting into a Waymo, you know, it's like, uh, it like feels

5:34

SPEAKER_01

pretty good, but I still think like Opus is the Tesla, you know, it's like a little dangerous. Um, and, uh, when I'm micromanaging tasks, I, I still use Opus as my daily driver, like if I'm heavily involved, uh, but I would say for, um, tasks where I know I'm not going to be able to pay that much attention, uh, and I just needs to get done. I want to make sure it's safe. I can delegate it without, you know, with peace of mind. I think like this model has it. Um, and give me, give me an example of like the kind of tasks that you might want to, uh, use this model for that, uh, and we just got Kieran Klassen, Kieran GM of every, of Cora at every, and, uh, the creator

6:15

SPEAKER_01

of compound engineering. Kieran, welcome. Hello everyone. Um, uh, so Mike, just give us, give us an example of a task where you're like, this is actually a really big difference. I might, I'm, I really want to try this today using this model that you might not realize. Yeah. So one thing I've been using it for this week is, uh, creating the curriculum for training material, uh, because this is, uh, quite an onerous task. Uh, like you have, uh, you know, tons of call notes from different people across the organization. Uh, like we're, we're working on AI adoption, uh, with a few companies trying to get, uh, the teams on blocks so they can use more AI.

6:50

SPEAKER_01

And, uh, it just requires a lot of diligent work of like, we need to go through all the call notes, make sure that like, uh, all the different themes and issues that we found in those notes have represented in the curriculum. And, uh, it like hasn't failed on that once. Uh, like I've been, uh, and, and, and also I would say that with, with Opus, uh, I was using Opus for that previously. And with Opus, I always felt like I had to like go back through it line by line. Uh, and it came up with some really sharp stuff and like some of the titles were like, cool. I'm like, oh, I'd love to

7:21

SPEAKER_01

teach that. But, but actually like, you don't want cool titles all the time when you're doing corporate training, right? They actually want something dependable, reliable that like, you know, regular people will find accessible. So I would say like if I was writing marketing copy, it would be Opus probably. But, uh, if it was, uh, something where I need it to be, uh, like, uh, unoffensive, I need it to be, um, reliable, I needed to capture all the main notes and not be like a little bit wild. And, uh, yeah, this is, this is the model. Great. Um, so we've got a couple more people to, to add

7:51

SPEAKER_01

into this vibe check. So we've got, as I said before, we've got Kieran Klassen, the GM of Cora and the creator of compound engineering, and we've got Naveen, the GM of Monologue. Hello, Naveen. How's it going? Hi, hi. Good. I'm really excited. So, okay. So I, we, I, we've got, if you look at our vibe, if you look at our vibe check, um, what we do on these vibe checks is every time we publish one, uh, this comes from, um, three weeks of internal testing. We always do the reach test, which is, do you reach for it every day for? And if you do, what do you reach for it for? And what do you reach for it over? So

8:29

SPEAKER_01

GBT 5.5 is absolutely my, my daily driver model. I'm a green. Kieran though is a yellow. So Kieran, you have some, I think you have some, some nuanced thoughts on the strengths and weaknesses of this model. Can you talk to us about what you found? Yeah, absolutely. So for, for me to use a model and use a daily, uh, it has to help me build Cora. I'm building a new version of Cora. Um, and it's a lot of product work. It's a lot of coding. It's a lot of testing. Uh, it's very wide work, lots of work. So front and back end, everything. And what I noticed is that while 5.5 is a very good model, like in my benchmarks that I ran, um, like when, when I ran it in 4.7, I was like,

9:18

SPEAKER_01

oh, this is the best coding model out there. And 5.5, GBT 5.5 scored the same on that specific benchmark. So it's a very, very good coding model. Um, but why, why don't I use it every day is that it's, it's also, it's also, it feels more like a specialist and less like a generalist. That's kind of how I look at it. Like Claude is the generalist that is a very good coding model, but it's also very good at product work. It's good at going into details. It's, uh, it's going, it's good at like looking at the big picture and for, um, 25.5, it's very good in execution and going into details, but sometimes it like breaks down if you look at it from far away and you just see

10:10

SPEAKER_01

things not being coherent. And I need that. I need a generalist. I need some, something that is working for everything. Although I'm using GPT 5.5 for review for execution. Like I use it in my flow, but it's not the daily driver that I go for first. Like the daily driver is still 4.7 because it's better of a generalist. And since I do product work and more general work, uh, it's, it's, uh, yeah, it's, I think it's still the same philosophy as we ever always seen. Open AI just like have a different perspective on what engineering work is than anthropic. And you see that in the model. It's not

10:54

SPEAKER_01

good or bad. It's just, some people are more aligned with one than the other. And I think

11:00

SPEAKER_07

Naveen, for example, is more aligned with the open AI take on what engineering work is than my take. Like I'm more of like a product engineer generalist. Naveen is maybe more of like an engineer engineer. And like, it's just for different people. It's not better or worse, but for me personally, um, while it's a very strong model and I've one shot at like amazing things with, uh, 5.5, all the models are very good, which is a very good luxury we have. Like there's nothing bad about any of these models. They're amazing. Yet like these little things make you go for one or the other. Um, so on that note, uh, Naveen, GM of monologue, uh, you're, you're usually a big GBT fan.

11:51

SPEAKER_07

Um, so I'm curious what you think of this release and, and sort of how you compare and contrast Kieran's feeling about this model. Cause I, I'm having this experience where I feel like I totally see some of the things that Kieran says, and I have a slightly different feeling. And I think that you have some different feelings. So there's like, I think there's a lot of different viewpoints and different types of people doing different types of work that make this model feel

12:13

SPEAKER_01

like, Holy fucking shit. I've never, I I've never seen a model do this versus, um, maybe it's not as, as good of a generalist for product work, like, like Kieran's finding. So Naveen, tell us what you found in your testing. So I, I used to agree with Kieran, with the older models of GPT, when he says this, yeah, I am with you because I also used to reach out to cloud code, uh, like, you know, opus models, uh, but coming to this 5.5, it just feels really well rounded model here in my case, I am like, you know, work, uh, in Python code base or shift code base, like native Mac app, uh, when it coming to web app codec, like codecs

12:55

SPEAKER_01

is not really good. So when you're building front end or anything, I always reach out to cloud code or some other model, but with 5.5, I didn't reach out to any other model. Like I know, Opus 4.7 released, I tried it a couple of times, but I just went back to 5.5 because it's really

13:13

SPEAKER_07

great for like the everything before, uh, I, I also do support, right. From on log, I do ton of, uh, replies, uh, obviously we use AI to do that before I used to use cloud code because it's like really good at writing, but with 5.5, I just completely moved to move that as well. So, and main drawback with the older models is vibe coding where I know for sure I can't pick codecs as my vibe coding tool. Uh, if I get come across like any other, like, you know, any idea, I just always go to cloud code, but in this case as well, I vibe coded three different apps in last couple of years.

13:53

SPEAKER_07

Can you show us some of the apps? Cause, cause one of the thing, one of the cool things is that, uh, Naveen used like 900 million tokens on GPT 5.5 over the last couple of weeks. Uh, so he's really, really, uh, making those GPS go burr. So Naveen, can you show us, I know you did an app called Bayline. So I'd love for you to show us an app that you are able to vibe code with GPT 5.5 on the side, while you're also shipping a bunch of features for monologue while you were also, you had pink eye. So you were pretty, you're kind of out of commission for a little bit there. I think I really helped me

14:26

SPEAKER_07

wipe code this because I'm just, you know, resting, but I have monologue dictating into, you know, uh, codecs. Let me share my screen. I'm sharing my screen. Hey, uh, so first one second, one second. Okay. Um, I removed myself and then I'm gonna add you. Good. Okay. Now we can see your screen. Can you zoom in a little bit? Okay. Nevermind. Oh, okay. Uh, let me open my downloads. I just like tried to download recent things.

15:02

SPEAKER_07

So one, okay. Vipe coding day line. Yeah. I want to show you guys, uh, this thing where it's a Raycast alternative. I personally use a Raycast notes like this notes where you can just bring this whenever you want to on top, but the thing is, this is like a plain text, right? I don't want this kind of plain text. So I thought, okay, let me make the to-do list, which is day to day. Like I can just have this like to-dos on a daily basis. So you can see how I'm already using it. Um, and yeah, this whole thing, I didn't look at single code. It just completely, I just gave a screenshot, uh, of Raycast.

15:43

SPEAKER_07

Like, okay. I really like Raycast. Like just go take, uh, do, do it. And then it wipe coded. What's really cool thing about it is though, the Mac app that you see here, like you can see, I like, I hit enter, you can type it, I hit enter. You can go back all these like interactions, minor interactions that I'm using. It's like difficult to implement, you know, in a Mac app. And it did that all in one single thread. This is one single, I think 200 million token thread that you're seeing right now.

16:19

SPEAKER_01

That's 200 million tokens. Oh my God. Yeah. It's all in one single thread. Uh, here I used, uh, build iOS apps, uh, plugin. That's the plugin that, uh, like, you know, codex comes with. This is my long prompt. I want to implement it today. This is what I was doing when I got pinged. So yeah. And then it implemented the first thing you can see 49 messages. Like, wow. It just went at it. It implemented the initial version. I was like started, Oh, can you actually take a screenshots on the Mac app and then, uh, do it for iOS as well. Now there is iOS app as well that syncs automatically. I'm just like blown away with, like, you can see,

17:02

SPEAKER_00

I just keep on having this to and fro and I'm just now regularly using it. I don't think I will be releasing it for anyone else, but it's just fun to test it, test it out. And today I experienced one bug, uh, and I just said that it fixed it and I'm just running it again. That's it.

17:21

SPEAKER_01

So that's really, that's really interesting. Yeah. I think there, there's something really interesting here. I want to pull out. Um, because, uh, because I've done a lot, I've done a lot of testing with this and I think some of your results are similar to mine and the, the, some of the things that I've found are, so I have this senior engineer benchmark. And like I said, if you test GPT 5.5 on the senior engineer benchmark versus Opus 4.7, um, 5.5 scores about a 62 and four, seven scores about a 33. So there's a 30 point difference, but the really interesting thing is that 5.5 scores

17:59

SPEAKER_01

the best. It gets that 62 when it uses an Opus level plan, when it uses a plan literally written by Opus 4.7. So that's really interesting. Right. And I sort of tried to tease apart what, what was going on. So, um, the thing that 5.5 can do that four, seven cannot do is if you give it a prompt that says, Hey, I want you to like go and rewrite a major part of this code base. I don't, I don't know what it is, but this is just a vibe. This is a vibe coded slop code base, like figure out how to, um, how you would rewrite it from first principles and then do it both four, seven and five, five can figure out,

18:36

SPEAKER_01

okay, here are the core principles. Here are the core invariants that I would rewrite and write a plan for it. But what happens is when you then say, okay, go execute it for Opus 4.7 says to itself, says to you like, Hey, actually like this is way too big of a project. I'm just going to pick off like a little piece. And even if you kind of push it a little bit to

18:56

SPEAKER_04

like, no, do the whole thing, it will still only do a little bit of a patch over the problem. It doesn't want to go and like actually rewrite the whole thing and actually delete a bunch of code. So it sort of gets distracted and a little bit almost intimidated by a really big code base and a big rewrite. Whereas GPT 5.5 is really interesting for its ability to take that plan and actually execute it over many turns, over many, many, many hours, over many, many tokens. And for having this sort of like ability to, um, uh, to like have the almost like the courage or the assertiveness to be like, okay, I'm going to go and like delete a bunch of code and I'm really going

19:40

SPEAKER_04

to think about this from first principles without getting as distracted by the existing code. And that's just something that that's new. Like I just have not seen that in, in an Opus model. I've not seen that in GPT 5.4. It's even less present in GPT 5.5 on high reasoning. This is only really on extra high. Um, and, and some of the things when you think about when you, when I looked at, okay, it's, it happens on extra high and it happens on extra high with an Opus 4 7 plan. What's different about the Opus 4 7 plan? The Opus 4 7 plan is, is very, it's very terse and very spec like and very contract

20:17

SPEAKER_04

heavy. It says things like a good rewrite will take this gigantic file and get it down to 500 lines. Like it has almost like contracts that it's giving to 5.5. And even though Opus won't itself execute those contracts, 5.5 is totally happy to go do that. And I think that's a really, that's a show, something really interesting about the character of this model. And I think it's part of what you're seeing to be in, in your vibe coding results because you're giving it, you're not doing a detailed engineering spec or whatever, but you are, you know, if you go look at the original prompt,

20:52

SPEAKER_04

you are giving it a lot of detail that a, a like baby vibe coder maybe wouldn't. And it, what it's able to do is take that detail and keep it in its head and actually like turn it into like a full thing from start to finish. Does that make sense? I see that.

21:11

SPEAKER_04

Do you want to talk a little bit about that? I would love that yeah. Yeah so I have a, one of my benchmarks is, is like purposely very under specified. Like I don't use plan mode and it's just create a version, like a clone of type form, but call it talk form. And the goal is like the backend is exactly the same as type form, structured data but the front end is you're just talking to the person you're interviewing them to get the answers for that structured data because it's like i don't think this exists uh but i think it would be cool if it did exist and i've vibra-coded that a few times in workshops and um and

21:45

SPEAKER_04

and the prompt is like four lines uh on purpose uh because i what i really like to see is like from scratch with no repo with a very underspecified prompt like as a beginner like is this model accessible and uh so i we we talked about this because i was like oh it didn't do very good on vibe coding and then you guys are what are you talking about this is the best vibe coding model ever and i think that was the difference um i think opus is like very ambitious to start uh and then it has like i don't know if you've noticed it says like okay are you guys ready to wrap this up uh you know like if i'm having a long conversation with it it almost like feels like

22:19

SPEAKER_04

panics that it's using too many tokens yeah i hate to anthropomorphize but uh that's how it feels it's like okay do you want to wrap this up and i'm like no no not yet and it's like okay but now do you want to wrap it up and and so i wonder if that's the difference between uh what we're doing was like codex seems pretty chill you know it's like dogged and determined it's just gonna like keep pounding it out you know time after time after time so i think like it really depends on what shape of job that you're doing uh when you're using these models 100 percent naveen any any reactions there

22:50

SPEAKER_04

um yeah i i i don't completely like you know i'm not seeing that so maybe it says that i do a lot of big brain dump prompts and then recently uh like i'm with dan what you observed which is uh i recently had to add mcp support remote mcp support to monologue i initially implemented really bad mcp server i think uh i just did it like uh so what happened is last week we are like you know doing the launch i just asked gpt 5.5 to come up with a plan this is where i use the plan and it added a plan across front end back end client side both ios mac os right and then i just said i'm like too lazy so i said implement this

23:41

SPEAKER_04

plan and then it did that like it continuously kept on working across all multiple code bases and that's when i know okay uh this model is really good at managing that multiple code like you know across context i think that's why i'm like really curious to test this uh single thread uh white coding app as well because it kind of like uh i'm not sure what they're doing with compaction i think the context between compaction or the context between each thread that uh models are like you know using it's really good so it's not losing that context i think the product shape that uh you know i think

24:23

SPEAKER_04

that's what makes this model really good yeah i actually have like a side project where i have a one

24:29

SPEAKER_01

of the carpathy style and knowledge bases uh that's but i'm doing it away from keyboard like it's just running in the background and uh i i initially set it up as a ralph wiggum loop right like uh every every time uh claude stops it uh you know it checks uh to see uh you know like what's the next task we have to do and then uh it takes that task down and then it commits and then it stops and then it starts up again and uh and i switched to codex uh on on that uh a while back and um and and uh i've just just been running uh and and uh i realized that i'm like it doesn't need that uh that like ralph

25:06

SPEAKER_01

wiggum loop anymore so i just switched it basically to just let it compact itself and and uh it's been going a lot faster like i need i needed less harness essentially which was pretty nice um do you i'm curious mike do you have examples to show us of of more knowledge-type tasks that you worked on it with yeah yeah i can pull some stuff up let me hear my okay that'll be great and then while he's doing that kieran what else what else do you have to say about this model what else do you think um you you've experienced that we haven't talked about yet yeah so one observation was uh that design looks

25:44

SPEAKER_01

a little bit less good than the previous model so um it just looked a little bit chaotic but uh design

25:54

SPEAKER_00

in sense of structure and typography look better so if you if you if you like typography it's better and so so this is the app where i try to just one shot from one prompt an entire app with like a rubber ducky store that where you can customize ducks and all of that like it nailed it everything worked which is like this is the first time with 4.7 this ever worked so 4.7 and 5.5 are beasts this is a web app this is react like a next app um it's very very very good at this so if you do web you uh need stuff to work in full detail um this is great uh i've seen here so i run the lfg bench which

26:42

SPEAKER_00

is using compound engineering to uh just run all the steps by itself and it took three times more tokens in the planning and the review state similar to 4.7 so it takes the tokens this was extra high but as you can see here like yeah it looks cool the typography side and the alignment and the padding but the the right side like is weird like what is that even like i don't even understand what it is so there is some some degradation like the previous models were better at this in design uh and yeah like i'm sure they're going back and forth like they're like okay now we need to structure more make

27:26

SPEAKER_00

it less chaotic and more structured um this one is the island test and i said just make a cozy island and uh it's pretty good like it did the the ground was blue instead of brown um and animations were not uh working completely correctly but it's very close like 4.7 opus did better there like the uh but yeah it's it's still like it's getting better every time and it's cool that you can see where they push the model towards um and yeah so design wise this one is actually good only the colors of here like the the nice thing like what i've seen before with gpt models is like it adds stuff it

28:14

SPEAKER_01

doesn't really need to add and it does a little bit too much in the ui that that's not there anymore so like that extra like it's it's more like opus where it just shows the stuff it needs to show which is which is nice it doesn't do extra things um and yeah so design but like one-shotting um like i did a i did a library and i just uh here i can share that i had an idea one morning where i was like oh it's like actually one of the more important things now with ai are like doing a call with someone and going through a product and sharing what you like about it and not like about it and um

29:03

SPEAKER_01

so i created this uh library call it's i shared my screen um but basically in the morning i was having coffee and i just pulled up this and and i was like hey what what if we have like a react package if you can share my screen then you can see the package uh got it yeah so i created this uh like riff rack as like let's go riff together and record it and what i did was i just record screen sharing

29:33

SPEAKER_00

and then feed it to a model to see what was told but i was like well what if we record like the dom interactions as well in addition to um all the other stuff so a model can very precisely see when you say oh this is dumb or this should move somewhere else you click on something it has that element and then that can be extracted very well into uh like an improvement well as you can see it's three commits yesterday and it's basically initial commit one uh one review commit and one more review commit i

30:10

SPEAKER_01

didn't do anything this is just lfg in one go like one prompt in monologue uh copy into uh like compound engineering lfg flow into codex five point or in gpt 5.5 and it did it and it's super detailed like it had oh yeah there was one iteration where i said hey like maybe we should have a consent screen it actually proposed that it was like hey like maybe we should do a consent screen that's probably something you want that's missing here came up in the reviews of it yeah sounds good um but it has all these things and um basically it has these things hey you can enable recording for urls by doing this and you can even

30:53

SPEAKER_01

um create your own param where you can do your own so like there are these details that are super nice that it adds that i haven't thought of but it's not too much it's not sloppy it feels neat and organized and so it's very good at these things this is typescript i use ruby in cora it's just not good at ruby

31:19

SPEAKER_00

so it's just very hard for me to use because it's not good at ruby it just does typescript stuff but then in ruby and that's not how you write ruby so for me that is honestly the biggest blocker like if you do typescript or anything open ai trained on like it's probably amazing but i built cora and use ruby and ruby is not liking the model as much so you can see like this is like an amazing experience and and i i i rolled this out in cora i will try it out basically it will just record your screen record all your events and you get a zip file stored somewhere you can share it with me and then

32:00

SPEAKER_00

feed it and just say lfg ideate uh come up with the the things we need to build here and uh yeah that's just a morning one shot and that's the world we live in and i'm not even surprised like some people say oh wow wow amazing guy for me that's the norm one-shotting things for me the next thing is like how can we one shot 10 things and release them automatically so i'm more like for me the bottleneck is not one-shotting things because that's kind of working for me the bottleneck is like how do you collaborate as a human with the model because that's the unblock and i think um like the ideation the

32:41

SPEAKER_00

polish uh at the beginning and the end like that's where you need most um help as a human still i love it um thank you for sharing that karen i'm gonna now bring in our head of growth austin austin welcome what's up um and austin has been testing this model for three weeks alongside us and the codex app and he is in an interesting place because he's using it for head of growth type stuff so for more knowledge work type pass and he's coming from cloud code so awesome give us the headline on um i think you started this experiment using cloud code and co-work mostly and you ended uh using

33:23

SPEAKER_00

mostly codex tell us about your experience yeah over the last few weeks the codex app in particular and think it's really because of like the the power and speed of the codex app relative to the cloud desktop app has become my daily driver for everything i do for someone like me without like a technical or engineering background i feel way more comfortable doing any kind of engineering work with the latest open ai models that can be everything from like like running code to build like underlying dashboards across multiple tools shipping landing pages that look at like our figma design or whatever doing

34:05

SPEAKER_00

analysis creating plans and when i you first like were nudging me to at least try codex a few months ago dan and i i did try it and i was like either i'm too stupid to use this it thinks i'm too stupid to use this or it's talking to me like like i can't hang at its level and um i was like i i kind of wish i could get as much out of this as as you did or as i knew like naveen was a big fan of it and i just kept going back to cloud code in the cli because i was like okay it's executing it at like a level i like and um it's it's like nice to me and and engaging and as someone who is still relatively new to this stuff

34:46

SPEAKER_00

that was important to me yeah you're like codex makes me feel like i'm dumb makes me feel like i'm dumb it makes me feel bad about it and um this came up a lot in the previous open ai models in codex because i would make a plan and in planning mode it would say like i have four questions for you and

35:04

SPEAKER_01

for each question for the options all the options i was like i have no idea what the hell you're saying

35:08

SPEAKER_00

please explain it to me and the tone of the open ai model was kind of like why why would i have to explain this to you why don't you just do what's recommended and especially in knowledge work when i'm trying to run like a campaign or planning like it's a weird feeling to have and all of that is gone for me in in the new model that we've been playing with for the past three weeks i both understand what it's saying i also really trust what it's saying the um the best example i have here is we are about to launch our plus one product um in a few weeks and uh we're sprinting to do it i'm sprinting to get a

35:45

SPEAKER_00

big like campaign together and i pointed um in the codex app i pointed the model at our notion our slack where um like all of our meeting notes and conversations have and i basically just like i did a monologue brain dump of like the high level version of what i wanted the goals of the campaign to be and a previous campaign plan we had and i said go make a campaign plan looking at this and it's informed by like my taste and the team's taste and thoughts because it looks at our meeting notes where we talked about what we want to do and the plan that we're going to go to market with is like

36:22

SPEAKER_00

90 of what it came up with i was like so impressed and it was so easy because it grabbed all of this context it asked me questions that made sense to clarify and the most important thing to me is that inside of the codex app it's so easy to use and it's just so fast i find the claw desktop app impossible to use due to the like clunkiness and slowness after having experienced how fast for knowledge work codex is with with these models um and then the the other thing that i really really like that i think is driven a bit more by how the codex app works with what i feel rather than the models itself is the

37:01

SPEAKER_00

way that it recommends automations to me for knowledge work and then just builds them and starts starts using them if i if i ask for them is um so powerful and impressive and they're automations that i rely on every day so um to me like the only way i would switch back to claude and opus whatever new model is my daily driver is if the if the cloud desktop app got like dramatically faster and better so as kind of like a an easy way to work it it was competitive with the codex app in in that way or if like something about the the model for knowledge work like far surpassed where the open ai models are and

37:47

SPEAKER_01

to me they're they feel about even for most stuff besides design i agree with karen that like um we use ai for a lot of like video clipping image generation stuff and over and over again i find myself still reaching for claude for that kind of stuff yeah that makes sense i i do i think it's really important just to like talk about how remarkable that shift has been like even three months ago definitely six months ago but even three months ago it was like it was one of those things where you just couldn't use codex for any kind of knowledge work um it was too slow and it treated you like and now it's super fast and it's

38:28

SPEAKER_01

just good for the kinds of things that if you're ahead of growth or you're like me you're doing like analysis of like some of our metrics or you're writing investor updates or whatever it's just like it's the thing that i want to reach for i think um obviously we've been talking about opus 4.7 has some unique powers so for example i think it's planning is much better we think it's design is much better um but uh it's also even in the places where it's roughly equivalent which i think there are a decent number of places like that it is actually a lot slower um and 5.5 is just a fast model and that

39:00

SPEAKER_01

that makes a huge difference i think speed is in some ways it's it is a type of intelligence um and it gives gives you a lot more power to have it be so fast yeah it passes this like i don't have a benchmark version of me yet i kind of want to make one of like my like uh do my knowledge work while i'm in meetings test where i'm like that or like i did this on sunday where i was watching the nba playoffs and i'm mostly watching the game and i'm just like i'm like nudging it along and my brain is only maybe like 10 working and when i'm in between meetings i'll check the outcome of whatever it is and this

39:31

SPEAKER_01

is a lot of like planning execution stuff and it moves so fast that i like can't keep watching the basketball game or i can't like i like during the meeting i'm like oh i like want to keep go looking at it and then it has something good that i can actually go do the work and i don't feel like i have that experience with opus that's really interesting okay so we've got i have a question is it okay then yeah like how is this up with remotion i know you've been playing with remotion a lot so how is gpt 5.5 with pre-motion i have so a thing that i do a lot as someone who like needs to make

40:12

SPEAKER_01

social assets um for us but has no real design ability or sense is i like to point these models at designs our designers have already made primarily through the figma mcp or to screen record how our gm's products already work like we launched a new version of sparkle with these models you can do a screen recording give it to give it to one of the models and say like hey make me a video with this storyboard and um over and over again with the exact same prompts the open ai models kind of just like take the recording and like zoom in a little bit it's like very very odd it's like the laziest

40:56

SPEAKER_01

possible approach and with the exact same prompts i find that uh opus and like all of the recent opus models have been this way for a few months they take this level of creativity and they they engage with you back and forth and questions for how to make it more creative that like i haven't found it close at all um i'm intrigued to see how much open ai catches up here because dan and i were messing around with the new image generation tool for youtube thumbnails and it was really good and it's really fast it's so good it's crazy um i was kind of hesitant to try it because i i'm like in my head

41:31

SPEAKER_01

i'm like this is the one thing the open a models can't do and um right before i jumped on i knocked one out for the the the thumbnail for this live stream was generated on the new open ai image generation tool in like three minutes and i was like i don't have to give it any notes see yeah like gpt image too is really cool can you show us austin yeah i think so um let me because i i haven't seen it yet so i want i want to see it but yeah let me see let me see if i can find the the chats uh so you can see like the conversation and just like all it took um cool all right i've got it here let me show my

42:17

SPEAKER_01

screen um i would say all right cool can you all see this uh yep we can all right cool so what i gave it was um two thumbnails actually that that dan had made and with gpt image with gpt image too so like it has it has this baseline this is for the video we recorded before the uh can you just click on one of them i just want to show people how good oh yeah for sure yeah so this is this is one that dan made i think very quickly right i made this really quickly this is like essentially one shot with uh the new gpt image model giving it a thought giving it a picture of me and this is a little bit more

42:56

SPEAKER_01

massaging with gpt 5.5 um and the new image model and it's like that's that's kind that's just crazy that it does that um and it's very it's very good at like um oh i want you to change this one little thing but keep the rest the same and it's also good at it looks like me which a lot of these models have not been that's been harder for them yeah it looks like you this is like i would say like a 95 approximation of our logo which sometimes these things are very bad at um so i gave it that and then i gave it a few more um just like screenshots of the video and basically said um hey this is the text for the youtube video

43:35

SPEAKER_01

about to do give me three rips on this and uh it made this one uh which is i believe what we're using it also made this one which is uh like a little busy but really good like in the busy state it's like quite quite good um typically these kinds of icons these things make would look absolutely horrible and they look pretty clean and it has the right number of fingers and that's not from a picture my fingers are not in there that's crazy um and this was the this was the third one just off of that prompt i gave it and the interesting thing for me with this is that my my process in um in cloud

44:14

SPEAKER_01

code with opus has previously been that i made a skill um for like a youtube kit skill and i had to preload it with all this context where it has the youtube api it has the figmcp connection it has all this stuff so it can go look at the internal taste that we've created and what's worked for us previously and the analytics behind our youtube thumbnails and all this was was like hey we're making this video and then it creates this thumbnail and um you know when we do these live streams and we do these vibe checks and reactions we have to move so fast we're doing so much stuff that um the only way

44:51

SPEAKER_01

we're able to react quickly with real insights and have stuff that performs well is tools like this that get us from zero to one in seconds i love it i love it okay thank you austin uh that's amazing so we've got some really cool stuff coming up we've got we're gonna have mike who's our head of consulting in our tech practice uh talk to us more about uh how he uses gpd 5.5 for knowledge work and his vibe

45:21

SPEAKER_07

check on that um the good the bad the ugly all that kind of stuff we're gonna have some special guests joining so there's more special guests non-every guests uh i will not say exactly who it is but you'll want to stick around for that before we get there what you should know is you got to stay hydrated it's important to here mike you got to stay hydrated it's important to keep yourself and your agent liquid cooled you should also know that every is the only subscription you need to stay at the edge of ai and i want to show you a little bit about every in case you don't know who we are and what we're about we are every um we do these

46:09

SPEAKER_07

vibe checks every time a new model drops we've got one this is our vibe check for today i want to show it to you look at this it's beautiful we've got a full team across coding writing design we we build tons of apps we're we're you know we're building when using open claw and what happens is we get access to these models so we got it we got access to 5.5 about three weeks ago and we get to put it through its paces and so you get the benefit of if you're in every subscriber of getting our take on day one from uh from our hands-on testing a couple if you're just getting here um a couple things that

46:48

SPEAKER_07

are important to know about this model we found it's it's really spikes on on some of its senior engineering ability especially in this in the sort of back-end uh department and especially if you give it a good plan so on our senior engineer benchmark which measures how good a model is at taking a vibe coded slop code base written by yours truly and rewriting that code base in the way a senior engineer would comparing it to real senior engineer code and on that benchmark gpt 5.5 scored a 62 out of 100 by comparison gpt uh opus 4.7 scored a 33 so there's about a 30 point swing there um and so this

47:31

SPEAKER_07

model has some special characteristics that we have not found in any other model that we've tested um a couple other things that are that are really interesting about this model it's it has that spiky senior engineer ability and it's also super fast and and and pretty good to

47:50

SPEAKER_01

talk to which is a it's rare to see a model improve on both dimensions at once i think you can you can see

47:56

SPEAKER_00

that a little bit with um with opus 4.7 it got a lot better at programming it's it and it but it's also much more terse and a little bit more like it feels like it lost some of its empathy um and i think open ai did a really fantastic job of making this much smarter but also keeping it uh pretty pretty personable and pretty fast so it's usable for the kind of like day-to-day knowledge knowledge work tasks that that were that we're used to when we at every when we do these um these vibe checks we always do a reach test which is are you using this every day uh are you reaching for it over other things and if so what

48:28

SPEAKER_00

are you reaching for it uh i'm a i'm a green on this model this is my daily driver for pretty much everything the only thing that i still use opus for is every once in a while when i'm not talking to my claw i'll i'll talk to uh claude on mobile i think it's i i really like claude on mobile mobile but other than that it's it's pretty much my go-to for kieran kieran is a yellow i think there's some things kieran you like some things you don't like but it's it seems like uh opus 4.7 is still really your your go-to model your your daily driver with using this this model for certain other things

49:01

SPEAKER_07

naveen uh you are a green you're usually an open ai stan so uh so you know know that but i think that you you just you love this you love this model um mike is a yellow and uh katie parrot who's our uh lead staff writer is a green so it's not all greens but there's there's a lot of people psyched about this model and even if you're a yellow i think i think pretty much everyone agrees there's a lot of power here um a couple things that we've so that we saw so again like it's really good at doing that senior engineer benchmark um it's really good at uh if you give it an idea for a big thing it's going to

49:37

SPEAKER_07

build like pushing that idea through over many rounds and many many tokens um it's a little bit weak it's it's really good at some visual design stuff but it's a little bit weaker at other things so for example kieran did this uh benchmark test where he asked it to to design this karen do you

49:54

SPEAKER_01

want to give us just a little bit of a uh a refresher on exactly what we're seeing here and what you think it says about the character of gp 5.5 as a model yeah and it's also good to understand what it is the benchmark because there i also test the harness here so like it is testing the the the

50:15

SPEAKER_04

codex cli harness and it's testing the plugin framework and it's testing compound engineering so there's a lot of things going on here it's not just a prompt in uh result out so it's also like how does gpt 5.5 work within the codex harness like what kind of features does the codex harness have and that is also represented in here but what i've seen change from previous things for design is like it's better at typography it's less fiddly like weird details that were in in the previous version but it's not as good at like visual so you can see that the the left side looks good with the typography

51:00

SPEAKER_04

but the right side looks it doesn't doesn't make any sense like there's like there does not really make a lot of sense what it even is or what it's trying to do there and that was better in other models but also i do see it improve in structure and maybe you can argue actually structure is more important and finishing things and not having any bugs and like staying on course is more important because these visuals are probably done by a designer or by an image model which they have which is very good so like you could even argue that i'm not saying this is bad or good i'm just

51:38

SPEAKER_04

observing this changed and normally in these changes if you see them in different places you can see what they try to do and this is a great model but also i test the model within the harness so i when i say yellow it means yellow for also the harness in which i use it and the harness is is very good um but there are other harness that also very good so uh also what what i'm what i'm uh what i like to do is i also recently like to try models against each other in the same harness in like cursor for example and and just see them head to head i will do more of that i remember last time i did a vibe

52:25

SPEAKER_04

check for a cursor i say yellow and now i'm using cursor every day so like me saying yellow on 5.5 means maybe that i need to adjust how i work or learn the powers or it just means i need to change something to see it work and i'm not saying i won't try it i'm actually running right now i'm running like three codex things in the background uh doing stuff so like i'm using this model for sure and but it's a big thing to say daily driver and go super green so like i'm very very enthusiastic about this model it's super good and i'm very excited that i can do things in one shot because that was never really

53:10

SPEAKER_04

the case with like older versions where it just stopped working or got trapped at some point i think 5.3 was the first time where i was like okay open ai you got my attention this is actually like very very very usable and 5.5 is just that but then way better like it's like if they keep going this is great this is the best thing any engineer or creator can have we're just like they're so good now that like you yeah it's hard to compete and they're so good and you can all use all the models so really the vibe check here is for me like what is different and what are the differences between

53:52

SPEAKER_04

all the models it doesn't mean it's good or bad i love it um okay so we've got about three minutes before our next guest uh mike taylor our head of technology consulting at every do you want to tell us you've been using this for some of uh some knowledge work use cases do you want to give us some uh some thoughts on on what you found in doing that i'm just going to open open up uh some of the things that you probably going to want to be looking at yeah do you want to open the dashboards yeah oh yeah let's do let's we'll do the dashboards okay yeah so this is a opus dashboard and you'll probably

54:30

SPEAKER_04

recognize that well maybe like uh introduce the benchmark first yeah yeah so the benchmark here is uh this is actually something that uh comes up again and again in every workshop i do uh can we make a custom dashboard from and this is something that used to take you know a lot of time a lot of it tickets to accomplish and now people are just vibing it and this looks vibed right this is uh opus um and you can see like you probably recognize the color scheme if you've uh done a bunch of vibe coded dashboards uh like if you just scroll down kind of you've got like uh uh yeah it's like very very bright like

55:05

SPEAKER_04

dark dark background and you've got the purple at the bottom as well the vibe coded purple oh yeah the purple always yeah uh now if you switch over to the other dashboard and so this this was uh five five this is gpt five five right so um i mean it's like way better it looks more professional it looks more professional like if you scroll down it starts to get really nice uh you know the color scheme was like more tasteful uh this the spacing was better um i actually like wanted to read it so i would say like this is the one this is the first benchmark i ran i was like oh my god this is like agi and and

55:40

SPEAKER_04

it's like it's a little bit more nuanced after after that but just to point that one out because uh that's a like that comes up in every workshop and i would say like this is like a very very clear win and and it surprised me because i uh i just saw that like with opus it did the best powerpoint deck i've ever seen and the reason was that it was better at design and it had better vision than previous models so i wasn't expecting uh this um weirdly uh you know gpt 5.5 failed the powerpoint test versus opus but it beat opus on the dashboard so it's not a clear picture that it's like worse at

56:15

SPEAKER_04

design better at design i think uh depends on what you're designing i love it thank you for thank you for sharing that we have now i'm very excited to welcome to the stage a very special guest romaine are you ready are you ready to be uh on the stage no he's already at mine because you are

56:36

SPEAKER_01

i'll take you off if you're not romaine is not ready okay uh so uh just uh just a little heads up romaine hewitt is the head of developer experience at open ai and he's joining us to talk to us about this model that's coming up soon um and while we wait i'm going to share my screen again and i'm going to just say every is the only subscription every is the only subscription that you need to stay at the edge of ai we get these models uh way before they're released and on release day you can expect a

57:09

SPEAKER_03

vibe check like this from us where we use it for everything from coding to writing to design to marketing to strategy and uh this page is not loading but this page does load we've got a really really sweet long vibe check here it's uh several thousand words a lot of it's free a lot of it some of it's for paid subscribers but if you're in every subscriber you get a lot of other cool stuff so as part of every we have a bundle of apps that we build um some of the apps are built by some of the the amazing people on this call on this live stream we've got quora which is an ai agent for your email

57:46

SPEAKER_03

and soon an inbox kieran i would love for you to show us some uh some inbox type stuff if you feel ready to do that um we also have monologue built by naveen it's a smart dictation app and it we recently launched monologue notes so um monologue uh started as you it's like a speech to text uh uh app sort of like whisper flow or super whisper but we just added notes so you can transcribe a note on on your walk you can you can record a note in a meeting you can uh brain dump into it it's really great i love it and it's very available to all your ai agents so it's really easy to get your notes where you want them

58:21

SPEAKER_01

so we've got quora we've got monologue we've got sparkle which organizes your desktop we've got spiral

58:27

SPEAKER_03

which is a uh agentic um ghostwriter we've got proof which is a markdown editor for your agent and we've got plus ones which is our um hosted open claw so this is all available for every subscriber as you pay one price you get access to everything that we write everything that we make and all of the we do a lot of um we do a lot of live streams and camps and stuff to teach you how we how we use stuff at the edge and we also have a pretty cool consulting practice uh where we work with big companies work with executive leadership teams to get their hands on um get their hands on stuff like

59:03

SPEAKER_03

cloud code or codex and get them using agents and get agents into the org um so if you have not checked out every you should go to every.to and you should subscribe and uh i think we're still waiting i think we're still waiting on romaine while we're doing that mike anything you want to add on some of the benchmarks you did or uh or maybe anything i haven't talked about on the consulting side yeah so um maybe if you want to bring up uh those two files called persona uh one for each model uh so i'll tell you a little bit about this benchmark uh so last year i worked on the startup uh failed uh unfortunately uh

59:41

SPEAKER_03

but uh which is why i'm working for a living uh but uh but i mean your loss is my game uh it's still going as a side project yeah we have we have some customers and stuff but um uh but but uh it's now largely away from keyboard i have uh claude and codex working on it while i'm here uh but uh yeah basically it's uh it was a virtual focus group uh startup so we used ai to roleplay as your customers uh so i got very very deep last year into cloning people like i you know we built a whole panel of 300 people that we uh cloned based on user interviews uh and then this is my clone uh just

1:00:17

SPEAKER_03

for ethical concerns don't share anyone else's clone uh but um this is uh this is a clone of me uh just a snapshot taken just before i moved here to new york um and uh yeah if you're seeing that on the screen here it's just a great benchmark because obviously like we have spiral right where it's like it learns to write like you but the way that i write isn't the way that i talk um and especially once it's gone through editorial it's a little bit different and i think it's maybe an easier benchmark to uh to to uh like write uh like uh an editorialized piece of content because that's more consistent it's more

1:00:53

SPEAKER_03

standard the way i talk is is really odd it has like inflections it has a lot more personal information so what this is is i i provide it with a transcript of this long half an hour interview i did um and that's the prompt and then i have some extra stuff on top to say basically impersonate me and um actually the first time i tried it uh it refused to impersonate me uh i think some alignment checks alignment is uh good and well but uh uh so so the first version failed on this bench but uh that that was quickly

1:01:23

SPEAKER_01

um fixed and and uh it worked beautifully um so i don't know if you have the this is the opus one is it yeah the opus one if you if you look especially if you scroll down to like maybe the second or third question um it's it's it's trying to too hard uh like there's a lot of like ums and ahs oh mate um a lot actually like i think exactly maybe i do talk like i mean i like there's some some of that but not not all of that yeah yeah and i i think it's just like feels so eager to be me that uh it's almost like a super version of me um like it's really emphasizing my ticks uh whereas if you go to gpt uh 5.5 um it's

1:02:07

SPEAKER_01

just much cleaner um it's just uh like a much smoother experience and it picked out a bunch of things that i actually uh really do think or or like very very plausibly like could think um which which is like you know as good as you can get really for this type of benchmark great um so i i agree i think we found the writing to just be a little bit smoother good good at voice but not overdoing it maybe a little bit less ai ish um and now i am very excited to welcome romaine hewitt and dominic dominic how do you pronounce your last name kunda but you can just say dom it's fine okay i'll say i'll say dom welcome

1:02:47

SPEAKER_03

both of you can you introduce yourselves yes thank you so much for having us uh this is roman i lead developer experience and uh i found uh dominic join me uh who leads a lot of the dx efforts for codex and so we're both very excited to talk to you today i'm excited we're psyched uh we've been we've been using this model for the last three weeks we love it uh it does some special things that we've never seen a model view before uh what like tell us about like what what you've seen internally and what you're excited about for this model i mean it's pretty amazing how much now like uh the work has

1:03:21

SPEAKER_03

changed at the point i like relying on this model i think we we've seen like across the board how

1:03:26

SPEAKER_01

much the model is smarter at any kind of professional task right like we talk a lot about writing software

1:03:31

SPEAKER_00

and of course we hear from engineers that the the model is writing better code it's like more thorough

1:03:37

SPEAKER_03

when it comes to like analyze analyzing a repository of complex code bases but also takes on like more and more of the other tasks that we do every single day and and i think what's very exciting to me is not just the model itself like gpt 5.5 but it's also the combo with codex because when you have such a great model you also have to have the right software and environment to to make it shine and and we're seeing like so much of that with codex right because codex can help the model not just like figure out the files and the context but also access all of the tools that it needs to really carry the work

1:04:12

SPEAKER_03

check-in sort of output and and really like deliver things that you can just ship right yeah i mean like i i i feel like 5.5 has especially in a combination with like the features we shipped last week in the codex app and some of the features we shipped today um has drastically changed how i work like compared to even just a few weeks ago um like i know that i can have codex handle everything from like the documentation to like i i think i've built like 50 uh demo apps in the last week um just sort of like trying to push codex to its limits and like like there's surely no way that it can do it and it just

1:04:49

SPEAKER_03

sort of like just hammers through it like last night i think it was like 11 pm or something i was like let let me see if i can like one shot like a full 3d modeling software um and like it it fully did it um like i i remember one of our colleagues was like well can you add like download uh like support for like the file type uh for like 3d files and i was like it already had it like it was already there

1:05:13

SPEAKER_01

so like the level of like attention to detail that like the model puts in is is incredible dom do you

1:05:20

SPEAKER_03

have any of those demos that you can show uh i don't have the 3d modeling one but i do have some let me um share my screen here sweet uh very impromptu demo by the way uh so we just like got out of the launch room uh flipping all the switches publishing the blog post and we just showed up so amazing i we're doing it live we're doing it live that's what we like that's how we do we want it we want it like this okay don you're live we've got your screen this was one of the one of the moments where i truly had

1:05:50

SPEAKER_01

sort of like a you know mind-blowing moment um and it like the final result is actually in the blog post as well if you want to take a look at it later but my colleague sent me this picture right like this is just a screenshot of some software um that he had been building and uh it's mostly like for like illustrative purposes so i thought like okay let's take this and i gave it to um five five and i simply told it like you know i had a couple of other folders in the project that like i didn't want it to deal with so i totally ignored those but then use webgl to and like the real data from the artemis 2 mission

1:06:26

SPEAKER_03

um to actually implement this um and it went ahead and um like fully built this as a working app and it went to the um you know it cited its sources it actually went to nasa downloaded like the flight data of the artemis 2 mission and then rendered all of this and some of the attention to detail that originally i was thrown off by was um it had um if i open this in our built-in browser here it had this like slightly weird curve where i was like why is this what like this doesn't look fully like how the data like how the curves looked when i saw like animations from other people um and the result was

1:07:12

SPEAKER_01

actually that like it saw in my image that i had given it that these two planets were like you know

1:07:20

SPEAKER_03

scaled like this so try to like replicate the scale but that's not real scale in space um so it had to like take the real data uh trajectory and then like fix it um for those uh for those moments i had i asked it later on to like add a true scale moment so like here you can see the actual sort of like slingshot and then how it goes over to the moon and if we go towards the end is also the flyby which is we have like the full flyby and like how it like sends the moon around it um and like all of the data like it tells me here like what were like the mission moments and stuff like this was like one of those

1:07:57

SPEAKER_03

moments where i'm like i i didn't give it a lot of instructions on what i wanted this to look like i just wanted it to look like the picture and it was like cool i got you um and so i've been experimenting a lot with this and so um one of the things that we sort of silently launched because

1:08:13

SPEAKER_01

i definitely want you to try it but like i didn't um like we just sort of like shipped this is um inside our like build if i find it build web apps like build web apps plugin if you have that already installed or if you install it now we added a new skill here um called front-end app builder and this one actually uses now the the design capabilities of image gen um that we launched this week uh together with like five fives attention to detail to build apps for you so i told it earlier like right before this like you know gpt5 is out here's the blog post but like a minimalist well-designed uh page and it

1:08:55

SPEAKER_01

went in and like you know this is one shot like the other things is me just asking to start the app um but this is the app that it built wow with like the data from the from the website um and the design if you look at it is like pretty spot on with um you know minor details of like image and generating the individual assets later on in slightly different positions but like it's pretty spot on for like a one shot that's really interesting so is this like is this a new workflow for doing front-end design is like have the image model like render a website and then have 5.5 turn that into html and css yeah i

1:09:35

SPEAKER_01

mean it's a great one i mean we've heard like you know from the community that one of the things that we still have to improve and push the frontier on is really like the front end of the taste kind of capabilities but now and we're still pushing on that very hard but at the same time like with imagine uh 2.0 the quality of the designs that you can create is pretty outstanding but also gpt 5.5 is so thorough at taking an image input and implementing it that you can now iterate on your design first and when you're happy with it just get it done with gpt 5.5 to implement it like this

1:10:08

SPEAKER_01

one which is pretty magical yeah and like i think i think one of the interesting things there is like it really brings together a lot of the capabilities that the model has gotten better at like one like the attention to detail but then also like it went through the entire like like loop here so like it used one of the things we added uh today as well is browser use inside this in-app browser um so it actually controlled the in-app browser to like take screenshots try things out um and then it would like do checks here on like you know does it work on mobile does it work on desktop take screenshots and

1:10:43

SPEAKER_01

like verify that work and really put in so that attention to detail before it felt comfortable with like all right this is good enough now um that's really the magic of bringing like the codex harness the codex surface and all of these tools right like i think if the model does not get everything right from the first try you don't even have to interrupt it because it checks it won't work like now with computer use outside of the codex app but even inside the codex app you get the model to just like keep on iterating until it's satisfied with the answer that's really interesting what i what i'm really curious about

1:11:16

SPEAKER_01

is i feel like you guys are on this crazy arc this crazy run right now where if i think about codex three months ago or six months certainly six months ago it was like very uh there was no desktop app and it was very like senior engineering uh both in terms of the the user base and what it was like to talk to it was like super slow it would get very very technical answers austin who's our head of growth was on the user base and what i was saying was saying like it just made me feel stupid um and i feel i feel like over the last especially the last three months kieran was saying since 5.3 which i agree with

1:11:53

SPEAKER_01

you've been on this run where like every release it just it's gotten smarter it's gotten um more personable and and really good for senior engineering work like it's it it's tops our senior engineer benchmark for what it's able to achieve but it's also uh just really good for just general agentic knowledge work so tell us about that shift and and the trajectory you're on yeah i mean like uh it's it's almost something that has been like very surprising to us in a great way like it's the like even at openia internally the adoption of the codex app right like it started with us being

1:12:28

SPEAKER_01

engineers and building with it and of course as you know we kind of started with the cloud and then the cli and then the id extension so we kind of had like different manifestations of the coding agents um but we always believed into this like agentic delegation vision and we truly needed a surface with the codex app to like achieve that that that vision of like being able to send complex tasks and have the model handle them but what's happened in the past like few weeks even at open ai and outside of open ai is like not just the best engineers in the world like moving to codex and taking on this

1:13:00

SPEAKER_01

like agentic delegation workflow it's also more and more of the like adjacent teams adjacent roles like

1:13:07

SPEAKER_04

actually using the codex app for any parts of their work um i mean like dom and i have like i've completely changed the way we work at open ai like most of us here like not just for building software but also to connect to like your slack and your notion with like you can show the plug-in screen but we have now like like 100 plugins where you can connect like all of your tools and all of the context to bring into this codex app and uh as of today also with the launches we also introduced like the artifacts so the ability to like hey if you're working with spreadsheets or with if you're working

1:13:38

SPEAKER_04

with like any kind of documents and files or if you produce like assets you can now visualize those in the codex app so it's really turning into a very advanced kind of like productivity app uh for much more than just the developer yeah go ahead i was just gonna say like i think one of the interesting things is that like um as i mentioned like over the last couple of weeks like the way i work has drastically changed like with these launches like especially we had a big launch last week we had a have a big launch this week and we had smaller ones along the way right like you know image gen was

1:14:13

SPEAKER_04

certainly a big launch but then we had things like chronicle um launch and research preview which um inside codex like can watch your screen so you can learn more about how you're using other tools and all of these sort of different capabilities have really helped um me like stay on top of all of these things where like the amount of times that i would kick off a task because of some feedback and i wouldn't even actually go in and um you know describe in like long words what i need to get fixed out like often a prompt would look much more like you know um like go uh like go look at like the last feedback that like

1:14:53

SPEAKER_04

roman sent me on slack and like go fix that and send him a send him a message when the pr is up and like that's the that's the full prompt and five five is intelligent enough to like figure out with like my memories and my context window like which uh and my like chronicle history like which roman am i talking about um it knows to use slack to like pull that message and then like knows possibly like what are the docs that are related to the uh to the launch and then actually like what uses an automation to like keep a track to keep track of the pr until the deploy preview is live and slack

1:15:28

SPEAKER_04

it back to him and like i i just send off one message and like it's wild for me to think that like it can just do that in like a message yeah and i think what's really magical too with the launch of last week to put things in perspective is that now you have the codex app but the codex coding agent historically could only manipulate like your files in the repo for instance and then connect to some skills or plugins to kind of bring that context but there's a lot of work you do on your computer that is not always just an api or a plugin and now if you combine like chronicle as a research

1:15:58

SPEAKER_04

preview we kind of like which kind of understands a little bit how you work on your computer but also

1:16:03

SPEAKER_01

computer use i don't know if like you all uh watching this i've tried it but i really recommend trying it because you can't really describe until you experience it like how good it is because to my knowledge there is no other computer use implementation out there where you can simply ask a task to codex and then it would kick off its own like app in the background and then its own cursor and you can still use your computer and so you can still like send like multiple of these tasks in the

1:16:30

SPEAKER_07

background for those that don't have let's say an api or a plugin the model can actually figure everything out uh including your your native app on a mac let's say like the reminders app or the notes app like anything that you use on your computer codex can also drive so now if you bring all of these things together you you have this like combo of codex and gpt 5.5 to really drive any task that you do in your days which is pretty pretty outstanding really cool kieran uh ask a question yeah i i would love to like it's it's really cool these examples and this inspires me to like actually like really kick the

1:17:08

SPEAKER_07

tires and clearly you know the model better than me since you used it even more than we did and um one question is with these examples which are very cool like can you also if you are allowed to share more concrete examples like what you do in like actual day-to-day work and how you use 5.5 because yes this is a cool example but like do you ship codex or do you uh like use harness engineering or like maybe maybe harness engineering changed with 5.5 like i would love to hear some of the like the inside stuff that you can share um around these yeah i mean um i think uh one of the i'm like trying

1:17:54

SPEAKER_07

trying to think back of like the last couple of weeks of like really leveraging 5.5 whether like some of the standout um examples but um i think one of the one of the parts is right now i've spent a lot of time sort of improving um like the codex uh documentation developers at openair.com um was like one of one of those examples and um we had this on what like sunday night um before we launched chronicle the next morning where i thought it would be helpful for people to understand the impact that chronicle has

1:18:35

SPEAKER_07

win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win win

1:18:57

SPEAKER_07

pull it into the developer's website, which is an entirely different code base, implement it, and then like read through the docs that I had drafted in a Google doc to figure out how it

1:19:08

SPEAKER_01

actually has to like document these examples and like create them. And all of those things were just sort of like running in the background. I think like one of the interesting things is that like I can contribute way more features now, even sort of like smaller delightful features to the app or like try out an idea in the app. Not all of them necessarily ship, but sort of like inspire people inside the team without necessarily doing a lot of context switching, if that makes sense. And I was chatting like on top of that, I was chatting with some early developers like you, Dan, who have had access to the model for like a couple of weeks. And like the things that came

1:19:48

SPEAKER_01

coming back in those chats were like, I am not even specifying like much of the task anymore. I just trust the model much more to figure things out, like navigate large, complex code bases and like extract what's needed, solve complex bugs that were kind of still hanging around, but now like five, five could solve them. And to your question, Kieran, on harness engineering, I was chatting with Will at an engineer at RAMP and they're using their own harness. So I was kind of curious how that worked for them. And he told me actually like complete plug and play, like from GPT 5.4 to 5.5, but better yet,

1:20:26

SPEAKER_01

5.5 started to discover new tools that other models could not discover. It's like all of the sudden it realized how to access the database, how to fetch data. It was complete plug and play, but also like new magical powers they did not have with any other model before. I love it. Thank you guys for sharing. I know you only have like a minute left, but a big question that we have not gotten to yet is GPT 5.5 is out right now in codex and in chat GPT, but it is not available in the API. Can you talk about why and when we can expect to be able to use it in the apps we're building? Yeah. We are taking like a very, you know, safe approach for this launch and we are very,

1:21:10

SPEAKER_01

very eager to give like 5.5 to everyone. And so we expect that to come extremely soon, hopefully in days maximum. You know, we're just like making sure that like the rollout goes well and that we have all of the safeguards in place for everyone to benefit from this model. But again, back to our OpenAI mission, we want to make sure that everyone can put their hands on this model. Awesome. Thank you guys so much for joining. It was a pleasure to chat with you and thank you for the model. It's great. Thank you for having us. Thank you for having us. Enjoy GPT 5.5. Of course. Thanks. See you guys.

1:21:45

SPEAKER_01

All right. So that was the OpenAI team. That was pretty cool. Kieran, Naveen, Mike, any thoughts, anything you want to share? And Laura, welcome. Laura is a staff writer at Every. Yeah. Thank you. Thank you. Laura is a staff writer at Every.

1:22:01

SPEAKER_07

Thank you. I'm not ready at all, but thank you. You're not ready. I will take you off then. I'll take you off.

1:22:10

SPEAKER_07

So Mike, Kieran, Naveen, any reactions to the OpenAI team? This is why I come in the room, Dan, so you can't surprise me. Yeah. I really want to see more computer use stuff. I think that the Chrome extension for Claude is one of the main reasons I use the model still as well. So I haven't had a chance to really kick the tires on that. But it's also the thing that all of my enterprise clients are most scared of. Yeah. Yeah. So it'll be really interesting to see how that rolls out. Which is an interesting sign. Like if people are scared of it, it might actually be the most useful

1:22:48

SPEAKER_07

thing. Like I feel like one of the things that held OpenAI back for a while is that they were super scared of unhoveling the model and just like letting it do stuff on your computer. And now it's just like, oh yeah, it just does stuff. So that's really interesting. I love, I love this. Okay, you're going to do an image generation with this sick image model and then we're going to turn it into a front end thing because it totally short circuits the issue that that we have with, okay, the taste of this model isn't as good. And so they don't even necessarily have to improve along the dimension of

1:23:18

SPEAKER_07

can it come up with a good design on its own? They just like let the image model team do that. And I think that's amazing. Yeah. Yeah. On that one, it's kind of interesting because I have a benchmark where I have an image and I say generate the website for this image, but it was so bad. It just put that image as a background. And that's what 5.5 did. So I'm just like, I'm like, oh, me too. But that's what most more. Are you sure it is thinking on? Because I've had that happen. Yes. Yes. Extra high. Yeah.

1:23:52

SPEAKER_01

It's on, but like it means that it's not always the model and it's also the harness and how you

1:23:59

SPEAKER_07

prompt. So for me, this is again like, hey, can I rewrite that in a way or like improve the harness to see, to get everything out of it. And that's why it's so great to talk with OpenAI because they have the best knowledge. Like we played with it so we can share our opinions and findings, but also they, if they can show you, hey, this is possible, then, then we can now fine tune how we do it and

1:24:25

SPEAKER_01

learn from that, which is like, I would love to see this evolve into a daily driver by just learning how to use it. That's my take on it. I really want to try my PowerPoint again now. Yeah. Yeah. Yeah. Yeah. Yeah. Yeah. Could smash it. I, I, I, I think the image, following the image is something that I observed, it has improved a lot because in the past it didn't do it. I have a couple of examples that

1:24:52

SPEAKER_05

I did recently, let me show my screen. I, I was quite shocked how well it's able to uh follow the images so i've had an opposite experience so daniel who is our good designer

1:25:07

SPEAKER_01

for monologue and every he gave me this design and this is the design oh sorry i think let me show you the old design that we have this is the old design right this is completely wipe coded it didn't fall and then this is the final design can you see any difference between the daniel's

1:25:26

SPEAKER_05

design and this one wait show us show it to us side by side if you can so that's a screenshot of the actual developed the actual app yeah yeah you can see it's a local host one right yeah yeah yeah but uh show us the the original design that i was working from yeah this is the original design oh how can i let's see if i can actually open this yeah sorry maybe let me just put it there yeah i'm now i think so it just did it was just a one shot thing or or what oh so there's a good harness here so i actually used it uh asked it to take screenshots so it started the web server it used playwright browser and then actually it said browser binary is missing

1:26:19

SPEAKER_05

because i never actually used codecs for building web apps because i always thought it's bad at it i reach out to cloud code for this kind of task but i tested it it actually downloaded the playwright i think somehow images are not loading here but it just took ton of screenshots you can see the whole

1:26:37

SPEAKER_06

thing is changed now and then i started giving some feedback it pushed it and then i gave some feedback feedback and then finally i think it's the assets so i gave a lot of uh like here it's not there i really love the comments thing uh codex has this in-app comments on the ui directly so i think i did couple of iterations to and fro but finally it's like all done in off an hour or something so that's very cool yeah all right it's like following the images really that's really cool so if you just got here we just had the open ai team on they shared a lot about a lot about gbt 5.5 um and uh and

1:27:21

SPEAKER_06

we have found that what they taught us is you can use the new image model to make front-end designs and have gbt 5.5 implement them i think that's really awesome uh now we have representing writers from all around the world we've got katie parrot staff writer at every uh katie wrote the vibe check so uh if if you saw the vibe check that we published today uh which is around here on one of my tabs anyway this is the vibe check look at that look at that see there's katie parrot is that's her name that's because

1:27:54

SPEAKER_05

she wrote it um so uh katie thank you so much for writing this for doing a lot of the testing i would love for you to walk us through what you found when you use this product for writing and i'm going to

1:28:07

SPEAKER_01

move over move our screen over to some of the stuff that some of the stuff you found yeah totally so just to set the scene a little bit um i have been i haven't touched open ai

1:28:20

SPEAKER_06

models for writing in almost a year um around the time that i think it would have been sonnet 3.5 i've dropped i just found the writing quality on claude models to be so much better that um i just didn't feel the need to use chat gpt and in fact i actually um did what i thought was breaking up with chat with chat gpt i canceled my personal subscription about three months ago and then um and honestly when i first started doing the tasks for this model i was like okay nothing new here i don't need to change anything it really wasn't until i got my hands a little bit more dirty and like went back and forth

1:28:58

SPEAKER_06

with it and saw how it took direction that i was like wow this is really good um it's um so some of the things that we found and mike can speak to his experience too um is just that it's it's it's smoother um so uh it like the logical progression from idea to idea flows like flows more naturally it it's less clever um i think the that's something that i found with opus and if you scroll down to the third bright pink um you can't read it on the screen but um yeah if you go in there and zoom really close you can see um that this is the gpt551 um and um you know at the end of q1 i sat down to review my okrs and

1:29:46

SPEAKER_06

discovered that i feel i have roughly half of them like i first read that i was like well you know there's not really much there so that's the kind of thing that i might want to like make more clever and more neurotic but um you see and like and i've just told that like capture more of my neurotic energy please and and it does that um i can't show you that because it might reveal some details about secret information about the model um but um yeah and like and and for me like a really big thing that is that's important to highlight is the speed because when you're writing with ai you know you're

1:30:20

SPEAKER_06

you're not going to get it's not going to one shot anything you're going to go back and forth you're going to be like this like i like the direction of this but it's not quite right like this hook isn't the right hook um and so you want that speed of that and that ability to take feedback and go in the direction that you want with it and i've just found that um gpt55 is faster and better at picking up on what i mean when i give a direction even when that direction is like extremely um extremely poorly phrased and very like minimal in terms of like you know direction which is something that we found with

1:30:56

SPEAKER_06

opus 4 7 is that it needs a lot of specificity in how you prompt it in order to you know follow your directions so i find gpt55 to be a more intuitive model as well so are you breaking up with claude i am probably i am thinking about breaking up with claude um the beautiful thing is because i work at every i i have subscriptions to both thank you every um but for my personal use i will probably bring my subscription level down just because i don't see myself using it as much um mike what do you think yeah for writing in particular i thought this is a very dependable model uh like there's a lot of

1:31:41

SPEAKER_06

things that i write where i just need to get the job done and and i i just want to rely on it to not say something stupid uh and uh and it does a great job of that i still prefer opus writing just because i

1:31:54

SPEAKER_01

like the neuroticism uh element like it is it's very witty right opus is very witty and and um but the problem with opus 4 7 i found is that it would say something clever and then just stop and it's like

1:32:08

SPEAKER_05

very like i'm like i asked for like a section and you've given me a paragraph uh almost like again anthropomorphizing the models but but like it feels like it's like great i've said something clever now i can like move on to the next thing uh whereas uh you know this is uh a workhorse is is the good is a good term for it i think you guys picked a really good title uh because uh it was probably katie deserves the credit there but uh but uh it is a workhorse it just like gets the job done without fuss like it just kind of churns through the work and um you know for most writing tasks i don't need like

1:32:42

SPEAKER_06

witticisms uh i need just like you know job done katie any uh any reactions to what mike just said or any anything that you found in your testing that we have not talked about yet um i think um you know like i think like um let me see i don't know um that's okay if not i can um yeah no i think that i mean that's more or less caught it's so hard to talk about writing it's like up short of just reading out the prose that it produced um you know it's you know doesn't make for good streaming but um what i will say is that i did write the um the vibe check itself with um with codex um with with gpt55

1:33:29

SPEAKER_06

and that was an awesome experience not just because of the model although well this may actually be about the model's capabilities the slack integration being able to pull information straight out of slack

1:33:41

SPEAKER_05

and when dan did what dan does and zoomed in at like 5 p.m the night before the review had to go out to flag an issue and talk it out i was able to pull the additional messages talk to to codex about what needed to change in the in the draft and get like new sections for dan to be happy with hopefully it

1:34:05

SPEAKER_01

was great yeah they came out decently um most of it made it into the into the doc um or into the published piece and so that the workflow like having that really powerful model sitting at the middle of all these integrations um is something else you know speaking more to the general knowledge worker angle which you know writing sort of nests inside of you know i just feel so much more resourceful sitting inside of this um this ecosystem which now also includes g gpt agents which i am also extremely excited about um have you tried i have not tried yet i so i built i've been working on a project to

1:34:46

SPEAKER_01

get ai to manage my okrs for me um because i failed at half of them last time and i was like can open like and i have an extreme weakness at project management like can't do timelines can't do dependencies can't do milestones can't do any of that so i just gave my okrs to the model and was like can you help me do this and so like i um i basically i talked to the i had the agent i had an agent build me a project manager um and yeah it connects to my notion task list it connects to my calendar so i can

1:35:20

SPEAKER_05

just say hey what should i do today and it will tell me what needs to ship which is great for my brain and hopefully really cool right well vibe check of the agents coming soon but uh we are we are approaching time here folks we've been streaming for about an hour and a half if you just got here you should know every is the only subscription you need to stay at the edge of ai we have published a detailed several thousand word vibe check of our three weeks of testing written by katie parrot um the overall take that we have is this model's a beast it's really good it it is both a step change in senior

1:35:58

SPEAKER_05

engineering capabilities and it is quite good at just being a fast personable like working collaborator which is a pretty rare thing to like increase your efficacy along both of those dimensions at once in a model um it scored a 62 on our senior engineer benchmark which measures 60 62 out of 100 which measures how well a model how good a model is at cleaning up a vibe coded slop code base in a way that a human senior engineer would it scored a 62 opus 4 7 scored a 30 32 33 ish um so it's like a 30 point swing which is pretty it's pretty crazy one caveat uh the on the senior engineer

1:36:40

SPEAKER_05

bench the uh 5.5 scored the best when it used an opus 4 7 plan so um so katie may be breaking up with opus or or with claude but if you really want the best get the best out of your coding capabilities i actually still think you kind of want both of these models in the mix um uh but anyway this is a really good model if you're into it you should definitely check out our vibe check uh every every dot to slash vibe check slash gpt55 or you can just go to every.to and it's linked in the home page um every like i said every is the only subscription you need to stay at the edge of ai you should subscribe at every.to

1:37:18

SPEAKER_06

slash subscribe we do ideas apps and training on the idea side we have articles like this every single day we have a daily newsletter where we keep track for everything that's going on in ai you just read one thing and you are absolutely caught up you get all of the stuff that we're learning as we as we build companies with this as we design as we write as we code um we also have a suite of apps that you get as part of the every subscription it's apps that we build for ourselves to help us do better work in ai that we include for you in the subscription so we have spiral which is your ai writing assistant with

1:37:48

SPEAKER_06

taste we've got quora which is an ai email assistant it's an it's an agent for your email and when it's coming out with an inbox soon so you can i will be doing all my email through quora very soon which is great sparkle which is a file organizer uh you can see my desktop is completely empty that is technically my second desktop but my actual desktop is also empty and that's because uh sparkle organizes it for me uh we've got monologue which is a speech text app it's like super whisper whisper flow it's built by naveen uh it's it's really really good it's super fast we also just release notes for it so you can uh you

1:38:21

SPEAKER_06

can record any note you want a voice note on a walk an idea storm a brainstorm your meetings uh they're all transcribed and all available for your agents we've also got proof this is the app that i vibe coded a couple a month or two ago that we had to rebuild because it was this vibe code was slop which became the senior engineer vibe code benchmark proof is sick it is a markdown editor for your agent it's all web-based so it's really easy to share a markdown document between different agents or between you and your colleagues so like coding plans and stuff like that so you should definitely check

1:38:50

SPEAKER_05

out proof proof editor.ai also plus ones it's our hosted open claw one click it's in your slack all of this is available as part of your every subscription every.to subscribe all these app plus all the articles we write plus we do a bunch of trainings we do live streams we do camps we do courses um we do all that stuff ideas apps and training all part of one subscription we also have an enterprise offering so the the the base every subscription is really intended for individuals and teams but uh if you're uh part of a part of a larger company a big startup trying to trying to go

1:39:27

SPEAKER_05

ai native or agent native uh we've got you covered mike do you want to talk a little bit about the consulting offering yeah sure yeah so uh so i'm a recent addition to the consulting team uh just joined in february but actually i've been uh writing for every for the past couple of years and uh before i joined every i like to joke that i kind of invent reinvented every from first principles like i was playing around with ai i was writing about it and then i was consulting so uh now i'm just doing it as part of a bigger team and more professionally uh hopefully natalia will tell you uh it's been to change i'm

1:40:03

SPEAKER_05

wearing a shirt now you know i've had a haircut no it's uh it's not that corporate here uh no but um yeah in the in the consulting team what we do primarily is uh ai transformation adoption so uh quite often

1:40:17

SPEAKER_01

a company will come to us and say we just bought cloud code for this whole team uh how do we get

1:40:22

SPEAKER_04

them to use it and how do we get them to be effective with it uh in a way that is like safe and matches like what we what we can do internally any kind of restrictions limitations uh so it's uh you know hands-on hands-on tactical stuff uh we're usually doing at least 50 of those uh workshops are uh like like actual actions like you're actually building stuff in the workshop uh so that's uh that's a big focus for us i think you can only really learn this stuff by playing around with it uh and we obviously have a little bit of theory uh to to go alongside that too uh one of the things we've been working on more

1:40:57

SPEAKER_04

recently is uh trying to take the bundle concept the membership concept uh from the the rest of the

1:41:04

SPEAKER_06

business and then thinking about what would uh subscription look like for consulting uh because time of materials doesn't really like make any sense uh anymore when uh you know my agent can be working overnight on something for you uh so uh that's something we if you're interested in that like we're trying to explore that at a minute we're talking to four or five companies uh about what that looks like uh so uh yeah happy to have that conversation so if you are uh if you want us to come train your executive leadership team if you want us to come set up agents inside your company

1:41:33

SPEAKER_06

you know where to find us otherwise you should subscribe to every and uh and if you don't that's okay we'll be back the next time there's a model release so you'll definitely see us anytime there's a new model check out every we'll tell you what we think of it um it's it's always a pleasure to do these i i love model release days honestly um super fun i never eat lunch because i'm too excited um and so it's like 3 45 here so i'm gonna probably eat some lunch but what you should know is you should always make sure to stay hydrated cheers keep your agent liquid cooled not hydrating i saw that coke zero katie that doesn't i was looking for my water bottle and i

1:42:14

SPEAKER_06

couldn't find it i got it i'll be hydrated next time all right thank you all for joining check out every and we'll see you next time

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note