SPEAKER_00
It's model release day and I'm wearing sunglasses because it is the release of GPT 5.6 Sol. Sol? It's a good model. Wow, it's bright. So put on your sunglasses, get your suntan lotion, and get ready for the vibe check. We've been testing it for about a month now internally at Every with a pause in between because of the federal government. Thank you, Howard Lutnick, but we got it back. Thank you again, Howard. Today's Sol is coming out in the new ChatGPT desktop app, which is a merge of the ChatGPT desktop app and the Codex desktop app. Now they're one app and you can see it's ChatGPT work and then ChatGPT codex. And as far as I can tell, these are the same thing on the work side. It just hides the code from you. And on the codex side, it looks a little bit more developer-y, but it's the same thing under the hood. These two together, 5.6 Sol and ChatGPT codex or ChatGPT work, whatever you want to call it, are the current gold standard for best model in best harness, in particular for knowledge work. There's a few different categories of this vibe check. There's coding, there's writing, there's design, there's knowledge work. We're going to go through all the tests that we did internally at Every on this model and tell you how it stacks up with other models like Fable and Opus 4.8 and what you can do with it if you push to the limit. Because I think this model introduces a new way of thinking about and doing knowledge work that is going to be familiar to you if you're a programmer, but now it makes it available for everybody inside of ChatGPT work and ChatGPT codex. So first, who am I? And how do you already have access to the model? It just came out a minute ago. Well, I'm Dan Shipper. I'm the co-founder and CEO of Every. Every is the only subscription you need to stay at the edge of AI. You can think of us as a frontier lab for the future of knowledge work. We spend all of our time using new models before they come out, putting them through their paces, using them for everything from coding to knowledge work, to design, to writing, to marketing, to growth. We use it for everything internally. And on the day it comes out, we give you these vibe checks. They happen on YouTube. They also happen on our website, every.to. And we have a bunch of other amazing stuff as part of the subscription. So if you like this, you should go to Every and you should also subscribe down below on YouTube. Let's talk about GPT 5.6 Soul from a high-level perspective. What's the vibe check? I think of it as a Porsche. I asked 5.6 to describe a Porsche. It's a low-slung German sports car built around precision. It feels elegant rather than flashy. It's fast. It's tightly controlled. It's engineered to make every curve inviting. A Porsche is both a luxury vehicle that can go really fast and it's designed to be something that you use every day. That's what I think 5.6 Soul is. It's really powerful. It's really fast. It's relatively inexpensive. And no matter who you are and what you're using it for, whether it's coding or knowledge work or any of the other things you might use AI for, it's going to be pretty usable and it's going to impress you with its power. We'll go into that in a bit. Okay. First, the moment you've all been waiting for, GPT 5.6 Vibe Check on coding. I'm going to do an S to F tier on this model for coding. It's an A tier model. It's really, really good. It's not an S tier model that's reserved for Fable. My feeling about getting an S tier label for a coding model these days is it has to be made illegal for a little while before you can get the S tier. Howard Lutnick is the only guy that can rate something an S and 5.6 is powerful, but not that powerful. I use it by default for almost everything. And then there are some coding tests that are just big and complicated and they require a lot of power. And those are the things I flip into Claude. I spin up a Fable task and then I go back to codecs. Let's talk about the benchmarks. Our senior engineer benchmark tests how good models are at senior engineer type tasks. In particular, it asks models to rewrite a vibe coded slop codebase from scratch. And it's intended to measure how it does at that task at taking a vibe coded slop codebase and re-imagining it from first principles to be well done. GPT 5.6 got a 56 out of 100 on this benchmark. Fable got a 91 out of 100. I think the 56 is a little bit low. It undersells how good this model is. In fact, 5.5 on its best run got a 62.5. So there's a lot of variance in the score on this benchmark. It's a very hard benchmark. My takeaway from it is 5.6 actually did a very good job of rewriting the codebase. It didn't do it in a particularly senior engineer type way to the level that Fable was able to do it. For example, Fable just created a much simpler codebase with much fewer abstractions. And 5.6 did the rewrite, but it was more complicated than it needed to be. One way to get an even better sense for how this works is to look at what I've been calling Babelbench, which is I just feed the model a prompt that says build the library of Babel from the Borges story as a video game, and then we play it. That'll give you a flavor, because one number sometimes doesn't totally capture the feeling of using a model like this. Okay, this is the library of Babel game. This is a one-shot prompt. I just said build the library of Babel from the Borges story. You can examine the volumes. You can go back and forth. There's some good stuff here. Let's compare this to the same prompt to Fable, and I think that'll give you a good idea of some of the differences here. This is Fable's version, and you can just see the graphics are just a bit better. There's more details. If you go into this menu, there's an about, which I didn't tell it to write. It just decided to do. It just feels a little bit more well-considered. That's what you're going to get. 5.6 can do the job. Fable just has a little bit of that extra oomph that makes it unprecedented for really, really hard coding tasks, and 5.6 is not quite there. One thing you should know, though, is it's not necessarily either or. One of my favorite things to do is to go into Fable for a difficult coding task and say, I want you to use GPT 5.6 Sol as a sub-agent, and that's a really good way to get a lot of the smartness of Fable and then also get the power, speed, efficiency, cheaper token costs.
SPEAKER_00
feels a little bit more well-considered. That's what you're going to get. 5.6 can do the job.
SPEAKER_00
Fable just has a little bit of that extra oomph that makes it unprecedented for really, really hard coding tasks, and 5.6 is not quite there. One thing you should know, though, is it's not necessarily either or. One of my favorite things to do is to go into Fable for a difficult coding task and say, I want you to use GPT 5.6 Sol as a sub-agent, and that's a really good way to get a lot of the smartness of Fable and then also get the power, speed, efficiency, cheaper token costs of 5.6 because your Fable credits run out. You say one thing and it spins up a fleet of 100 agents and then suddenly you're out of credits. If you use 5.6 with it, not the case. They're a match made in heaven. Okay, next, writing. I think it's a better writer than 4.8. I think it's a better writer than Fable. Both 4.8 and Fable have this tendency to over-explain, to be a little bit literary. Fable in particular often runs for so long that it creates almost its own private language that it ends up speaking in. And GPT 5.6 is just to the point. It's simple. It clearly expresses the thing it needs to express. It doesn't have a ton of AI-isms. I use it, for example, to compose emails. Here's an email. Tucker, 4:30 ET on Tuesday, the 14th works for me. Hope that still works on your end. Looking forward. It just gets it. Our head of growth, Austin, uses it to do marketing emails and he's like, this is the first time that I can actually one-shot marketing emails with this thing, with any model. I'm very impressed by its ability to be such a good programmer and to have this level of—yeah, I actually want to talk to it. It's a good writer. It doesn't overthink things. It doesn't overdo it. If I'm looking for a metaphor, if I'm looking for a tagline, or if I'm asking it to reflect back to me something in a piece of writing that I'm trying to massage, it's my go-to model. Even better is it's just fast. With Opus or Fable, you're just going to be waiting for a long time. And 5.6 is just like, okay, answer, answer, simple. It's good. Okay, next, design. So design is something they've talked about a lot and they're very proud of, and the designs are better. We can go back to the Library of Babel video game. This has some design taste. It's thinking about things. One thing that you'll notice is when you ask it to design a website, for example, it'll say, let's think through the concept for this design. And then it'll be like, okay, I want it to be like a cream-colored, warm paper type feel or aesthetic. And then it'll go and do the thing. It is actually good. It's a step up from 5.5, which was notoriously bad. It is definitely not as good as Opus or Fable for the same task. I'll give you an example. I was trying to use it to make an image of this way of working that I think 5.6 makes available. And it made this. And it's fine, but it's too complicated. It doesn't look well considered. Same prompt. Look at what Fable did one shot. Same image model too. That's what Fable looks like. Fable is just playing on a different level. They're both using the GPT image model, but the way they prompt it is so different that it makes Fable make stuff like this. 5.6 makes stuff like this. So if you care about design, you're still going to need some Fable, some Opus 4.8 in your life. Okay. Now let's get to the most important part, which is knowledge work in general—all the stuff that you're doing on your computer, whether that's writing documents, researching stuff, doing reports, being on Slack, all the stuff that you probably do in your job. 5.6 ushers in this new era where you can actually move in knowledge work in a lot of ways from doing all the work yourself to managing a system that does the work. This is something that managers have been doing for a long time. Founders have been doing for a long time. It's a classic truism to be like work on the company, not in the company. Coders have been doing this for a while in AI where they've been able to, instead of prompting the model to fix the issue, they set up a system where when someone reports an issue, it kicks off an agent, the agent does the work, then reviews it, then pushes the PR, and then pushes it to production. That is now becoming possible with 5.6. It's just smart enough, fast enough, and reliable enough, and a good enough writer that for a lot of knowledge work tasks, you can start to abstract yourself a level up and work on the system that does a lot of the more rote stuff so that you can do the more interesting stuff. So an example that I've used before on this channel and I think is amazing is I use it a lot to do my email. So I have this little app that I made called Tend. We'll be releasing this as a prompt that you can use to build this system for yourself. That'll come out tomorrow on Every. So if you want it, every.to slash subscribe. Basically, it turns all my emails into these little cards, and each card has an action. So a system like this means that instead of me going through all my emails and then typing responses to each one, I'm being presented with a set of emails that Codex has already processed and decided what it thinks I should do. And then I get to say this is good or this is bad. And over time, Codex gets better and better at doing that. And 5.6 is the intelligence under the hood here that makes that possible. And it's workable for more than just emails. I have the same process running for the whole company. It goes and looks at meetings and then tells me stuff that I need to know. So for example, this is a meeting I was in, I left early, Codex read the transcript and grabbed for me what happened after I left. And it keeps me updated so that I know, here's a decision that might need to be made. And then Codex and 5.6 are going to go tell the people that need to know, here's what I might think. So really, what I'm doing is I'm the one who's tuning 5.6 inside of Codex to tell it what to pay attention to, what the next actions are. And then when it presents me with decisions to make, I just make decisions. And I think this is going to be more and more common over time. And 5.6 is the first model you can trust for this. And it's bigger than just knowledge work. It also works for your personal life. For example, I have this app right now that
SPEAKER_00
keeps me updated so that I know here's a decision that might need to be made. And then Codex and 5.6 are going to go tell the people that need to know, here's what I might think. So really, what I'm doing is I'm the one who's tuning 5.6 inside of Codex to tell it what to pay attention to, what the next actions are. And then when it presents me with decisions to make, I just make decisions. And I think this is going to be more and more common over time. And 5.6 is the first model you can trust for this. And it's bigger than just knowledge work. It also works for your personal life. For example, I have this app right now that it just takes all the meals I've eaten, anything that I leave in a voice note in monologue, or anything I take a picture of in Apple Photos, it just grabs it from my photos, figures out the macros and then records it for me. And 5.6 just does this working in the ChatGPT Codex app in a loop. I do this all the time. I use it for buying stuff on Facebook Marketplace or helping me decorate my apartment. These are all the things that you can do if you have it running in a loop doing work for you that were previously not possible or it would require too much work. It just wouldn't be worth it. Now we're getting to the end of it. 5.6 in the ChatGPT Codex app, ChatGPT Work app, it's the gold standard. It's where I and most of the team spend most of our time working with AI. I think it's a great model. And I think it's a great harness. I think you're also starting to see that OpenAI and Anthropic are taking slightly different approaches to how they do model releases and the kinds of models they build. Fable has a lot of big model smell. It's chunky. You can tell it's gigantic. And because of that, it's hard to use. And it's slow and it's expensive. So if you know how to use it, and if you have the money, you can do crazy things with Fable. I think OpenAI is going a different direction. I don't know for sure, but I'm pretty sure Sol is just a smaller model than Fable. And it's really well post-trained. It's designed to be extremely ergonomic, extremely fast, still extremely smart, but just not at the same level of world-changing genius that Fable is. And I actually think that's a very smart bet because for most people, most of the time, I want a model that I can collaborate with really well. That's not too expensive. And that's still very powerful. And I think 5.6 is that model. So that's the vibe check. If you want more, you should go to every.to and read our full vibe check and get the rest of the stuff that comes with the every subscription. We've got a lot of stuff for you over there. And if you're using 5.6 Sol today, I would love to get your vibe check. Leave feedback in the comments, tell me what you like, tell me what you don't. And we'll see you next time.
SPEAKER_00
Howard Lutnick, but we got it back. Thank you again, Howard. Today's Sol is coming out in the new ChatGPT desktop app, which is basically a merge of the ChatGPT desktop app and the Codex desktop app. Now they're one app and you can see it's ChatGPT work and then ChatGPT codex. And as far as I can tell, these are basically the same thing on the work side. It just hides the code from you. And on the codex side, it like looks a little bit more developer-y, but it's basically the same thing under the hood. These two together, 5.6 Sol and ChatGPT codex or ChatGPT work, whatever you want to
SPEAKER_00
call it, are the current gold standard for best model in best harness, in particular for knowledge work. There's a few different categories of this vibe check. There's coding, there's writing, there's design, there's knowledge work. We're going to go through all the tests that we did internally at Every on this model and tell you how it stacks up with other models like Fable and Opus 4.8 and really what you can do with it if you push to the limit. Because I think this model introduces a new way of thinking about and doing knowledge work that is going to be familiar to you if you're a programmer, but now it makes it available for everybody inside of ChatGPT work
SPEAKER_00
and ChatGPT codex. So first, who am I? And how do you already have access to the model? It just came out a minute ago. Well, I'm Dan Shipper. I'm the co-founder and CEO of Every. Every is the only subscription you need to stay at the edge of AI. You can think of us a little bit like a frontier lab for the future of knowledge work. We spend all of our time using new models before they come out, putting them through their paces, using them for everything from coding to knowledge work, to design, to writing, to marketing, to growth. We use it for everything internally. And on the day it comes out, we give you these vibe checks. They happen on YouTube. They also
SPEAKER_00
happen on our website, every.to. And we have a bunch of other amazing stuff as part of the subscription. So if you like this, you should go to Every and you should also subscribe down below on YouTube. Let's talk about GPT 5.6 Soul from a high-level perspective. What's the vibe check? I think of it kind of like a Porsche. I asked 5.6 to describe a Porsche. It's a low-slung German sports car built around precision. It feels elegant rather than flashy. It's fast. It's tightly controlled. It's engineered to make every curve inviting. A Porsche is both a luxury vehicle that can go really fast and it's designed to be something that you use every day. That's what I think
SPEAKER_00
5.6 Soul is. It's really powerful. It's really fast. It's relatively inexpensive. And no matter who you are and what you're using it for, whether it's coding or knowledge work or any of the other things you might use AI for, it's going to be pretty usable and it's going to impress you with its power. We'll go into that in a bit. Okay. First, the moment you've all been waiting for, GPT 5.6 Vibe Check on coding. I'm going to do an S to F tier on this model for coding. It's an A tier model. It's really, really good. It's not an S tier model that's reserved for Fable. My feeling about getting an S tier label for a coding model these days is it has to be
SPEAKER_00
made illegal for a little while before you can get the S tier. Howard Lutnick is the only guy that can rate something an S and 5.6 is powerful, but not that powerful. I use it by default for almost everything. And then there are some coding tests that are just big and complicated and they require a lot of power. And those are the things I flip into Claude. I spin up a Fable task and then I go back to codecs. Let's talk about the benchmarks. Our senior engineer benchmark tests how good models are at senior engineer type tasks. In particular, it asks models to rewrite a Vibe Coded Slop codebase
SPEAKER_00
from scratch. And it's intended to measure how does it do at that task at taking a Vibe Coded Slop codebase and re-imagining it from first principles to be well done. GPT 5.6 got a 56 out of 100 on this benchmark. Fable got a 91 out of 100. I think the 56 is a little bit low, like it a little bit undersells how good this model is. In fact, 5.5 on its best run got a 62.5. So there's a lot of variance in the score on this benchmark. It's a very hard benchmark. My takeaway from it is 5.6 actually did a very good job of rewriting the codebase. It didn't do it in a particularly senior engineer type way to the level
SPEAKER_00
that Fable was able to do it. For example, Fable just created a much simpler codebase with much fewer abstractions. And 5.6 did the rewrite, but it was more complicated than it needed to be. One way to get an even better sense for how this works is to look at what I've been calling Babelbench, which is I just feed the model a prompt that says build the library of Babel from the Borges story as a video game, and then we play it. That'll give you like a little bit of a flavor, because one number sometimes doesn't totally capture the feeling of using a model like this. Okay, this is the library of Babel game. This is a
SPEAKER_00
one-shot prompt. I just said build the library of Babel from the Borges story. You can examine the volumes. You can go back and forth. There's some good stuff here. Let's compare this to the same prompt to Fable, and I think that'll give you a good idea of some of the differences here. This is Fable's version, and you can just see the graphics are just a bit better. There's more details. If you go into this menu, there's an about, which I didn't tell it to write. It just decided to do. It just feels a little bit more well-considered. That's kind of what you're going to get. 5.6 can do the job.
SPEAKER_00
Fable just has a little bit of that extra oomph that makes it unprecedented for really, really hard coding tasks, and 5.6 is not quite there. One thing you should know, though, is it's not necessarily either or. One of my favorite things to do is to go into Fable for a difficult coding task and say, I want you to use GPT 5.6 Sol as a sub-agent, and that's a really good way to get a lot of the smartness of Fable and then also get the power, speed, efficiency, cheaper token costs of 5.6 because your Fable credits run out like that. You say one thing and it spins up a fleet of 100 agents and then suddenly you're out of credits. If you use 5.6 with it, not the case.
SPEAKER_00
They're kind of a match made in heaven. Okay, next, writing. I think it's a better writer than 4.8. I think it's a better writer than Fable. Both 4.8 and Fable have this tendency to over-explain, to be a little bit literary. Fable in particular often runs for so long that it creates almost its own private language that it ends up speaking in. And GPT 5.6 is just to the point. It's simple. It clearly expresses the thing it needs to express. It doesn't have a ton of AI-isms. I use it, for example, to compose emails and like, here's an email. Tucker, 4.30 ET on Tuesday, the 14th works
SPEAKER_00
for me. Hope that still works on your end. Looking forward. It just sort of gets it. Our head of growth, Austin, uses it to do marketing emails and he's like, this is the first time that I can actually one-shot marketing emails with this thing, with any model. I'm very impressed by its ability to be such a good programmer and to have this level of, yeah, I actually want to talk to it. It's a good writer. It doesn't overthink things. It doesn't overdo it. If I'm looking for a metaphor, if I'm looking for a tagline, or if I'm asking it to reflect back to me something in a piece of
SPEAKER_00
writing that I'm trying to massage, it's my go-to model. Even better is it's just fast. With Opus or Fable, you're just going to be waiting for a long time. And 5.6 is just like, okay, answer, answer, simple. It's good. Okay, next, design. So design is something they've talked about a lot and they're very proud of, and the designs are better. We can go back to the Library of Babel video game. This has some design taste. It's like thinking about things. One thing that you'll notice is when you ask it to design a website, for example, it'll say, let's think through the concept for this
SPEAKER_00
design. And then it'll be like, okay, I want it to be like a cream-colored, warm paper type feel or aesthetic. And then it'll go and do the thing. It is actually good. It's a step up from 5.5, which was notoriously bad. It is definitely not as good as Opus or Fable for the same task. I'll give you an example. I was trying to use it to make an image of this way of working that I think 5.6 makes available. And it made this. And it's fine, but it's, I don't know, it's too complicated. It doesn't look well considered. Same prompt. Look at what Fable did one shot. Same image model too. That's what Fable looks like. Fable is just playing on a different level. They're both using
SPEAKER_00
the GBT image model, but the way they prompt it is so different that it makes, Fable makes stuff like this. 5.6 makes stuff like this. So if you care about design, you're still going to need some Fable, some Opus 4.8 in your life. Okay. Now let's get to the most important part, which is knowledge work in general, like all the stuff that you're doing on your computer, whether that's writing documents, researching stuff, doing reports, being on Slack, all the stuff that you probably do in your job. 5.6 ushers in this new era where you can actually move in knowledge work in a lot of ways
SPEAKER_00
from doing all the work yourself to managing a system that does the work. This is something that managers have been doing for a long time. Founders have been doing for a long time. It's a kind of classic truism to be like work on the company, not in the company. Coders have been doing this for a while in AI where they've been able to, instead of prompting the model to fix the issue, they set up a system where when someone reports an issue, it kicks off an agent, the agent does the work, then reviews it, then pushes the PR, and then pushes it to production. That is now becoming possible with 5.6.
SPEAKER_00
It's just smart enough, fast enough, and reliable enough, and a good enough writer that for a lot of knowledge work tasks, you can start to abstract yourself a level up and work on the system that does a lot of the more rote stuff so that you can do the more interesting stuff. So an example that I've used before on this channel and I think is amazing is I use it a lot to do my email. So I have this little app that I made called Tend. We'll be releasing this as a prompt that you can use to build this system for yourself. That'll come out tomorrow on Every. So if you want it, every.to
SPEAKER_00
slash subscribe. Basically, it turns all my emails into these little cards, and each card has an action. So a system like this means that instead of me going through all my emails and then typing responses to each one, I'm being presented with a set of emails that Codex has already processed and decided what it thinks I should do. And then I get to say this is good or this is bad. And over time, Codex gets better and better at doing that. And 5.6 is the intelligence under the hood here that makes that possible. And it's workable for more than just emails. I have the same process running for the whole
SPEAKER_00
company. It goes and looks at meetings and then tells me stuff that I need to know. So for example, this is a meeting I was in, I left early, Codex read the transcript and grabbed for me what happened after I left. And it keeps me updated so that I know, here's a decision that might need to be made. And then Codex and 5.6 are going to go tell the people that need to know, here's what I might think. So really, what I'm doing is I'm the one who's tuning 5.6 inside of Codex to tell it what to pay attention to, what the next actions are. And then when it presents me with decisions to make, I just make decisions. And I think this is going to be more and
SPEAKER_00
more common over time. And 5.6 is the first model you can trust for this. And it's bigger than just knowledge work. It also works for your personal life. For example, I have this app right now that it just takes all the meals I've eaten, anything that I leave in a voice note in monologue, or anything I take a picture of in Apple Photos, it just grabs it from my photos, figures out the macros and then records it for me. And 5.6 just does this working in the ChatGPT Codex app in a loop. I do this all the time. I use it for buying stuff on Facebook Marketplace or helping me decorate my
SPEAKER_00
apartment. These are all the things that you can do if you have it running in a loop doing work for you that were previously not possible or it would require too much work. It just wouldn't be worth it. Now we're getting to the end of it. 5.6 in the ChatGPT Codex app, ChatGPT Work app, it's the gold standard. It's where I and most of the team spend most of our time working with AI. I think it's a great model. And I think it's a great harness. I think you're also starting to see that OpenAI and Anthropic are taking slightly different approaches to how they do model releases and the kinds of models they build. Fable is like, it has a lot of big model smell. It's like,
SPEAKER_00
it's chunky. You can tell it's like, it's gigantic. And because of that, it's hard to use. And it's slow and it's expensive. So if you know how to use it, and if you have the money, you can do crazy things with Fable. I think OpenAI is going a little bit of a different direction. I don't know for sure, but I'm pretty sure Sol is just a smaller model than Fable. And it's just really well post-trained. It's just designed to be extremely ergonomic, extremely fast, still extremely smart, but just not at the same level of world-changing genius that Fable is. And I actually think that's a very smart bet because for most people, most of the time,
SPEAKER_00
I want a model that I can collaborate with really well. That's not too expensive. And that's still very powerful. And I think 5.6 is that model. So that's the vibe check. If you want more, you should go to every.to and read our full vibe check and get the rest of the stuff that comes with the every subscription. We've got a lot of stuff for you over there. And if you're using 5.6 Sol today, I would love to get your vibe check. Leave feedback in the comments, tell me what you like, tell me what you don't. And we'll see you next time.