I Tested GPT‑5.6 Sol for a Month
Description
Dan Shipper, CEO of Every, spent a month testing OpenAI’s GPT‑5.6 Sol. It is fast, powerful, relatively inexpensive and an excellent writer. Dan breaks down how GPT‑5.6 performs across coding, writing, design, and knowledge work, why he compares it to a Porsche, when to use it vs. Fable 5 and Opus 4.8, and how models like this let us move from doing every task ourselves to managing systems that do the work. Get the full Vibe Check https://every.to/vibe-check/gpt-5-6 How to Use Codex for Knowledge Work: https://every.to/guides/codex-for-knowledge-work?utm_source=youtube Every is the most AI-native startup on the internet. Through ideas, software and education, subscribers get the tools to work at the frontier of AI. Start your free trial today: https://every.to/subscribe?utm_source=youtube Follow Every: https://x.com/every Follow Dan Shipper: https://x.com/danshipper Timestamps: 00:00 GPT-5.6 Sol launch 00:32 New desktop app merge 01:10 Vibe check roadmap 02:19 GPT-5.6 is a Porsche 03:02 Coding tier rating 04:47 Babel bench demo 05:58 Pairing with Fable 06:31 Writing quality test 07:45 Design taste comparison 09:02 Knowledge work systems 10:04 Email and meeting agents 11:34 Personal life automations
Summary
Generated by claude-sonnet-4-5At-a-Glance
- Verdict: Watch fully
- Core thesis: GPT-5.6 Sol is the first AI model reliable and fast enough to shift knowledge workers from direct task execution to managing automated systems—the 'work on the company, not in it' paradigm now accessible beyond coding
- Why it matters: This represents a phase transition in AI utility: moving from 'AI as copilot' to 'AI as managed workforce,' particularly for non-coders in knowledge work domains like email, meetings, and decision routing
- Best use: Ken should watch fully to understand the benchmarking methodology (senior engineer test, BabelBench), see concrete workflow automation examples (Tend email system, meeting digestion), and evaluate the strategic positioning (OpenAI's ergonomic small-model bet vs. Anthropic's power-first approach)
Executive Summary
Dan Shipper (CEO, Every) positions GPT-5.6 Sol as the 'gold standard' for knowledge work—a Porsche analogy: daily-drivable luxury performance rather than exotic hypercar (Fable). Every has tested it internally for a month. The model launches in a merged ChatGPT desktop app combining 'Work' and 'Codex' modes (same underlying model, different UX). Shipper's central argument: 5.6 is the first model with sufficient reliability, speed, cost-efficiency, and writing quality to enable knowledge workers to transition from executing tasks directly to managing automated systems that execute tasks—a workflow shift coders have enjoyed but knowledge workers have not.
The video offers tier ratings across four domains: Coding (A-tier, not S-tier like Fable), Writing (better than Claude Opus 4.8 and Fable—less over-explanation, fewer AI-isms, one-shot marketing emails now viable), Design (improved over 5.5 but still inferior to Fable/Opus for visual taste), and Knowledge Work (the breakthrough category). On Every's 'senior engineer benchmark' (rewriting slop codebases from first principles), 5.6 scored 56/100 vs. Fable's 91/100, though Shipper notes variance and argues the score undersells usability. Fable produces simpler, better-abstracted code; 5.6 works but introduces unnecessary complexity. BabelBench (one-shot Borges Library of Babel game build) shows Fable delivers better graphics and unprompted polish (e.g., an About menu), but 5.6 is competent and much faster/cheaper.
Shipper's headline use case is 'Tend,' an email management system he built: Codex (powered by 5.6) processes emails into action cards; Dan approves/rejects decisions rather than drafting replies. He extends this to company-wide meeting digestion—5.6 reads transcripts, surfaces decisions, and routes updates. He also demos personal-life automation: macro tracking from photos/voice notes, Facebook Marketplace shopping, apartment decoration. The key insight: 5.6's speed and reliability make continuous-loop automation practical for tasks previously too marginal to justify AI involvement. Shipper advocates using Fable for hard coding tasks but delegating subtasks to 5.6 as a sub-agent to preserve Fable credits and gain speed.
Strategic framing: OpenAI is betting on smaller, heavily post-trained, ergonomic models vs. Anthropic's 'big model smell' (Fable is slower, pricier, spins up 100-agent fleets). Shipper believes most users most of the time prefer collaborable, fast, affordable power over world-changing genius. He positions 5.6 as the model that makes 'working on the system instead of in the system' accessible to non-coders. Every will release Tend prompts tomorrow; Shipper invites vibe-check feedback from users. No caveats about failure modes, edge cases, or risks of over-delegation are discussed.
Key Takeaways
- Claim: GPT-5.6 Sol enables knowledge workers to shift from task execution to system management—'work on the company, not in it' now available beyond coding | Evidence: Shipper's Tend app: email → action cards processed by 5.6; meeting transcripts → decision summaries routed automatically; personal life examples include macro tracking from photos, Facebook Marketplace automation. He states: 'It's just smart enough, fast enough, and reliable enough, and a good enough writer that for a lot of knowledge work tasks, you can start to abstract yourself a level up.' | Caveat: No discussion of failure modes, cases where 5.6 misroutes decisions, or the overhead of tuning/supervising systems. Also unclear how much prompt engineering or infrastructure (Codex loop, API access) is required to replicate these workflows. | Implication: Ken should evaluate whether this workflow model (AI-as-managed-workforce vs. AI-as-copilot) is already practical for agent systems, GTM ops, or content workflows. The bottleneck may have shifted from model capability to system design and supervision overhead. | Timestamp: 10:30
- Claim: On coding, 5.6 is A-tier (default daily driver) but not S-tier (Fable reserved for hardest tasks); scored 56/100 on Every's senior engineer benchmark vs. Fable's 91/100 | Evidence: Senior engineer benchmark measures rewriting 'vibe-coded slop' codebases from first principles. 5.6 completed rewrites but with more complexity/abstraction than needed; Fable produced simpler, cleaner code. BabelBench (one-shot Borges game): 5.6 delivers functional game; Fable adds unprompted About menu, better graphics. Shipper: 'I use it by default for almost everything. And then there are some coding tests that are just big and complicated and they require a lot of power. And those are the things I flip into Claude.' | Caveat: Shipper acknowledges 'a lot of variance' in benchmark scores—5.5 once scored 62.5, so 56 may undersell 5.6. Benchmark difficulty (rewriting slop codebases) may not reflect typical dev tasks. No discussion of where 5.6 fails or produces bugs. | Implication: For Ken's agent dev work, 5.6 is likely sufficient for 80% of tasks (speed/cost advantage) but Fable remains necessary for architectural rewrites or high-abstraction problems. Shipper's 'use Fable with 5.6 as sub-agent' pattern is worth testing to optimize credit spend. | Timestamp: 04:15
- Claim: 5.6 is a better writer than Claude Opus 4.8 and Fable—less over-explanation, fewer AI-isms, one-shot marketing emails now viable | Evidence: Email example: 'Tucker, 4:30 ET on Tuesday, the 14th works for me. Hope that still works on your end. Looking forward.' Shipper: 'Our head of growth, Austin, uses it to do marketing emails and he's like, this is the first time that I can actually one-shot marketing emails with this thing.' He contrasts with Fable: 'often runs for so long that it creates almost its own private language that it ends up speaking in.' | Caveat: No examples of where 5.6's brevity undershoots (e.g., misses nuance, tone, or context). Marketing email quality is subjective; no side-by-side comparison provided. Shipper's use case (internal/transactional emails) may not generalize to persuasive/creative copy. | Implication: Ken should test 5.6 for content ops (newsletters, GTM copy, social posts) where speed and 'just get it done' matter more than literary polish. If Shipper's head of growth one-shots marketing emails, this could compress content production timelines meaningfully. | Timestamp: 07:10
- Claim: 5.6 design is improved over 5.5 but still inferior to Fable/Opus—prompts image model better but lacks taste/polish | Evidence: 5.6 explains design concept before executing (e.g., 'cream-colored, warm paper aesthetic'). Image comparison: 5.6 produced 'too complicated, doesn't look well considered' image; Fable (same GPT image model, different prompting) produced polished, well-composed image one-shot. Shipper: 'Fable is just playing on a different level.' | Caveat: Both models use the same underlying GPT image model—difference is prompting strategy, not image generation capability. Shipper doesn't quantify 'better' or provide criteria (composition, color theory, branding alignment). Single example; no systematic design benchmark. | Implication: For Ken's content/brand work, 5.6 can handle functional design (UI mockups, wireframes) but Fable/Opus still required for high-stakes visual work (pitch decks, brand assets). The gap is prompting skill, not raw capability—Ken could potentially close this via custom instructions or fine-tuning. | Timestamp: 08:45
- Claim: OpenAI is betting on smaller, post-trained, ergonomic models vs. Anthropic's 'big model smell'—5.6 is likely smaller than Fable, optimized for speed/cost/collaboration over raw genius | Evidence: Shipper: 'Fable has a lot of big model smell. It's chunky. You can tell it's gigantic. And because of that, it's hard to use. And it's slow and it's expensive.' He contrasts: 'I'm pretty sure Sol is just a smaller model than Fable. And it's really well post-trained... designed to be extremely ergonomic, extremely fast, still extremely smart, but just not at the same level of world-changing genius that Fable is.' | Caveat: Shipper admits uncertainty ('I don't know for sure') about model size. No hard data on parameters, training compute, or post-training methods. 'Ergonomic' and 'world-changing genius' are subjective. Fable's agent-spawning behavior (100-agent fleets) may reflect scaffolding choices, not raw model size. | Implication: Ken should track this strategic divergence: if OpenAI wins the ergonomics bet, scaling AI ops may favor breadth (many 5.6 instances) over depth (few Fable calls). For investing thesis, question whether model size arms race plateaus and post-training differentiation becomes moat. For Ken's workflows, cost-per-task and iteration speed may matter more than peak capability. | Timestamp: 12:15
- Claim: Using Fable with 5.6 as a sub-agent is a 'match made in heaven'—preserves Fable's architectural intelligence while gaining 5.6's speed and cost efficiency | Evidence: Shipper: 'One of my favorite things to do is to go into Fable for a difficult coding task and say, I want you to use GPT 5.6 Sol as a sub-agent... your Fable credits run out. You say one thing and it spins up a fleet of 100 agents and then suddenly you're out of credits. If you use 5.6 with it, not the case.' | Caveat: No technical detail on how sub-agent delegation works (API calls, prompt structure, error handling). Unclear if Fable reliably delegates vs. attempting tasks itself. Cost savings depend on task decomposition—complex tasks may still burn Fable credits on orchestration. | Implication: Ken should experiment with this pattern in agent systems: use frontier models (Fable, O1) for planning/architecture, delegate execution to cheaper models (5.6, 4o). This could unlock cost-effective scaling for complex workflows where full-frontier-model inference is overkill. | Timestamp: 06:45
Detailed Brief
Coding Performance and Benchmarks
- Claims: 5.6 is A-tier for coding—default daily driver but not S-tier (reserved for Fable); Scored 56/100 on Every's senior engineer benchmark (rewriting slop codebases); Fable scored 91/100; BabelBench (one-shot Borges Library of Babel game) shows functional but less polished output vs. Fable; 5.6 completes rewrites but introduces unnecessary complexity/abstraction; Fable produces simpler, cleaner code
- Evidence: Senior engineer benchmark: rewrite 'vibe-coded slop' codebase from scratch, measuring first-principles re-imagination; 5.5 previously scored 62.5 on best run, indicating variance in benchmark; BabelBench comparison: 5.6 game has examine/navigation functions; Fable version has better graphics, unprompted About menu; Shipper uses 5.6 by default, switches to Fable for 'big and complicated' tasks requiring 'a lot of power'
- Caveats: High variance in benchmark scores—single 56 may not be representative; Benchmark (rewriting slop code) may not reflect typical dev workflows (greenfield, debugging, API integration); No failure mode examples—unclear where 5.6 produces bugs, security issues, or non-functional code; BabelBench is qualitative/subjective—no code quality metrics provided
- Implications: For Ken's agent dev, 5.6 is sufficient for most tasks; reserve Fable for architectural/high-abstraction problems; Cost/speed trade-offs favor 5.6 for iterative dev; Fable for one-shot complex builds; Shipper's 'Fable + 5.6 sub-agent' pattern worth testing to optimize credit spend and iteration speed; Every's benchmarks (senior engineer, BabelBench) could be useful for Ken's own model evals if made public
Writing Quality and Use Cases
- Claims: 5.6 is a better writer than Claude Opus 4.8 and Fable—less over-explanation, fewer AI-isms; First model enabling one-shot marketing emails (per Every's head of growth); To the point, simple, clearly expresses needed content without overthinking; Fast output compared to Opus/Fable (which require long waits)
- Evidence: Email example: 'Tucker, 4:30 ET on Tuesday, the 14th works for me...' shows natural, concise tone; Austin (head of growth) successfully one-shots marketing emails for first time with any model; Shipper uses it for emails, metaphors, taglines, reflecting on writing drafts; Fable 'often runs for so long that it creates almost its own private language'
- Caveats: No side-by-side writing comparisons provided—claims are subjective/anecdotal; Marketing email quality unverified (no response rates, conversion data, or external review); Use cases (internal/transactional emails) may not generalize to persuasive/creative copy; No discussion of where brevity undershoots (missing nuance, tone, context)
- Implications: Ken should test 5.6 for content ops where speed >> polish: newsletters, social posts, internal comms, GTM copy; If marketing emails are truly one-shottable, this compresses content production timelines significantly; Writing quality may be a latent differentiator for knowledge work automation—if 5.6 writes like a human, delegation becomes easier to trust; Potential to replace copywriters for certain categories (transactional, informational) but not creative/brand-critical work
Design Capabilities and Limitations
- Claims: 5.6 design improved over 5.5 (which was 'notoriously bad') but still inferior to Fable/Opus; Explains design concept before executing (e.g., 'cream-colored, warm paper aesthetic'); Both 5.6 and Fable use same GPT image model—difference is prompting, not generation capability; Fable 'playing on a different level' for design taste/polish
- Evidence: Library of Babel game shows some design taste—visual theming consistent with prompt; Image comparison: 5.6 produced 'too complicated, doesn't look well considered' image; Fable produced polished, well-composed image one-shot; Shipper: 'When you ask it to design a website, it'll say, let's think through the concept for this design'
- Caveats: Single image example—no systematic design benchmark or rubric; Subjective assessment ('well considered,' 'playing on a different level') without defined criteria; Both models using same image model means gap is prompting skill, not raw capability—potentially addressable via custom instructions; No discussion of whether design quality matters for Shipper's use cases (functional apps, internal tools)
- Implications: For Ken's brand/content work, 5.6 can handle functional design (UI mockups, wireframes, low-fidelity prototypes) but Fable/Opus required for high-stakes visuals (pitch decks, marketing assets); The gap being prompting (not generation) suggests Ken could close it via custom instructions, chain-of-thought prompts, or reference images; Design may be the domain where model choice matters least—external tools (Figma, Midjourney) still likely superior for serious visual work; Worth testing 5.6 for rapid prototyping where speed >> polish, then hand off to designer or Fable for refinement
Knowledge Work Automation—The Breakthrough Category
- Claims: 5.6 enables transition from executing tasks to managing automated systems—'work on the company, not in it' now accessible beyond coding; First model with sufficient reliability, speed, cost-efficiency, and writing quality to enable continuous-loop automation for marginal tasks; Use cases: email management (Tend app), meeting digestion, personal life automation (macro tracking, Marketplace shopping); Knowledge workers can now abstract themselves 'a level up'—tune the system, make decisions, let AI execute
- Evidence: Tend app: emails → action cards processed by 5.6; Dan approves/rejects rather than drafting replies. Over time, Codex learns preferences.; Meeting digestion: Dan leaves early, Codex reads transcript, surfaces what happened after he left, flags decisions; Personal life: photos → macro tracking; voice notes → monologue summaries; Facebook Marketplace automation; apartment decoration; Shipper: 'These are all the things that you can do if you have it running in a loop doing work for you that were previously not possible or it would require too much work. It just wouldn't be worth it.'
- Caveats: No discussion of failure modes—misrouted decisions, incorrect summaries, bad approvals; Unclear how much prompt engineering, system design, or Codex infrastructure is required to replicate workflows; Supervision overhead not quantified—how much time does Dan spend tuning vs. executing?; Examples are from Dan's specific context (CEO, small team at Every)—may not generalize to larger orgs, regulated industries, or roles requiring deep expertise; Every will release Tend prompts tomorrow but unclear if prompt alone is sufficient or if custom tooling/API access is required
- Implications: Ken should evaluate whether 'AI-as-managed-workforce' model is practical for agent systems, content ops, GTM workflows—or if supervision overhead negates gains; Bottleneck may have shifted from model capability to system design, tuning, and trust—Ken's differentiation could be in building reliable scaffolding around 5.6; For investing, question whether knowledge work automation is a horizontal platform play (ChatGPT Work) or vertical-specific (legal, finance, healthcare require specialized systems); Tend-style email/meeting automation could be high-leverage for Ken personally—worth testing to free up decision-making time; If continuous-loop automation is now viable, agent systems could shift from one-shot tasks to long-running processes (monitoring, routing, escalation)
Strategic Positioning—OpenAI vs. Anthropic
- Claims: OpenAI betting on smaller, heavily post-trained, ergonomic models vs. Anthropic's 'big model smell'; 5.6 likely smaller than Fable, optimized for speed/cost/collaboration over raw genius; Fable is 'chunky,' slow, expensive—requires expertise and budget to unlock value; Most users most of the time prefer collaborable, fast, affordable power over world-changing capability
- Evidence: Shipper: 'Fable has a lot of big model smell. It's chunky. You can tell it's gigantic. And because of that, it's hard to use.'; Fable spins up '100-agent fleets' and burns credits quickly; 5.6 is cost-efficient for continuous use; Shipper: 'I'm pretty sure Sol is just a smaller model than Fable. And it's really well post-trained... designed to be extremely ergonomic, extremely fast, still extremely smart, but just not at the same level of world-changing genius.'; Shipper and Every team spend most time in 5.6/ChatGPT Work—default choice, not Fable
- Caveats: Shipper admits uncertainty about model size ('I don't know for sure')—no hard data on parameters, compute, or post-training methods; 'Ergonomic' and 'world-changing genius' are subjective—no operationalized definitions; Fable's 100-agent behavior may reflect scaffolding/orchestration choices, not raw model size; Different strategies may serve different markets—enterprises may prefer Fable's peak capability; SMBs/individuals may prefer 5.6's accessibility
- Implications: For Ken's agent systems, scaling strategy may favor breadth (many 5.6 instances) over depth (few Fable calls)—distributed swarms of cheap agents vs. monolithic genius; For investing thesis, question whether model size arms race plateaus and post-training/RLHF becomes primary moat—implications for compute spend, data strategy, talent needs; If OpenAI wins ergonomics bet, they capture consumer/SMB market; Anthropic captures specialized/high-stakes enterprise use cases (research, compliance, high-risk decisions); Ken should track which labs invest in post-training vs. pre-training—signals about where differentiation lies; Cost-per-task and iteration speed may matter more than peak capability for most real-world workflows—implications for Ken's own tool selection and portfolio companies
Notable Concepts & Terms
- GPT-5.6 Sol: Latest OpenAI model, positioned as 'gold standard' for knowledge work—Porsche analogy (daily-drivable luxury performance). Smaller, faster, cheaper than Fable but still highly capable.
- ChatGPT Work / ChatGPT Codex: Merged desktop app—Work mode hides code, Codex mode shows developer view. Same underlying model (5.6). Shipper's preferred harness for AI work, enabling continuous-loop automation.
- Fable (Claude Opus 4.8 assumed): Anthropic's frontier model—'big model smell,' slow, expensive, S-tier for coding. Produces simpler/better code but burns credits fast. Shipper reserves for hardest tasks.
- Senior Engineer Benchmark: Every's internal eval: rewrite 'vibe-coded slop' codebase from scratch, measuring first-principles re-imagination. 5.6 scored 56/100; Fable 91/100. High variance.
- BabelBench: Qualitative coding eval: one-shot prompt to build Library of Babel game from Borges story. Tests end-to-end capability, design taste, unprompted polish. Fable outperforms 5.6.
- Tend: Email management app Dan built on Codex/5.6—turns emails into action cards, AI suggests responses, user approves/rejects. Example of 'managing the system' vs. 'doing the work.' Prompts releasing tomorrow on Every.
- 'Work on the company, not in it': Classic founder/manager truism—now enabled for knowledge workers via 5.6. Shift from task execution to system management: tune AI, make decisions, let AI execute rote work.
- Big model smell: Shipper's term for Fable—'chunky,' visibly gigantic, hard to use, slow, expensive. Reflects parameter count, training compute, inference cost. Contrast with 5.6's 'ergonomic' design.
- Sub-agent delegation: Pattern: use Fable for hard coding tasks, tell it to use 5.6 as sub-agent for execution—preserves Fable's architecture/intelligence, gains 5.6's speed/cost savings. 'Match made in heaven.'
- Post-training: Training phase after pre-training (RLHF, instruction tuning, distillation). Shipper speculates 5.6's edge is heavy post-training vs. raw scale—ergonomics/collaboration over genius.
Operator Notes / Why Ken Should Care
- Ken should test 'AI-as-managed-workforce' model for agent systems—shift from one-shot tasks to continuous-loop automation (email routing, meeting digestion, decision escalation). Tend-style workflow could be high-leverage for Ken personally.
- Fable + 5.6 sub-agent pattern worth experimenting with for cost optimization—use frontier models for planning/architecture, delegate execution to cheaper models. Could unlock scaling for complex workflows.
- 5.6 writing quality (one-shot marketing emails, natural tone, no AI-isms) suggests new viability for content ops automation—newsletters, social posts, GTM copy. Test for Every's content production pipeline.
- Strategic divergence between OpenAI (ergonomics, post-training, accessibility) and Anthropic (raw power, scale, expertise-required) has investing implications—Ken should track which approach wins and why. Affects moat sources (data, compute, talent).
- If 5.6 makes 'work on the system' accessible to non-coders, bottleneck shifts to system design and supervision overhead—Ken's differentiation could be in building reliable scaffolding (error handling, escalation, feedback loops).
- Every's benchmarks (senior engineer, BabelBench) are valuable—Ken should request methodology details or build similar internal evals for agent systems. Need operationalized rubrics beyond 'vibe.'
- Tend prompts releasing tomorrow on Every—Ken should review for applicability to his workflows. If prompt alone is insufficient, custom tooling/API access may be required (implies higher barrier to replication).
- Design quality gap (5.6 vs. Fable) is prompting, not raw capability—Ken could close via custom instructions or chain-of-thought. But for high-stakes visuals, external tools (Figma, Midjourney) still likely superior.
- 5.6 as 'Porsche' (daily-drivable luxury performance) is compelling positioning—Ken should evaluate whether this applies to agent systems (breadth of cheap agents vs. depth of expensive genius) or if peak capability still required.
- Macro tracking, Marketplace shopping, apartment decoration examples show 5.6 enables automation of previously marginal tasks—Ken should brainstorm low-value/high-frequency personal/business tasks to automate.
Watch Map
- 00:00: Introduction—GPT-5.6 Sol launch, merged ChatGPT Work/Codex app, Every's month-long testing, 'gold standard' claim
- 02:30: Porsche analogy—daily-drivable luxury performance, powerful/fast/affordable, positioning vs. Fable
- 04:15: Coding section—A-tier rating, senior engineer benchmark (56/100 vs. Fable's 91/100), variance discussion, BabelBench demo
- 06:45: Fable comparison—BabelBench side-by-side, Fable's unprompted polish, sub-agent delegation pattern (match made in heaven)
- 07:10: Writing section—better than Opus/Fable, one-shot marketing emails, email example, speed advantage
- 08:45: Design section—improved over 5.5, concept explanation, image comparison (5.6 vs. Fable), still inferior to Fable/Opus
- 10:30: Knowledge work section—transition to managing systems ('work on company not in it'), Tend email app demo, meeting digestion, personal life automation
- 12:15: Strategic positioning—OpenAI's ergonomic bet vs. Anthropic's 'big model smell,' smaller/post-trained vs. gigantic, cost/speed vs. genius
- 13:15: Conclusion—5.6 as gold standard, Every subscription pitch, request for user vibe checks in comments
Source/Metadata
- Title: GPT‑5.6 Sol Is the New Gold Standard for Work
- Transcript words: 5889
- Duration seconds: 814
- Timestamp note: Timestamps present and used throughout transcript; video duration 814 seconds (~13.5 minutes)
Transcript
It's model release day and I'm wearing sunglasses because it is the release of GPT 5.6 Sol. Sol? It's a good model. Wow, it's bright. So put on your sunglasses, get your suntan lotion, and get ready for the vibe check. We've been testing it for about a month now internally at Every with a pause in between because of the federal government. Thank you, Howard Lutnick, but we got it back. Thank you again, Howard. Today's Sol is coming out in the new ChatGPT desktop app, which is a merge of the ChatGPT desktop app and the Codex desktop app. Now they're one app and you can see it's ChatGPT work and then ChatGPT codex. And as far as I can tell, these are the same thing on the work side. It just hides the code from you. And on the codex side, it looks a little bit more developer-y, but it's the same thing under the hood. These two together, 5.6 Sol and ChatGPT codex or ChatGPT work, whatever you want to call it, are the current gold standard for best model in best harness, in particular for knowledge work. There's a few different categories of this vibe check. There's coding, there's writing, there's design, there's knowledge work. We're going to go through all the tests that we did internally at Every on this model and tell you how it stacks up with other models like Fable and Opus 4.8 and what you can do with it if you push to the limit. Because I think this model introduces a new way of thinking about and doing knowledge work that is going to be familiar to you if you're a programmer, but now it makes it available for everybody inside of ChatGPT work and ChatGPT codex. So first, who am I? And how do you already have access to the model? It just came out a minute ago. Well, I'm Dan Shipper. I'm the co-founder and CEO of Every. Every is the only subscription you need to stay at the edge of AI. You can think of us as a frontier lab for the future of knowledge work. We spend all of our time using new models before they come out, putting them through their paces, using them for everything from coding to knowledge work, to design, to writing, to marketing, to growth. We use it for everything internally. And on the day it comes out, we give you these vibe checks. They happen on YouTube. They also happen on our website, every.to. And we have a bunch of other amazing stuff as part of the subscription. So if you like this, you should go to Every and you should also subscribe down below on YouTube. Let's talk about GPT 5.6 Soul from a high-level perspective. What's the vibe check? I think of it as a Porsche. I asked 5.6 to describe a Porsche. It's a low-slung German sports car built around precision. It feels elegant rather than flashy. It's fast. It's tightly controlled. It's engineered to make every curve inviting. A Porsche is both a luxury vehicle that can go really fast and it's designed to be something that you use every day. That's what I think 5.6 Soul is. It's really powerful. It's really fast. It's relatively inexpensive. And no matter who you are and what you're using it for, whether it's coding or knowledge work or any of the other things you might use AI for, it's going to be pretty usable and it's going to impress you with its power. We'll go into that in a bit. Okay. First, the moment you've all been waiting for, GPT 5.6 Vibe Check on coding. I'm going to do an S to F tier on this model for coding. It's an A tier model. It's really, really good. It's not an S tier model that's reserved for Fable. My feeling about getting an S tier label for a coding model these days is it has to be made illegal for a little while before you can get the S tier. Howard Lutnick is the only guy that can rate something an S and 5.6 is powerful, but not that powerful. I use it by default for almost everything. And then there are some coding tests that are just big and complicated and they require a lot of power. And those are the things I flip into Claude. I spin up a Fable task and then I go back to codecs. Let's talk about the benchmarks. Our senior engineer benchmark tests how good models are at senior engineer type tasks. In particular, it asks models to rewrite a vibe coded slop codebase from scratch. And it's intended to measure how it does at that task at taking a vibe coded slop codebase and re-imagining it from first principles to be well done. GPT 5.6 got a 56 out of 100 on this benchmark. Fable got a 91 out of 100. I think the 56 is a little bit low. It undersells how good this model is. In fact, 5.5 on its best run got a 62.5. So there's a lot of variance in the score on this benchmark. It's a very hard benchmark. My takeaway from it is 5.6 actually did a very good job of rewriting the codebase. It didn't do it in a particularly senior engineer type way to the level that Fable was able to do it. For example, Fable just created a much simpler codebase with much fewer abstractions. And 5.6 did the rewrite, but it was more complicated than it needed to be. One way to get an even better sense for how this works is to look at what I've been calling Babelbench, which is I just feed the model a prompt that says build the library of Babel from the Borges story as a video game, and then we play it. That'll give you a flavor, because one number sometimes doesn't totally capture the feeling of using a model like this. Okay, this is the library of Babel game. This is a one-shot prompt. I just said build the library of Babel from the Borges story. You can examine the volumes. You can go back and forth. There's some good stuff here. Let's compare this to the same prompt to Fable, and I think that'll give you a good idea of some of the differences here. This is Fable's version, and you can just see the graphics are just a bit better. There's more details. If you go into this menu, there's an about, which I didn't tell it to write. It just decided to do. It just feels a little bit more well-considered. That's what you're going to get. 5.6 can do the job. Fable just has a little bit of that extra oomph that makes it unprecedented for really, really hard coding tasks, and 5.6 is not quite there. One thing you should know, though, is it's not necessarily either or. One of my favorite things to do is to go into Fable for a difficult coding task and say, I want you to use GPT 5.6 Sol as a sub-agent, and that's a really good way to get a lot of the smartness of Fable and then also get the power, speed, efficiency, cheaper token costs. feels a little bit more well-considered. That's what you're going to get. 5.6 can do the job. Fable just has a little bit of that extra oomph that makes it unprecedented for really, really hard coding tasks, and 5.6 is not quite there. One thing you should know, though, is it's not necessarily either or. One of my favorite things to do is to go into Fable for a difficult coding task and say, I want you to use GPT 5.6 Sol as a sub-agent, and that's a really good way to get a lot of the smartness of Fable and then also get the power, speed, efficiency, cheaper token costs of 5.6 because your Fable credits run out. You say one thing and it spins up a fleet of 100 agents and then suddenly you're out of credits. If you use 5.6 with it, not the case. They're a match made in heaven. Okay, next, writing. I think it's a better writer than 4.8. I think it's a better writer than Fable. Both 4.8 and Fable have this tendency to over-explain, to be a little bit literary. Fable in particular often runs for so long that it creates almost its own private language that it ends up speaking in. And GPT 5.6 is just to the point. It's simple. It clearly expresses the thing it needs to express. It doesn't have a ton of AI-isms. I use it, for example, to compose emails. Here's an email. Tucker, 4:30 ET on Tuesday, the 14th works for me. Hope that still works on your end. Looking forward. It just gets it. Our head of growth, Austin, uses it to do marketing emails and he's like, this is the first time that I can actually one-shot marketing emails with this thing, with any model. I'm very impressed by its ability to be such a good programmer and to have this level of—yeah, I actually want to talk to it. It's a good writer. It doesn't overthink things. It doesn't overdo it. If I'm looking for a metaphor, if I'm looking for a tagline, or if I'm asking it to reflect back to me something in a piece of writing that I'm trying to massage, it's my go-to model. Even better is it's just fast. With Opus or Fable, you're just going to be waiting for a long time. And 5.6 is just like, okay, answer, answer, simple. It's good. Okay, next, design. So design is something they've talked about a lot and they're very proud of, and the designs are better. We can go back to the Library of Babel video game. This has some design taste. It's thinking about things. One thing that you'll notice is when you ask it to design a website, for example, it'll say, let's think through the concept for this design. And then it'll be like, okay, I want it to be like a cream-colored, warm paper type feel or aesthetic. And then it'll go and do the thing. It is actually good. It's a step up from 5.5, which was notoriously bad. It is definitely not as good as Opus or Fable for the same task. I'll give you an example. I was trying to use it to make an image of this way of working that I think 5.6 makes available. And it made this. And it's fine, but it's too complicated. It doesn't look well considered. Same prompt. Look at what Fable did one shot. Same image model too. That's what Fable looks like. Fable is just playing on a different level. They're both using the GPT image model, but the way they prompt it is so different that it makes Fable make stuff like this. 5.6 makes stuff like this. So if you care about design, you're still going to need some Fable, some Opus 4.8 in your life. Okay. Now let's get to the most important part, which is knowledge work in general—all the stuff that you're doing on your computer, whether that's writing documents, researching stuff, doing reports, being on Slack, all the stuff that you probably do in your job. 5.6 ushers in this new era where you can actually move in knowledge work in a lot of ways from doing all the work yourself to managing a system that does the work. This is something that managers have been doing for a long time. Founders have been doing for a long time. It's a classic truism to be like work on the company, not in the company. Coders have been doing this for a while in AI where they've been able to, instead of prompting the model to fix the issue, they set up a system where when someone reports an issue, it kicks off an agent, the agent does the work, then reviews it, then pushes the PR, and then pushes it to production. That is now becoming possible with 5.6. It's just smart enough, fast enough, and reliable enough, and a good enough writer that for a lot of knowledge work tasks, you can start to abstract yourself a level up and work on the system that does a lot of the more rote stuff so that you can do the more interesting stuff. So an example that I've used before on this channel and I think is amazing is I use it a lot to do my email. So I have this little app that I made called Tend. We'll be releasing this as a prompt that you can use to build this system for yourself. That'll come out tomorrow on Every. So if you want it, every.to slash subscribe. Basically, it turns all my emails into these little cards, and each card has an action. So a system like this means that instead of me going through all my emails and then typing responses to each one, I'm being presented with a set of emails that Codex has already processed and decided what it thinks I should do. And then I get to say this is good or this is bad. And over time, Codex gets better and better at doing that. And 5.6 is the intelligence under the hood here that makes that possible. And it's workable for more than just emails. I have the same process running for the whole company. It goes and looks at meetings and then tells me stuff that I need to know. So for example, this is a meeting I was in, I left early, Codex read the transcript and grabbed for me what happened after I left. And it keeps me updated so that I know, here's a decision that might need to be made. And then Codex and 5.6 are going to go tell the people that need to know, here's what I might think. So really, what I'm doing is I'm the one who's tuning 5.6 inside of Codex to tell it what to pay attention to, what the next actions are. And then when it presents me with decisions to make, I just make decisions. And I think this is going to be more and more common over time. And 5.6 is the first model you can trust for this. And it's bigger than just knowledge work. It also works for your personal life. For example, I have this app right now that keeps me updated so that I know here's a decision that might need to be made. And then Codex and 5.6 are going to go tell the people that need to know, here's what I might think. So really, what I'm doing is I'm the one who's tuning 5.6 inside of Codex to tell it what to pay attention to, what the next actions are. And then when it presents me with decisions to make, I just make decisions. And I think this is going to be more and more common over time. And 5.6 is the first model you can trust for this. And it's bigger than just knowledge work. It also works for your personal life. For example, I have this app right now that it just takes all the meals I've eaten, anything that I leave in a voice note in monologue, or anything I take a picture of in Apple Photos, it just grabs it from my photos, figures out the macros and then records it for me. And 5.6 just does this working in the ChatGPT Codex app in a loop. I do this all the time. I use it for buying stuff on Facebook Marketplace or helping me decorate my apartment. These are all the things that you can do if you have it running in a loop doing work for you that were previously not possible or it would require too much work. It just wouldn't be worth it. Now we're getting to the end of it. 5.6 in the ChatGPT Codex app, ChatGPT Work app, it's the gold standard. It's where I and most of the team spend most of our time working with AI. I think it's a great model. And I think it's a great harness. I think you're also starting to see that OpenAI and Anthropic are taking slightly different approaches to how they do model releases and the kinds of models they build. Fable has a lot of big model smell. It's chunky. You can tell it's gigantic. And because of that, it's hard to use. And it's slow and it's expensive. So if you know how to use it, and if you have the money, you can do crazy things with Fable. I think OpenAI is going a different direction. I don't know for sure, but I'm pretty sure Sol is just a smaller model than Fable. And it's really well post-trained. It's designed to be extremely ergonomic, extremely fast, still extremely smart, but just not at the same level of world-changing genius that Fable is. And I actually think that's a very smart bet because for most people, most of the time, I want a model that I can collaborate with really well. That's not too expensive. And that's still very powerful. And I think 5.6 is that model. So that's the vibe check. If you want more, you should go to every.to and read our full vibe check and get the rest of the stuff that comes with the every subscription. We've got a lot of stuff for you over there. And if you're using 5.6 Sol today, I would love to get your vibe check. Leave feedback in the comments, tell me what you like, tell me what you don't. And we'll see you next time. Howard Lutnick, but we got it back. Thank you again, Howard. Today's Sol is coming out in the new ChatGPT desktop app, which is basically a merge of the ChatGPT desktop app and the Codex desktop app. Now they're one app and you can see it's ChatGPT work and then ChatGPT codex. And as far as I can tell, these are basically the same thing on the work side. It just hides the code from you. And on the codex side, it like looks a little bit more developer-y, but it's basically the same thing under the hood. These two together, 5.6 Sol and ChatGPT codex or ChatGPT work, whatever you want to call it, are the current gold standard for best model in best harness, in particular for knowledge work. There's a few different categories of this vibe check. There's coding, there's writing, there's design, there's knowledge work. We're going to go through all the tests that we did internally at Every on this model and tell you how it stacks up with other models like Fable and Opus 4.8 and really what you can do with it if you push to the limit. Because I think this model introduces a new way of thinking about and doing knowledge work that is going to be familiar to you if you're a programmer, but now it makes it available for everybody inside of ChatGPT work and ChatGPT codex. So first, who am I? And how do you already have access to the model? It just came out a minute ago. Well, I'm Dan Shipper. I'm the co-founder and CEO of Every. Every is the only subscription you need to stay at the edge of AI. You can think of us a little bit like a frontier lab for the future of knowledge work. We spend all of our time using new models before they come out, putting them through their paces, using them for everything from coding to knowledge work, to design, to writing, to marketing, to growth. We use it for everything internally. And on the day it comes out, we give you these vibe checks. They happen on YouTube. They also happen on our website, every.to. And we have a bunch of other amazing stuff as part of the subscription. So if you like this, you should go to Every and you should also subscribe down below on YouTube. Let's talk about GPT 5.6 Soul from a high-level perspective. What's the vibe check? I think of it kind of like a Porsche. I asked 5.6 to describe a Porsche. It's a low-slung German sports car built around precision. It feels elegant rather than flashy. It's fast. It's tightly controlled. It's engineered to make every curve inviting. A Porsche is both a luxury vehicle that can go really fast and it's designed to be something that you use every day. That's what I think 5.6 Soul is. It's really powerful. It's really fast. It's relatively inexpensive. And no matter who you are and what you're using it for, whether it's coding or knowledge work or any of the other things you might use AI for, it's going to be pretty usable and it's going to impress you with its power. We'll go into that in a bit. Okay. First, the moment you've all been waiting for, GPT 5.6 Vibe Check on coding. I'm going to do an S to F tier on this model for coding. It's an A tier model. It's really, really good. It's not an S tier model that's reserved for Fable. My feeling about getting an S tier label for a coding model these days is it has to be made illegal for a little while before you can get the S tier. Howard Lutnick is the only guy that can rate something an S and 5.6 is powerful, but not that powerful. I use it by default for almost everything. And then there are some coding tests that are just big and complicated and they require a lot of power. And those are the things I flip into Claude. I spin up a Fable task and then I go back to codecs. Let's talk about the benchmarks. Our senior engineer benchmark tests how good models are at senior engineer type tasks. In particular, it asks models to rewrite a Vibe Coded Slop codebase from scratch. And it's intended to measure how does it do at that task at taking a Vibe Coded Slop codebase and re-imagining it from first principles to be well done. GPT 5.6 got a 56 out of 100 on this benchmark. Fable got a 91 out of 100. I think the 56 is a little bit low, like it a little bit undersells how good this model is. In fact, 5.5 on its best run got a 62.5. So there's a lot of variance in the score on this benchmark. It's a very hard benchmark. My takeaway from it is 5.6 actually did a very good job of rewriting the codebase. It didn't do it in a particularly senior engineer type way to the level that Fable was able to do it. For example, Fable just created a much simpler codebase with much fewer abstractions. And 5.6 did the rewrite, but it was more complicated than it needed to be. One way to get an even better sense for how this works is to look at what I've been calling Babelbench, which is I just feed the model a prompt that says build the library of Babel from the Borges story as a video game, and then we play it. That'll give you like a little bit of a flavor, because one number sometimes doesn't totally capture the feeling of using a model like this. Okay, this is the library of Babel game. This is a one-shot prompt. I just said build the library of Babel from the Borges story. You can examine the volumes. You can go back and forth. There's some good stuff here. Let's compare this to the same prompt to Fable, and I think that'll give you a good idea of some of the differences here. This is Fable's version, and you can just see the graphics are just a bit better. There's more details. If you go into this menu, there's an about, which I didn't tell it to write. It just decided to do. It just feels a little bit more well-considered. That's kind of what you're going to get. 5.6 can do the job. Fable just has a little bit of that extra oomph that makes it unprecedented for really, really hard coding tasks, and 5.6 is not quite there. One thing you should know, though, is it's not necessarily either or. One of my favorite things to do is to go into Fable for a difficult coding task and say, I want you to use GPT 5.6 Sol as a sub-agent, and that's a really good way to get a lot of the smartness of Fable and then also get the power, speed, efficiency, cheaper token costs of 5.6 because your Fable credits run out like that. You say one thing and it spins up a fleet of 100 agents and then suddenly you're out of credits. If you use 5.6 with it, not the case. They're kind of a match made in heaven. Okay, next, writing. I think it's a better writer than 4.8. I think it's a better writer than Fable. Both 4.8 and Fable have this tendency to over-explain, to be a little bit literary. Fable in particular often runs for so long that it creates almost its own private language that it ends up speaking in. And GPT 5.6 is just to the point. It's simple. It clearly expresses the thing it needs to express. It doesn't have a ton of AI-isms. I use it, for example, to compose emails and like, here's an email. Tucker, 4.30 ET on Tuesday, the 14th works for me. Hope that still works on your end. Looking forward. It just sort of gets it. Our head of growth, Austin, uses it to do marketing emails and he's like, this is the first time that I can actually one-shot marketing emails with this thing, with any model. I'm very impressed by its ability to be such a good programmer and to have this level of, yeah, I actually want to talk to it. It's a good writer. It doesn't overthink things. It doesn't overdo it. If I'm looking for a metaphor, if I'm looking for a tagline, or if I'm asking it to reflect back to me something in a piece of writing that I'm trying to massage, it's my go-to model. Even better is it's just fast. With Opus or Fable, you're just going to be waiting for a long time. And 5.6 is just like, okay, answer, answer, simple. It's good. Okay, next, design. So design is something they've talked about a lot and they're very proud of, and the designs are better. We can go back to the Library of Babel video game. This has some design taste. It's like thinking about things. One thing that you'll notice is when you ask it to design a website, for example, it'll say, let's think through the concept for this design. And then it'll be like, okay, I want it to be like a cream-colored, warm paper type feel or aesthetic. And then it'll go and do the thing. It is actually good. It's a step up from 5.5, which was notoriously bad. It is definitely not as good as Opus or Fable for the same task. I'll give you an example. I was trying to use it to make an image of this way of working that I think 5.6 makes available. And it made this. And it's fine, but it's, I don't know, it's too complicated. It doesn't look well considered. Same prompt. Look at what Fable did one shot. Same image model too. That's what Fable looks like. Fable is just playing on a different level. They're both using the GBT image model, but the way they prompt it is so different that it makes, Fable makes stuff like this. 5.6 makes stuff like this. So if you care about design, you're still going to need some Fable, some Opus 4.8 in your life. Okay. Now let's get to the most important part, which is knowledge work in general, like all the stuff that you're doing on your computer, whether that's writing documents, researching stuff, doing reports, being on Slack, all the stuff that you probably do in your job. 5.6 ushers in this new era where you can actually move in knowledge work in a lot of ways from doing all the work yourself to managing a system that does the work. This is something that managers have been doing for a long time. Founders have been doing for a long time. It's a kind of classic truism to be like work on the company, not in the company. Coders have been doing this for a while in AI where they've been able to, instead of prompting the model to fix the issue, they set up a system where when someone reports an issue, it kicks off an agent, the agent does the work, then reviews it, then pushes the PR, and then pushes it to production. That is now becoming possible with 5.6. It's just smart enough, fast enough, and reliable enough, and a good enough writer that for a lot of knowledge work tasks, you can start to abstract yourself a level up and work on the system that does a lot of the more rote stuff so that you can do the more interesting stuff. So an example that I've used before on this channel and I think is amazing is I use it a lot to do my email. So I have this little app that I made called Tend. We'll be releasing this as a prompt that you can use to build this system for yourself. That'll come out tomorrow on Every. So if you want it, every.to slash subscribe. Basically, it turns all my emails into these little cards, and each card has an action. So a system like this means that instead of me going through all my emails and then typing responses to each one, I'm being presented with a set of emails that Codex has already processed and decided what it thinks I should do. And then I get to say this is good or this is bad. And over time, Codex gets better and better at doing that. And 5.6 is the intelligence under the hood here that makes that possible. And it's workable for more than just emails. I have the same process running for the whole company. It goes and looks at meetings and then tells me stuff that I need to know. So for example, this is a meeting I was in, I left early, Codex read the transcript and grabbed for me what happened after I left. And it keeps me updated so that I know, here's a decision that might need to be made. And then Codex and 5.6 are going to go tell the people that need to know, here's what I might think. So really, what I'm doing is I'm the one who's tuning 5.6 inside of Codex to tell it what to pay attention to, what the next actions are. And then when it presents me with decisions to make, I just make decisions. And I think this is going to be more and more common over time. And 5.6 is the first model you can trust for this. And it's bigger than just knowledge work. It also works for your personal life. For example, I have this app right now that it just takes all the meals I've eaten, anything that I leave in a voice note in monologue, or anything I take a picture of in Apple Photos, it just grabs it from my photos, figures out the macros and then records it for me. And 5.6 just does this working in the ChatGPT Codex app in a loop. I do this all the time. I use it for buying stuff on Facebook Marketplace or helping me decorate my apartment. These are all the things that you can do if you have it running in a loop doing work for you that were previously not possible or it would require too much work. It just wouldn't be worth it. Now we're getting to the end of it. 5.6 in the ChatGPT Codex app, ChatGPT Work app, it's the gold standard. It's where I and most of the team spend most of our time working with AI. I think it's a great model. And I think it's a great harness. I think you're also starting to see that OpenAI and Anthropic are taking slightly different approaches to how they do model releases and the kinds of models they build. Fable is like, it has a lot of big model smell. It's like, it's chunky. You can tell it's like, it's gigantic. And because of that, it's hard to use. And it's slow and it's expensive. So if you know how to use it, and if you have the money, you can do crazy things with Fable. I think OpenAI is going a little bit of a different direction. I don't know for sure, but I'm pretty sure Sol is just a smaller model than Fable. And it's just really well post-trained. It's just designed to be extremely ergonomic, extremely fast, still extremely smart, but just not at the same level of world-changing genius that Fable is. And I actually think that's a very smart bet because for most people, most of the time, I want a model that I can collaborate with really well. That's not too expensive. And that's still very powerful. And I think 5.6 is that model. So that's the vibe check. If you want more, you should go to every.to and read our full vibe check and get the rest of the stuff that comes with the every subscription. We've got a lot of stuff for you over there. And if you're using 5.6 Sol today, I would love to get your vibe check. Leave feedback in the comments, tell me what you like, tell me what you don't. And we'll see you next time.