Every

LET'S RIP FABLE TOKENS FROM THE JACUZZI

3911 summary words 17 min summary Watch video

Start with the signal

17 min read

Summary

At-a-Glance

  • Verdict: Skim
  • Core thesis: Claude Opus 5 (Fable) has returned after export controls were lifted, prompting immediate real-world testing for copy-editing benchmarks, token spend audits, and code review tasks—but new safety classifiers and the absence of EAP may limit production utility.
  • Why it matters: First live operator stress-test of Fable post-export controls; reveals actual capabilities, cost discipline, and workflow integration challenges that EAP insiders skipped.
  • Best use: Ken should watch the first 15 minutes for Fable's verdict on Every's copy-edit experiments and benchmark architecture (Kate Bench), then skip to operator notes for workflow integration lessons and cost/safety trade-offs.

Executive Summary

Dan Shipper (Every CEO) streams live from Cabo as Claude Opus 5 (Fable) returns after U.S. export controls are lifted. He immediately sets Fable to work on two high-stakes internal projects: reviewing four months of failed copy-editing experiments (Kate Bench—an attempt to replicate Kate Lee's human editorial judgment with 30,000 historical edits) and auditing three months of personal token spend across ChatGPT, Claude, Codex, and Google AI Studio. Fable delivers a sharp diagnosis within minutes: Every's benchmark team spent 180 agent threads optimizing for 70% recall against Kate's exact historical edits without first validating that Kate herself is consistent. Fable estimates Kate's self-consistency ceiling is ~47–50%, meaning the models weren't failing—they were near the human limit. It recommends switching from strict historical recall to a live suggestion loop where Kate accepts/rejects edits in real time, creating a continuously improving label factory.

The stream reveals three immediate operator concerns: (1) Fable's new safety classifiers triggered on innocuous commands (process kill + browser automation) and silently downgraded to Opus 4.8 mid-task; (2) memory persistence and context bloat are now cost-sensitive at Fable's price point, pushing operators to disable memory; (3) concurrent model releases without early-access programs (EAP) mean users encounter rough edges—Sonnet 3.5 v2 (released the same week) received universally negative internal reviews except from one editor using it in short collaborative loops, not long agentic runs. Mike Taylor (head of AI consulting at Every) joins from AI Engineer Summit and plans to use Fable for DSPy book review, code review on a stalled open-source project (Rally), and rewriting a failed memetics book for broader appeal.

The Kate Bench case study is the stream's most substantive segment. Every has been fine-tuning Qwen, GPT-4, and Claude models on 30,000 of Kate's historical copy edits, attempting to replicate her editorial judgment. None have exceeded 50% recall. Fable's analysis suggests the team conflated two distinct problems: (1) repair at a known site (which is learnable cheaply—proven by a $5.29 experiment that worked) and (2) detection and selection (which remains unsolved and may require a different architecture). Fable recommends deploying the model in Kate's live workflow so she can accept/reject suggestions concurrently with her own edits, generating higher-quality training labels than historical forensic recall. Kate (VP of Editorial) expresses concern about reading unfiltered drafts first to preserve her independent editorial judgment, but agrees to a concurrent workflow where the model runs in parallel and results are compared post-edit.

Dan successfully restarts a Fable task after it downgrades to Opus 4.8, suggesting that Lutnik's safety classifier can be bypassed by flipping models mid-conversation. The classifier appears more sensitive to computer-control actions (process termination, browser automation) than to text-only reasoning. No blog post clarifies the new policy beyond 'bigger safety margin.' The stream underscores a broader operator lesson: Anthropic's cadence of releasing models without EAP (4.7, Sonnet 3.5 v2) correlates with lukewarm or negative initial reception, suggesting internal confidence issues or distraction (Fable export-control crisis). However, models often age well—4.7 was panned, then rehabilitated after users discovered new use cases. Fable's token cost disciplines operators: Mike notes he's 'famously against memory,' Dan cancels running agents to free compute, and Ariel (head of ops) demands token-spend audits. The era of 'ripping tokens unaccountably' is over.

Key Takeaways

  • Claim: Fable diagnosed that Every's copy-editing benchmark (Kate Bench) was optimizing for an unvalidated target: 70% recall against Kate's exact historical edits, when Kate herself may only hit ~47–50% self-consistency. | Evidence: Fable: 'It spent 5 weeks, 180 agent threads optimizing a target that was never validated. Strict recall against Kate's exact historical edits with a 70% goal set before anyone measured whether even Kate can hit 70% against herself. The evidence scattered through the repo says she can't. The sealed failures at 47.5 and 50 may actually be near the human ceiling.' | Caveat: Fable's estimate is inferred from scattered repo evidence, not a controlled study. Kate has only self-duplicated a handful of edits, so the sample size is tiny. | Implication: Operator lesson: Benchmarks that target exact human replication may be chasing noise, not signal. Shift to acceptance-based evals (would the human accept this edit?) rather than forensic reconstruction. This applies to any domain where human judgment varies with context, mood, or time—legal review, medical triage, content moderation. | Timestamp: ~8:00–12:00
  • Claim: Fable identified that Every's two successful experiments—span-conditional repair ($5.29 cost) and a live suggestion loop (52/53 useful)—were sidelined in favor of expensive, low-signal optimization runs. | Evidence: Fable: 'The two things that did work—span conditional repair for $5.29 and a live suggestion loop Dan rated 52 out of 53 useful—were sidelined. The single cheapest high-information action, one human spending two hours judging 60 cases, has been sitting idle since June 23rd.' | Caveat: The stream does not explain why these experiments were deprioritized. Possible reasons: organizational inertia, overcommitment to fine-tuning pipelines, or difficulty operationalizing the live loop at scale. | Implication: Operator anti-pattern: Complex, expensive optimization can crowd out simple, effective interventions. Fable's recommendation—use the live loop as a label factory—suggests that production deployment is often the best training data generator, not offline fine-tuning. Ken should ask: Are we running experiments because they're tractable, or because they're effective? | Timestamp: ~10:00
  • Claim: Fable's new Lutnik-era safety classifier silently downgraded Dan's task to Opus 4.8 mid-execution when he attempted to kill a process and automate a browser action. | Evidence: Dan: 'I have just hit our first Lutnik guardrail. I asked Fable to put together a report on all of my token spending… It got triggered on a command that was basically killing a process on my computer and trying to do some work with Chrome. That silently tripped it up about 20 minutes ago. Now it's back to Opus 4.8.' | Caveat: Dan was able to restart the task by manually flipping the model back to Fable, and it continued. Unclear whether this is reliable or context-dependent. The classifier's trigger threshold is opaque. | Implication: Fable's computer-use and browser-automation safety rails are stricter than text-only reasoning. Operators using Fable for agentic workflows (process management, file I/O, web scraping) will hit silent downgrades. Workaround: Flip models manually mid-conversation. This will be a persistent UX pain point until Anthropic clarifies policy. Ken's agent systems should log model downgrades as a reliability signal. | Timestamp: ~35:00–38:00
  • Claim: Anthropic released Claude Sonnet 3.5 v2 without early access, and Every's internal vibe check was universally red except for one editor (Jack Chang) who used it in short collaborative loops, not long agentic runs. | Evidence: Kate: 'I've never seen a response like this where it was just not just meh of like, oh, it's not better than this. It was like, it's actually just worse and it's not good at this and it's not good at this. Like, who is this for?' Dan: 'The last model that we all didn't really like from Anthropic was 4.7, and that was the last model that they dropped without an EAP, which they did for this one too.' Jack's exception: He uses Sonnet in 'much shorter chunks of work… working collaboratively with it,' not long-running tasks. | Caveat: Models often age well. Every's coverage of 4.7 was lukewarm initially, then rehabilitated after users discovered new use cases. Sonnet 3.5 v2 may follow the same trajectory. No EAP means no pre-tuned workflows or best-practice guides. | Implication: Operator heuristic: Anthropic models released without EAP are less reliable for day-one production use. If a model is 'clearly amazing,' Anthropic runs EAP. If it's dropped cold, expect rough edges. Jack's collaborative-loop success suggests Sonnet 3.5 v2 may excel at interactive editing/refinement rather than autonomous long-context tasks. Ken should test in both modes before ruling it out. | Timestamp: ~45:00–50:00
  • Claim: Fable's token cost is forcing operational discipline: operators are disabling memory, killing agents to free compute, and auditing token spend for the first time. | Evidence: Dan: 'I've been hounded. We're sort of a real company now, and so I can't just rip tokens unaccountably anymore.' Mike: 'I'm famously against having memory turned on… the tokens are so expensive. I want to be really judicious about what tokens go into Fable.' Dan kills running agents mid-stream to fix screen-share lag. Ariel (head of ops) demanded a three-month token-spend report. | Caveat: Dan is CEO of a funded startup and was previously comfortable with untracked spend. Smaller operators or bootstrapped teams have always faced this constraint. The shift is cultural, not technical. | Implication: Fable's pricing ($15/M input, $75/M output per unofficial reports) is high enough to force ROI conversations at Every's scale (~$40M Series A per Crunchbase). Operators will shift from 'spray and pray' token usage to targeted, high-confidence prompts. Memory persistence becomes a cost center, not a convenience. Ken's systems should track token spend per task type and auto-disable memory for low-stakes workflows. | Timestamp: Throughout; explicitly ~22:00, ~38:00

Detailed Brief

Kate Bench: Fable's diagnosis of a failed copy-editing benchmark

  • Claims: Every spent 5 weeks and 180 agent threads optimizing recall against Kate's exact historical edits without validating that Kate herself is consistent.; Failures at 47.5% and 50% recall may be near the human ceiling, not model failures.; Repair at a known site is learnable cheaply ($5.29 experiment succeeded), but detection and selection remain unsolved.; The forward path is to split the objective: shift from 'recall versus history' to 'acceptance precision in the live loop,' using production deployment as a label factory.
  • Evidence: 30,000 historical copy edits from Kate across ~4,000 documents over four years.; Fine-tuning attempts on Qwen, GPT-4, and Claude models; none exceeded 50% recall.; Span-conditional repair cost $5.29 and worked; live suggestion loop rated 52/53 useful by Dan.; Two hours of human judging on 60 cases has been idle since June 23rd.; Fable's recommendation: Deploy model in Kate's live workflow; she accepts/rejects suggestions concurrently; results are compared post-edit.
  • Caveats: Kate has only self-duplicated edits a handful of times; Fable's consistency estimate is inferred, not measured.; Kate expressed concern about reading drafts with pre-existing AI edits ('I want to read it my way… I want an unfiltered reaction'), but agreed to a concurrent workflow as a compromise.; Unclear whether the live loop scales beyond Kate—Every has multiple editors, and training data generation may bottleneck on Kate's throughput.; Fable did not specify how to solve detection/selection; it only confirmed it's the hard part.
  • Implications: Benchmark validity: If the human ceiling for self-consistency is 50%, targeting 70% is chasing noise. Operators should validate human baselines before optimizing.; Deployment as training: The live loop (production use → accept/reject → retrain) is a continuous learning system. This is more capital-efficient than offline fine-tuning if human-in-the-loop bandwidth exists.; Detection vs. repair: Two-stage architectures (detect → repair) may outperform end-to-end models. Fable's $5.29 repair success suggests span-conditional edits are tractable; detection is the bottleneck.; Editor workflows: Kate's 'unfiltered first read' preference suggests AI-assisted editing must preserve editorial agency. Concurrent workflows (AI runs in parallel, not inline) may be the only acceptable UX for high-judgment roles.

Fable's Lutnik-era safety classifiers and operator workarounds

  • Claims: Fable's new classifier silently downgrades tasks to Opus 4.8 when triggered by computer-control actions (process kill, browser automation).; Dan successfully restarted the task by manually flipping the model back to Fable mid-conversation.; Anthropic's blog post mentions 'a bigger safety margin' but provides no specifics on what triggers downgrades.; Memory persistence exacerbates classifier sensitivity—Mike notes biology content in memory triggered guardrails for one user.
  • Evidence: Dan's token-spend audit task triggered on 'killing a process and doing something in my browser'; silently flipped to Opus 4.8 after ~20 minutes.; Dan flipped model back to Fable, and it resumed on the same thread.; Mike: 'If they had something in memory that would trip it quite often as well. There's a case of this guy on Twitter saying that because he is a biologist, it was remembering biology stuff and then tripping some of those safeguards.'; Dan and Mike both disable memory to control context and cost.
  • Caveats: Unclear whether manual model-flipping is reliable across all contexts or just low-risk tasks like token audits.; No clarity on whether the classifier is stricter for computer-use actions vs. text-only reasoning.; Anthropic's policy is opaque; no public documentation on trigger criteria or downgrade logic.
  • Implications: Agentic workflows that require file I/O, process management, or browser automation are at higher risk of silent downgrades. Operators should log model switches as a reliability metric.; Memory-off is now a cost and safety optimization. Ken's systems should default to memory-off for low-stakes tasks and memory-on only for high-context, trusted workflows.; Workaround: If downgraded, flip models manually mid-conversation. This is a UX anti-pattern but appears functional as of July 1, 2025.; Lutnik's 'bigger safety margin' likely means lower risk tolerance than pre-export-control Fable. Operators should expect more false positives.

Sonnet 3.5 v2: Universally panned except in short collaborative loops

  • Claims: Every's internal vibe check on Sonnet 3.5 v2 was red across the board except for Jack Chang, who uses it in short, interactive editing sessions.; The model was released without early access (EAP), which correlates with rough edges and negative first impressions.; Jack's success: 'He is less likely to be giving Sonnet a very long-running task to just set it off and do something. He's working much more collaboratively with it, and he found that in working collaboratively with much shorter chunks of work that it actually was really good.'; Anthropic's previous cold-drop model, Opus 4.7, was also panned initially but rehabilitated after users discovered new use cases.
  • Evidence: Kate: 'I've never seen a response like this where it was just not just meh… it was like, it's actually just worse.' Uniform red rankings except Jack's green.; Dan: 'The last model that we all didn't really like from Anthropic was 4.7, and that was the last model that they dropped without an EAP, which they did for this one too.'; Jack rebuilt Codex on his own and is 'really deep into this stuff,' so his eval carries weight.; Dan notes models 'often age well' as users find new workflows.
  • Caveats: No EAP means no pre-tuned prompts, best practices, or known strengths/weaknesses. Users are discovering the model's fit in real time.; Jack's collaborative loop may not generalize to other operators or domains.; Anthropic may have prioritized Fable's return over Sonnet 3.5 v2's polish, so rough edges could be timing artifacts, not capability limits.
  • Implications: Model selection heuristic: If Anthropic drops a model without EAP, assume it's less production-ready. If they run EAP, it's 'clearly amazing.'; Task type matters: Sonnet 3.5 v2 may excel at interactive refinement (editing, debugging, iterative prompting) but fail at long-running autonomous tasks (codegen, research, agentic loops).; Patience pays: Models often age well as users discover fit. Ken should re-eval Sonnet 3.5 v2 in 3–4 weeks, especially if Jack's collaborative loop generalizes.; Every's vibe-check process (red/yellow/green/gold rankings + thorough documentation) is a repeatable eval framework Ken can adapt for his own systems.

Token cost discipline and the end of unaccountable usage

  • Claims: Dan is being audited by Ariel (head of ops) for three months of token spend across ChatGPT, Claude, Codex, and Google AI Studio.; Fable's cost is high enough that operators are disabling memory, killing agents mid-task, and scrutinizing ROI for the first time.; Mike: 'I'm famously against having memory turned on… the tokens are so expensive. I want to be really judicious about what tokens go into Fable.'; Dan: 'We're sort of a real company now, and so I can't just rip tokens unaccountably anymore.'
  • Evidence: Dan kicked off a Fable task to audit token spend across all platforms for the last three months, including personal and company workspaces.; Dan killed running agents mid-stream to fix screen-share lag, suggesting compute resource constraints.; Mike disabled memory to control context bloat and avoid triggering classifiers.; Ariel has been 'bothering [Dan] all week' for the token-spend report.
  • Caveats: Every raised a $40M Series A (per Crunchbase), so cost discipline is relative—smaller operators have always faced this constraint.; The shift is cultural (transition from 'move fast' to 'run like a real company'), not necessarily driven by Fable's absolute cost.; No public pricing for Fable, but industry estimates are $15/M input, $75/M output.
  • Implications: Fable's pricing tier will force ROI conversations at mid-market scale. Operators will shift from 'spray and pray' to targeted, high-confidence prompts.; Memory persistence is now a cost center. Ken's systems should track memory overhead per workflow and auto-disable for low-context tasks.; Token-spend audits should be automated: track cost per task type, model, and outcome. Every is building this in-house with Fable; Ken can replicate with structured logging.; The 'rip tokens unaccountably' era (2023–2024) is over. 2025+ operators will treat LLM calls like cloud compute: metered, audited, and optimized.

Notable Concepts & Terms

  • Kate Bench: Every's internal benchmark attempting to replicate Kate Lee's (VP of Editorial) copy-editing judgment by fine-tuning models on 30,000 of her historical edits across 4,000 documents. Failed to exceed 50% recall; Fable diagnosed the target as unvalidated because Kate's self-consistency may only be 47–50%.
  • Span-conditional repair: A $5.29 experiment in Kate Bench where the model was given a known edit site and asked to repair it. It succeeded, proving that repair at a known site is learnable cheaply. Detection and selection (identifying which spans need editing) remain unsolved.
  • Live suggestion loop / label factory: Fable's recommended architecture for Kate Bench: Deploy the model in Kate's production workflow; she accepts/rejects suggestions in real time; accepted edits become high-quality training labels. Rated 52/53 useful by Dan but was sidelined in favor of offline fine-tuning.
  • Lutnik guardrails / classifier: New safety mechanisms in Claude Opus 5 (Fable) post-export-control reinstatement. Silently downgrades tasks to Opus 4.8 when triggered (e.g., process kill, browser automation, biology content in memory). Anthropic's blog mentions 'bigger safety margin' but provides no specifics. Workaround: Flip models manually mid-conversation.
  • EAP (Early Access Program): Anthropic's private beta for new models. Models released with EAP (e.g., Opus 5 initially) tend to be 'clearly amazing' and production-ready. Models dropped without EAP (e.g., Opus 4.7, Sonnet 3.5 v2) correlate with rough edges and lukewarm reception. Dan uses this as a heuristic for model readiness.
  • Memory-off / memory persistence cost: Operators (Dan, Mike) are disabling Claude's memory feature to control token cost and avoid triggering classifiers (e.g., biology content in memory tripped guardrails). Memory is now a liability, not a convenience, at Fable's price point.
  • DSPy (Demonstrate-Search-Predict): A prompt-optimization framework that Mike Taylor is writing a second book about. He plans to use Fable to review the manuscript for missing concepts and to test whether Kate Bench's architecture is 'the right shape' for DSPy or fine-tuning.
  • Rally (focus group AI project): Mike's side project (with co-founder Uto) to automate focus groups. Mike is afraid to commit code because Uto (a better engineer) set up complex infra. Plans to use Fable for code review to increase confidence before submitting PRs.
  • Concurrent workflow (for AI-assisted editing): Kate's compromise for Kate Bench: The model runs in parallel while she edits; results are compared post-edit. Preserves her 'unfiltered first read' preference while generating training labels. Alternative to inline AI suggestions, which she finds intrusive.
  • Ripping tokens: Every's informal term for heavy, unthrottled LLM usage. The 'rip tokens unaccountably' era (2023–2024) is over; 2025+ operators treat LLM calls like metered cloud compute.

Operator Notes / Why Ken Should Care

  • Benchmark validity: Validate human baselines before optimizing. If Kate's self-consistency is 50%, targeting 70% is noise-chasing. This applies to any high-judgment domain (legal, medical, editorial).
  • Deployment as training: Live loops (production use → accept/reject → retrain) are more capital-efficient than offline fine-tuning if human-in-the-loop bandwidth exists. Every's $5.29 span-conditional repair succeeded; 180 agent threads of fine-tuning did not.
  • Detection vs. repair: Two-stage architectures may outperform end-to-end. Fable confirmed repair at known sites is learnable cheaply; detection/selection is the bottleneck. Ken's systems should separate these concerns.
  • Fable's classifier: Computer-use and browser-automation tasks are high-risk for silent downgrades. Log model switches as a reliability signal. Workaround: Flip models manually mid-conversation. Memory-off is now a cost and safety optimization.
  • Model selection heuristic: Anthropic models without EAP are less production-ready. If they run EAP, it's 'clearly amazing.' Sonnet 3.5 v2 was panned except in short collaborative loops—task type matters more than raw capability.
  • Token cost discipline: Fable's pricing forces ROI conversations at mid-market scale. Automate token-spend audits: track cost per task type, model, and outcome. Memory persistence is now a cost center; disable by default for low-stakes workflows.
  • Editor workflows: AI-assisted editing must preserve editorial agency. Concurrent workflows (AI runs in parallel, not inline) are the only acceptable UX for high-judgment roles like Kate's. Inline suggestions feel intrusive.
  • Code review with Fable: Mike plans to use Fable to review PRs and increase confidence before committing to production. If Fable approves, blame shifts to the model ('who are we to judge?'). This is a repeatable pattern for less-confident engineers contributing to high-quality codebases.
  • DSPy + Kate Bench: Mike will test whether Kate Bench's architecture is 'the right shape' for DSPy or fine-tuning. Row-level optimization (test each edit in isolation) may outperform corpus-level ranking.
  • Every's vibe-check process: Red/yellow/green/gold rankings + thorough documentation is a repeatable eval framework. Ken can adapt this for his own model evals.

Watch Map

  • 0:00–5:00: Intro: Fable is back; Dan sets off first prompts (copy-edit review, token-spend audit) from jacuzzi in Cabo.
  • 5:00–15:00: Kate Bench deep dive: Fable diagnoses failed benchmark—team optimized for 70% recall without validating Kate's self-consistency (~47–50%). Recommends live suggestion loop as label factory.
  • 15:00–25:00: Kate (VP of Editorial) joins; discusses workflow concerns ('I want an unfiltered first read'); agrees to concurrent model deployment.
  • 25:00–35:00: Every's Fable prompt library plug; intro to Every's subscription (ideas, apps, training).
  • 35:00–42:00: First Lutnik guardrail: Dan's token-spend audit silently downgrades to Opus 4.8 after process-kill command. Workaround: flip models manually.
  • 42:00–52:00: Mike Taylor (head of AI consulting) joins from AI Engineer Summit; discusses hit list (DSPy book review, Rally code review, memetics book rewrite). Explains Kate Bench's detection vs. repair problem.
  • 52:00–58:00: Sonnet 3.5 v2 vibe check: Universally red except Jack Chang (senior editor), who uses it in short collaborative loops. Dan notes models without EAP correlate with rough edges.
  • 58:00–60:00: Token cost discipline: Dan audited by Ariel for three months of spend; operators disable memory to control cost. 'Rip tokens unaccountably' era is over.

Source/Metadata

  • Title: LET'S RIP FABLE TOKENS FROM THE JACUZZI
  • Transcript words: 14565
  • Duration seconds: 3582
  • Timestamp note: Approximate timestamps inferred from speaker transitions and topic flow; video duration is 3,582 seconds (~60 minutes). No chapter markers or precise timestamps in transcript; segment boundaries are editorial estimates.
Full transcript 7894 words · 64 min read
0:05

SPEAKER_04

Hello, everybody.

0:19

SPEAKER_04

Hello. Fable is back. We're here. We're live in Cabo. I'm on vacation, but I could not let the return of Fable go by without a live stream. We are here. We've got hydration. We've got a mojito.

0:48

SPEAKER_04

We'll have some guests here in a little bit, but for now, I want to show you something incredibly special. So check this out. Check this out. We've got Fable in our model selector. Let's fucking go.

1:09

SPEAKER_04

[SPEAKER_01] Let's fucking go. And I'm apparently not sharing. So let's see.

1:17

SPEAKER_01

[SPEAKER_04] Share screen.

1:21

SPEAKER_04

Entire screen. Share. You're not able to see my screen. All right. We've got some screen sharing issues happening here, but what you need to know is Fable is back.

1:34

SPEAKER_04

And what we're going to be doing here for the next hour or so is ripping tokens. So if you're here, tell me what you're ripping tokens on. I need to figure out what I want to build. I've got a long list of stuff because I wasn't making things as diligently for the last couple of weeks because I needed Fable to come back. So here we are. I'm going to see if we can get some Ellevery people on here. And I'm going to see if I can fix my screen. Let's go. PE just fired off a massive three-page prompt for an auto-research ML system as soon as it came back. Please let me know how that goes. Okay.

2:24

SPEAKER_04

So a couple of big things.

2:32

SPEAKER_04

One is it's awesome that it's back. We love it. But the second thing that I think is really important is we're going to have to figure out if Fable Lutnik's version is totally lobotomized. I believe it's available for everyone now. So the question is, is Fable Lutnik's version totally lobotomized or not? Even before this whole thing happened, Fable was pretty hard to use. My computer is way overextended right now. Fable was pretty hard to use, actually. There were a lot of times where it was completely off. Wait, I figured it out. Hold on. [SPEAKER_01] Let's go. [SPEAKER_01] Let's just delete all of this stuff. [SPEAKER_00] This will make my performance much better.

3:21

SPEAKER_04

[SPEAKER_01] Oh. All right. We're getting closer here. Don't save. Trying to figure out how to do this. All right. So my big question here is how lobotomized is it? Are the safety classifiers going to be a problem?

3:40

SPEAKER_01

[SPEAKER_04] Because they were already going off a lot with the previous version of Fable for innocuous requests. [SPEAKER_04] So I'm really interested in whether Fable Lutnik's version is actually usable or not. [SPEAKER_04] And we will see.

3:47

[SPEAKER_04] So let me see if I can get my screen share working.

3:52

SPEAKER_00

[SPEAKER_04] I'm in Dia right now. [SPEAKER_04] And there's something about running agents on this computer that's interesting.

3:59

SPEAKER_00

[SPEAKER_04] Let me see if I can get this thing working.

4:04

SPEAKER_01

[SPEAKER_03] Hmm. [SPEAKER_04] Yes, Sam McAllister.

4:14

SPEAKER_04

Claude Cabo Fott. Actually, that's a great model name. You should take that to heart.

4:23

SPEAKER_04

Maybe the next model bigger than Fable is the Claude Cabo.

4:30

SPEAKER_04

All right. I can't share my screen. So I'm just going to narrate what I'm doing. The big thing that I am going to try right now is I've been trying to get a model to be good at copy editing. We do a ton of copy editing at Every. And models are really bad at this. And I have a whole set of benchmarks that measure how good they are because we have a ton of data. We have a lot of data on doing good copy editing. Because at Every, we're constantly editing articles. So I have a ton of that. So I've been trying to put together a benchmark that tests different models on how good they are at copy editing. And then I've been trying to fine-tune on that benchmark.

5:03

SPEAKER_04

And none of the models are very good at it. So what I'm going to do is set Fable off on a fine-tuning task. Or actually, what I'm going to do is have it review all of the work that I've done with Codex and tell me what it thinks. Let's see.

5:18

SPEAKER_03

[SPEAKER_01] Let's see.

5:20

SPEAKER_04

[SPEAKER_01] Let's see. All right. So here we go. What I need is to try Quen. J Unlimited asks, have I ever tried Kimi K2 on this? No. I was fine-tuning Quen on thinking machines. But it did not work very well, unfortunately. Let's see if this works. I'm really having trouble sharing my screen here. [SPEAKER_01] Give me a second. Let's see if I can get someone else from Every to be on this stream with me. All right. I'm going to set off Fable. My first Fable prompt since it came back. So J Unlimited asks, have I ever tried Kimi K2 on this? No. I actually was fine-tuning Quen on with thinking machines. But it did not work very well, unfortunately. Let's see if this works.

6:04

SPEAKER_04

I'm really having trouble sharing my screen here. [SPEAKER_01] Give me... Let's see if I can get someone else from every to be on this stream with me. Let's see. [SPEAKER_04] All right. I'm going to set off Fable. My first Fable prompt since it came back. Ooh, it's got a new effort slider, which is nice. That's nice. I haven't seen this before. Okay. Let's see.

6:40

SPEAKER_04

Okay. So since you were gone, and I know that you don't know. I don't think it knows about export controls. Since you got export controlled, I've been doing a ton of experiments to do better copy editing with the data corpus that we have.

7:00

SPEAKER_01

[SPEAKER_04] I want you to review the experiments because they haven't been working and come to your own conclusion about the quality of the experiments and what we should do next. [SPEAKER_04] What might actually work.

7:11

SPEAKER_04

Do a full review. All right.

7:18

SPEAKER_04

[SPEAKER_04] So we're setting it off on that. Let's see. Heather, I did something similar with Fable before I had it look at my content style guide, Lexicon, and center and create a skill with an eval rubric to run against every time I need a doc creator review. That's really interesting. I'm really curious if you felt that actually worked really well. Because my experience is there's a limit to how good prompt tuning is for these models. And especially if you have a lot of copy rules, it gets lost. But Fable might be different. I honestly didn't really try Fable directly on a set of a really specific style guide because it's so expensive. But maybe I should try that.

7:44

SPEAKER_04

We have one of those.

7:47

SPEAKER_01

[SPEAKER_04] Let's see.

7:51

SPEAKER_04

All right. It's off.

8:09

SPEAKER_04

It's off. It has not been tripped back to Opus 4.8. So that's great. All right. What are my next things that I want to be doing? Let's see. If you guys have Fable use cases that you're kicking off, let me know. I want to talk about them. Okay. Another big thing. Oh. Kate's on, I think. [SPEAKER_02] Hi. [SPEAKER_02] Hi. How are you doing? [SPEAKER_02] Good. [SPEAKER_02] It looks like you're doing better than me. I'm doing pretty good. I've got a mojito. I've got a virgin pina colada. It's 90 degrees out here.

9:07

SPEAKER_04

I'm in the jacuzzi. It's pretty great. How are you doing? [SPEAKER_02] Incredible. [SPEAKER_02] I'm good. [SPEAKER_02] It's also at least 90 degrees here in Manhattan. [SPEAKER_02] I heard that. [SPEAKER_02] I heard that. [SPEAKER_02] But I'm not in a jacuzzi, nor do I have a drink, virgin or otherwise. [SPEAKER_02] But we're good. [SPEAKER_01] We got to keep ripping back.

9:57

SPEAKER_04

[SPEAKER_02] I know. [SPEAKER_02] It is closing on 4 o'clock here in the east, so closing time calls for that. [SPEAKER_02] But we're good. [SPEAKER_02] It's a big day. Do you have any Fable token-ripping use cases that you're excited about, or are you just along for the ride? [SPEAKER_02] I'm just along for the ride right now. [SPEAKER_02] But I'm eagerly watching what everyone else is doing and seeing how it really felt there was a collective depression when it was taken away. Honestly, yeah.

10:17

SPEAKER_04

[SPEAKER_02] I'm so very eager to see now how people are doing. I felt a little bit those old Geico commercials where it's the caveman going around?

10:31

SPEAKER_04

I felt a little bit a caveman. I was just I'm in the Stone Age here using very simple tools when I could be in the Space Age. And it's good to be back. I'll say that.

10:37

SPEAKER_02

[SPEAKER_04] It's good to be back. Incredible.

10:38

SPEAKER_04

Incredible.

10:40

SPEAKER_02

[SPEAKER_04] Okay. [SPEAKER_04] Here's my next Fable prompt.

10:43

SPEAKER_04

So I've been having trouble. For some reason, I can't screen share. This is I really need to solve this.

10:53

SPEAKER_04

Something's going wrong with my computer where for weeks and weeks and weeks, I haven't been able to screen share. And I think it has something to do with the agents that I'm running. I don't know why. But the thing, so one of the big questions that we've had over the last several months is token spend, which is a really big deal now in AI, especially now that Fable's back.

10:59

SPEAKER_02

[SPEAKER_04] How much are we spending? [SPEAKER_04] So Ariel, who is our head of operations and it also happens to be my sister, has been bothering me all week to create a report on all my token spend for the last three months so that we can figure out how much money I'm costing us. [SPEAKER_04] And I've been trying to do it in Codex and it's been a little finicky. [SPEAKER_04] So I'm going to just kick that off right now. [SPEAKER_04] I want you to prepare a report on all of my personal token spend between ChatGPT, Codex, Cloudx, in the API, in the actual apps, in my personal workspace.

11:05

SPEAKER_02

[SPEAKER_04] And if to the extent we can get it in the company workspaces, for me specifically, also Google, I think I have some Google AI studio spend. [SPEAKER_04] So I want you to collect all of it into a spreadsheet.

11:13

SPEAKER_01

[SPEAKER_04] And then I want you to also make a website to display it for us.

11:15

SPEAKER_02

[SPEAKER_04] It should be all the token spend for the last three months starting today and three months backwards. [SPEAKER_04] Thank you. [SPEAKER_04] All right. [SPEAKER_04] Let's see.

11:27

SPEAKER_04

[SPEAKER_02] This is a very responsible initial way back into using Fable. [SPEAKER_02] I have to say.

11:34

SPEAKER_02

[SPEAKER_04] I know, right? [SPEAKER_04] And for Ariel. I'm very impressed. [SPEAKER_04] Also Google, I think I have some Google AI studio spend. [SPEAKER_04] So I want you to collect all of it into a spreadsheet.

11:54

SPEAKER_04

And then I want you to also make a website to display it for us. It should be all the token spend for the last three months starting today and three months backwards. Thank you. All right. Let's see. [SPEAKER_02] This is a very responsible initial way back into using Fable. [SPEAKER_02] I have to say. [SPEAKER_02] I know, right?

12:15

SPEAKER_02

And for Ariel.

12:16

SPEAKER_04

[SPEAKER_02] I'm very impressed. Well, you'll be glad to know that I set it off on evaluating all the copy edit experiments that we've been doing because obviously they have not been working or not been working as well as we would like. And I've just been waiting for Fable to come and save me. Who do we have in the background there? And is he excited about Fable? [SPEAKER_02] Nicholas? [SPEAKER_02] Oh, he's playing a game. [SPEAKER_03] Hi. [SPEAKER_03] Say hi. [SPEAKER_02] Yes. [SPEAKER_02] He will soon be better at using it than all of us given gen alpha. I certainly hope so. I certainly hope so. Oh, wow.

12:49

SPEAKER_04

I'm just checking my X and there is a lot of activity happening here. [SPEAKER_02] So I have a question for you. [SPEAKER_02] Sorry. [SPEAKER_02] There may be a little noise in my background. [SPEAKER_02] So it's included until July 7th. [SPEAKER_02] And then what? Yeah. And then you better be rich. [SPEAKER_02] And then you better pay out of pocket. [SPEAKER_02] That's it. I guess so. Is there a blog? I haven't looked at the blog post yet. Have you seen the blog post? [SPEAKER_02] Let me check. Let me check the channel.

14:00

SPEAKER_02

[SPEAKER_04] And I'm going to try, I've just been trying to get Anthropic to send someone for this. Yeah. Let me see. [SPEAKER_01] Let's see what we can get. [SPEAKER_04] I'm really curious for the inside scoop on how they're thinking, feeling about Fable what?

14:10

SPEAKER_04

[SPEAKER_02] I'm sorry. I didn't hear that. [SPEAKER_02] Name. [SPEAKER_02] We named variant after him but I think it's true. All right let's see let's look at their. I really love their new visual design language for their models that uses all these butterflies and nature stuff. I don't know what it's called when you, I don't know what it's called but it's really nice. It's beautiful. Very beautiful. Okay, Table Five is rolling out.

14:28

SPEAKER_04

It brings fifth generation intelligence to your most ambitious coding and professional work. I think that's the first time I've heard them say fifth generation intelligence. How is it possible? How are we already there? I don't know. Okay, wait. Okay, on Friday I'm going to read this post a little bit so we can talk about it on Friday. The U.S. government applied export controls to our newest Claude Five and Claude Mythos models. This required us to restrict access to foreign nationals whether inside or outside the U.S. Because the order affected immediately and we had no reliable way to verify nationality in real time, we suspended access to both models for all users. As of today, the export controls have been lifted.

14:31

SPEAKER_02

[SPEAKER_04] It will be available starting tomorrow, Wednesday, July 1st. Oh, this is, I think this is the same, sorry, this is the old announcement. Let's see. [SPEAKER_04] Yeah, this is the same announcement.

14:36

SPEAKER_03

[SPEAKER_02] So they're not really saying anything new. [SPEAKER_04] Yeah. Oh, I didn't really read this yesterday. One of the things that they're saying is the classifier has a bigger safety margin of what it appears to be potentially harmful. So again, we're going to have to see. Oh, let's see. I just got some results.

14:37

SPEAKER_02

[SPEAKER_04] Okay, so just to set the stage here, I just ran Fable Lutnik's version on our copy editing benchmark, which basically it's an attempt to get language models to be as good of copy editors as Kate is, and that has been perfectly extremely difficult. And basically I use a corpus of 30,000 of her historical copy edits and then I've tried a bunch of different versions to see if we can get these models to, given that corpus, perform as well as she does on a copy editing set of copy editing tasks. And then I just set off Fable the program to go and review all of the copy edit experiments that we ran, and Fable said:

14:40

SPEAKER_02

[SPEAKER_04] "The program's engineering and honesty are genuinely excellent. The corpus, the benchmark discipline, and the negative result hygiene are rare and valuable." So it's actually complimenting Codex a little bit, which is interesting to see. But it spent five weeks, 180 agent threads optimizing a target that was never validated. Strict recall against Kate's exact historical edits with a 70 percent goal set before anyone measured whether even Kate can hit 70 percent against herself. The evidence scattered through the repo says she can't. So the sealed failures at 47.5 and 50 may actually be near the human ceiling. Meanwhile, the two things that did work—banned condition repair for five dollars 29 cents and a live suggestion loop Dan rated 52 out of 53 useful—were sidelined. And the single cheapest high information action, one human spending two hours judging 60 cases, has been sitting idle since June 23rd. The forward path is to split the objective from recall versus history to acceptance precision in the live loop and use that loop as a label factory that I'm always needed. Really interesting.

14:46

SPEAKER_04

So it's saying a bunch of that I don't totally understand, but I think basically what it's saying is one thing that we're trying to see is how good is a model at making historical edits in the way that Kate would make them. What we have not measured is do you make the same edit every single time in the same way that we have to do it, and the answer will change a lot because if you're not, if you yourself are not consistent, then... [SPEAKER_02] Right. What's the expectation? Should the model still be more consistent than me or no?

14:50

SPEAKER_04

I mean, I don't know. I mean, hopefully we should get to a point where it is where it's Kate on her best day. That's what we would hope, and obviously the corpus, because it's four years of your edits, includes Kate on many different kinds of days. Who knows. There's been a lot of bad days in there. [SPEAKER_02] What would you take from this? What would you say would be a next step based on its response?

15:02

SPEAKER_02

[SPEAKER_04] Well, I asked it. So it said it needs you to judge the thing that I would ask you to judge. So maybe we could spend a little bit of time on that. Let me just see. [SPEAKER_04] cake on her best day [SPEAKER_04] that's what we would hope and obviously the corpus because it's four years of your edits includes Kate on many different kinds of days [SPEAKER_04] who knows, there's been a lot of bad days in there, who absolutely knows what would you take from this? What would you like to say would be a next step based on its response?

15:13

SPEAKER_04

well I asked it, so it said it needs, would you like plugging me, it needs you to judge the thing that you were at, that I would ask you to judge, so maybe we could spend a little bit of time on that. Let me just see a baseline. Is there a construction going on? I think we need a baseline from you of whether you would make the same edit. [SPEAKER_01] I can think about this but it won't hurt to get that from you

15:18

SPEAKER_02

[SPEAKER_04] and let's see what else. I think it also wants to be in your live loop of editing things, which shouldn't be too hard. So take a look at the documents that you're doing as you do it and see how close you are to what you would say is good. And I think that what we probably need to do is put it somehow into your loop so when someone sends you a document to copy edit, it goes through and does a pass first and then you go and accept or reject and then add things that you think it is right. Do you think that's doable?

15:20

SPEAKER_02

I think it's doable. I think the question that I have, or that I've had as we've talked about this over the past months and years is that at least more recently, since I have been supported by a team and I'm not doing everything myself, often the first time I'm reading something is when I'm getting it for a top edit. And that's not every piece. There are certainly pieces that I will read drafts and go through with our editors to make sure we're aligned on the direction and on the same page or help troubleshoot and problem solve. But typically it's the first time I'm reading it and so I don't, I want to read it my way. There's just something about it, I want to read it myself before before, and again I don't know how rational that is because I'm not the first person to read it, I am a lot of person to read it, everyone in the steps in our process. But I like to want it before I give it to another set of eyes, to like an AI set of eyes, so to speak. But that may just be my desire. It's not actually necessarily something that couldn't be done in the process.

15:21

SPEAKER_04

So do you want like, does it ruin it for you if when you get to a piece it has edits in it that you can accept or reject as you're reading it?

15:22

SPEAKER_04

[SPEAKER_02] No, because sometimes an editor will leave a few pieces that are like comments that are like I want Kate to chime in on here and so it will be left for me. So that's not totally uncommon. Yeah, I guess I just have to kind of think about it. It's just my own perception of what an unfiltered reaction is from me versus knowing that it's you want an unfiltered reaction, yeah, for myself. But again it's been filtered through my team who I trust very much, which is why they are the ones who are doing it. But I don't know, it's a really good question. At the same time, as you know, again we've talked about like if an AI can catch that missing Oxford comma, that's great, better that it does it than I have to do that. Yeah, it's hard to say. I mean, as you know, I am ultimately responsible for everything that I read and I'm responsible for every single thing we publish and so you know I take that seriously. I have to think about that. Okay, so at least for now, every time you get a draft we could set it off and it can do stuff and you can do your thing and then we compare at the end.

15:23

SPEAKER_04

[SPEAKER_02] That's actually a really good idea to have it running concurrently. Yeah, and then yeah I like that idea. Okay, let me just ask Table what it thinks. Okay, while you're doing that, one thing that Kate wants is she likes to look at a draft without seeing a bunch of copy edits from a model because she wants to get her own take on it first. So is it possible for us to have a workflow where when she gets a draft she sets off the model to do its own version and then she does her version and then we compare and then we improve that way? Is that possible or how would the loop work?

15:25

SPEAKER_02

[SPEAKER_04] You can do it also. You can do a production deploy and make sure it's like has a database. We don't want it to get lost. Yeah okay, we'll see. Oh we've got Mike Taylor. Mike, hello?

15:26

SPEAKER_04

[SPEAKER_03] Hey, hey, just jumped out of the conference to test. All right, please introduce yourself and tell people which conference you are at and tell us your reaction. They will be back.

15:33

SPEAKER_02

[SPEAKER_03] Yeah, I didn't realize you'd actually be in the hot tub. That's pretty funny. We don't miss, we don't mess around. Yeah, nothing gets in the way of people. Yeah, I was at the AI Engineering Conference in San Francisco. I think this is the fourth, fourth annual one. I went to the first one a while ago and it was crazy packed, like dense with every weird nerdy AI engineering thing that I was doing at the time and I learned a lot from that. So I was like I wanted to see how it differed and now it's at Moscow Center and it's huge. You know, pretty much everyone in AI is here, it seems it feels like, apart from the people in the hot tub. But yeah it's been really fun and like as soon as Fable came out, of course it was a stream of people.

15:34

SPEAKER_02

[SPEAKER_04] Did it empty out? What happened?

15:36

SPEAKER_01

[SPEAKER_03] No, I mean there's still yeah it's still a lot of people there. That's so funny. Yeah, but I was like, yeah I need to go test this thing. So yeah I didn't introduce myself, but I'm Mike Taylor. I work as the head of AI tech consulting, so work with Natalia, if you've seen her podcast recently, as she was on. And yeah, mostly doing a lot of AI enablement stuff, so helping people figure out how to get more AI-pilled and be a little bit more like us weird people who test Fable in hot tubs.

15:47

SPEAKER_04

Amazing. Awesome. And what's the first thing you're going to do with Fable?

15:54

SPEAKER_02

[SPEAKER_03] So I'm publishing a book on DSPy at the minute. It's my second book, first was on prompt engineering. And it's all drafted, so it's like ready, like I'm not allowed to use AI to write it. And actually AI wasn't that good at DSPy anyway when I started writing it, so it wasn't a big loss.

15:56

SPEAKER_04

[SPEAKER_03] and yeah mostly doing a lot of AI enablement stuff so helping people figure out how to get more AI pilled and be a little bit more like us weird people who test Fable and hot tubs.

16:11

SPEAKER_02

[SPEAKER_04] Amazing. Awesome. And what's the first thing you're gonna do with Fable? So I'm publishing a book on DSPy at the minute. It's my second book, first was on prompt engineering.

16:12

[SPEAKER_03] And it's all drafted so it's ready. I'm not allowed to use AI to write it. And actually AI wasn't that good at DSPy anyway when I started writing it, so it wasn't a big loss. But yeah, the main thing I want to do is figure out if there's anything I'm missing. I think one thing Fable was good at when we were testing it was figuring out from a lot of context what was meaningful, and I haven't really seen it like make good judgments with previous models. And Fable still, I would still prefer human judgment over Fable, but I would say that it could at least supplement and say like hey Mike, you kind of missed this major thing. So that was one of the first things I wanted to try.

15:57

[SPEAKER_04] Cool. Do you think how much have you taken a look at the Kate Bench stuff and do you think that Fable will help there? [SPEAKER_03] Yeah, that's another good one. So if people don't know, Kate Bench, we're trying to automate Kate. We just went through it. We just didn't call it Kate Bench. [SPEAKER_02] But yeah, yeah, we were calling it Kate. [SPEAKER_03] Kate Bench, yeah, there we go. You're famous. First person to be benchmarked. Yeah, so the stuff we're working on there, generally speaking, the edit corpus that we have is thousands and thousands of edits. But I think it's only from 14 documents. Is that right, Dan?

15:57

[SPEAKER_04] From 14 documents? No, it's from three to four thousand documents today.

15:57

[SPEAKER_03] Oh, okay, interesting. Yeah. So that was one of the things, validated edits or something from 14, but there's a lot of like that. Oh yeah, okay, for the validate. Okay, yeah, cool. So yeah, so one of the problems I was running into is the models just wasn't learning general rules or general patterns. It was basically over-indexing on a few small things and just not really predicting what the right text would be. So I think the main thing that Fable can help with here is actually redesigning the experiment to some degree because the experiment we're running is kind of complex. It has the evals. Basically, they happen all at once across the whole corpus. So when you're testing whether it did a good job of finding the same edits as Kate, one of the ways it tests it, as I understand it, is that it decides which ones Kate, which edits were the most meaningful, and because it has a ranking, that's one of the big things that's tripping it up. And when I've done optimization in the past, I've done it more row level where you can test each individual edit in isolation from all the other edits. So I think it'll be interesting to see what Fable thinks about the architecture of the experiment and whether we can make it the right shape for something like DSPy or fine-tuning. So that's something I want to ask him.

15:57

[SPEAKER_04] That makes sense. So I actually did ask it and it said something pretty interesting. It said the programs engineering and honesty are genuinely excellent, but it spent five weeks and 680 agent threads optimizing a target that has never been validated against Kate's exact historical edits with a 70 percent goal, except no one measured whether even Kate can hit 70 percent against herself. The evidence is scattered through the repo. She can't. The sealed failures at 47.5 and 50 percent may actually be near the human ceiling. Meanwhile, the two things that did work, which is span conditional repair and a live suggestion loop, were sidelined, and the single cheapest high information action, which is once a human spending two hours judging 60 cases, has been sitting idle since June 23rd. So it thinks that repair at a known site is learnable cheaply. So that means if we've identified that a particular site is something that Kate would fix, we'll know what the fix would be. But detection and selection is unsolved. So it's been really hard for us to detect which parts of a paragraph, for example, Kate would notice, and that doesn't seem like something that we can fix easily.

15:57

[SPEAKER_02] I would not be that I would be 50 percent accurate against myself, not 70 percent. [SPEAKER_04] Well, that's where we're getting that from. Do you make the same edit twice? [SPEAKER_02] But I've only done that a handful of times. Is that what it's taking from? [SPEAKER_04] I think it's just saying it doesn't know yet. [SPEAKER_02] Okay, not that I'm not accurate against myself?

15:57

[SPEAKER_04] Because we don't have that. Yeah. So okay. If you just got here, we are ripping Fable tokens from the hot tub in Cabo. We've got a Fable prompt library. So if you want to start ripping Fable tokens like we are, you should go to every.to slash p slash claw dash Fable dash five dash prompt dash library. I promise it's worthwhile. It's got all the prompts that we at Every are using to rip Fable tokens and it also has prompts from people like Mike Krieger who's on my podcast when Fable first came out talking about it. I think it was around the Fable time period. And it's pretty cool. You should check it out.

15:57

[SPEAKER_04] Yeah, if you don't know Every, we're the only subscription you need to stay at the edge of AI. We do ideas, apps, and training. On the ideas side, we have a daily newsletter. Whenever something happens in AI, like today, we publish five checks, we publish our own commentary and our own results using these models to figure out what works, what doesn't work, what they can be used for, and what's good for what use case. We also have apps. We have a suite of apps that we build internally with these models that help us work at the edge of AI, everything from Quora, which is an eight-day big email inbox, to Sparkle, which is a file organizer. We have a bunch more. We've got Monologue, which is a Whisper for speech-to-text mark application, which is all-time. You should check it out. And then we do training. So we do live streams like this one. We do courses. We do consulting with large companies. That's part of what Mark and Mike do. Anyway, if you don't know Every, you should check it out. Normally I'm not in a hot tub. This is a special event. So I don't want to promise more hot tub time. This is more of a one-off.

15:57

[SPEAKER_04] Have apps. We have a suite of apps that we build internally with these models that help us work at the edge of AI. Everything from Quora, which is an eight-day big email inbox, to Sparkle, which is a file organizer. We have a bunch more. We got Monologue, which is a Whisper flow speech-to-text app, and Mark, which is all time. You should check it out. And then we do training. We do live streams like this one. We do courses. We do consulting with large companies. That's part of what Mark and Mike do. Anyway, if you don't know Ever, you should check it out. Normally, I'm not in a hot tub. This is a special thing. I don't want to promise more hot tub time. This is more of a special thing for people. But in general, we have a lot of really cool stuff that we do. Now, where were we? Does anyone remember? I think getting to my head.

15:57

[SPEAKER_03] Because I don't remember where we're talking about. Cape Bench. And how to measure. We actually, one thing I wanted to point out here is that because I did a bunch of work with AI personas last year for my project, I was working on. And we found that we basically could never reconstruct what a certain person actually did. The accuracy was always pretty low, about 60 percent. Where we got much better results was in predicting something that the person could plausibly do. As in, something that that person would agree with. Yeah, this is something I might do. Because the world, you know, if you edit the same thing twice, there's going to be a different result in some cases depending on what mood you're in, if you had lunch today. There's a million studies like this out. Say with judges. One of my favorite studies is they found that judges are much stricter if they hadn't eaten lunch yet. Things like that. There's all sorts of variables that you're trying to capture, and they're not necessarily important variables to capture. So I would say that the important thing is, do you think Kate would agree with the edit? Once Fable has made it. I think that's maybe what we can design a study around, or maybe Fable can help us organize that.

15:57

[SPEAKER_04] I agree. Yeah, I think that the strict exact thing is probably worse than "Does Kate agree if this is a good idea or not?" I have just hit our first Lutnik guardrail. I asked Fable to put together a report on all of my token spending because I've been hounded. We're sort of a real company now, and so I can't just rip tokens unaccountably anymore. So I've been asked for a report of all my token spend, and I asked Fable to do it. It got triggered on a command that was basically killing a process on my computer and trying to do some work with Chrome. That silently tripped it up about 20 minutes ago. Now it's back to Opus 4.8. What I'm really curious about actually is if I flip the model back if it'll keep going and be okay. So I don't know if this would have tripped previously, but it's a very innocuous command.

15:57

[SPEAKER_02] What do you think that's saying in your mind about the kinds of guardrails it's now putting up? Sorry about the noise in my background.

15:57

[SPEAKER_04] It looks like the thing that I got tripped up on is ending a process and doing something in my browser. So there's something about maybe controlling your computer that makes it a little bit more sensitive, I would guess. But to be honest, I don't know. We're going to have to do some more thinking and talking. Luke Face says, "Lol, I flipped past this because I only read the title and it looked like an indie hacker bro building in public slot fest push-ups and tech on BS." But then he said, "But you dudes in hot tubs is awesome," which I thoroughly agree with. And he was requesting techno, but then said lawnmowers the same thing. So thank you, Luke Face. And I agree. Rich in Paradise, who is probably Howard Lutnik's handle, says, "Only hackers and processes." So that seems like something Howard Lutnik would say. Howard, if you would like to come on the show and talk to us about these new classifications, we're very interested in talking about it. It does seem that I asked Fable to continue after it got flipped back to Opus. I asked it to continue, and it does seem that it is back on the same thread. So if you're trying to use this and it gets flipped to Opus, you can flip it back, I would guess, depending on the conversation. In this conversation, I think it's a pretty low-risk conversation. I'm just asking it to tally up my token spend, and there was one command that flipped it. Then I just restarted, and it seems to work. So that's interesting.

15:57

[SPEAKER_03] One thing people saw from the last time we had Fable was if they had something in memory that would trip it quite often as well. So there's a case of this guy on Twitter saying that because he is a biologist, it was remembering biology stuff and then tripping some of those safeguards. So yeah, I'm famously against having memory turned on. I wonder if that's going to be more normal for people to turn it off, yeah, to turn it off just to control the context. Also, the tokens are so expensive. I don't know. I want to be really judicious about what tokens go into Fable, you know, when I'm paying for usage.

15:57

[SPEAKER_04] All right. What else, Mike? What else is on your mind? We haven't talked about it. [SPEAKER_03] Bringing up my hit list. I made a little note. Actually, it's funny. Me and Kieran were at the conference yesterday together. He gave a talk on compound engineering. And when they said they were going to release Fable, he bolted up and was like, "Let's go." And I was like, "Wait, hold on, hold on. It's not out yet." And then he showed me his hit list, and I showed him mine. We both had, you know, here's what we're going to test. What was on there? [SPEAKER_04] I think he may be traveling. [SPEAKER_03] Yeah, he's traveling back.

15:57

[SPEAKER_04] Is he on a plane or something or in a car? [SPEAKER_03] I'm not sure. [SPEAKER_04] Yeah, yeah. He'd be sad to be missing this. I know, right? Like, we need him here. This is like his time to shine. So what's on your list? And did you catch what's on his hit list? [SPEAKER_03] Yeah, yeah. So on my list, one of the other things I have, I wrote a really terrible book. Well, I say terrible. I think it's pretty good, but it didn't do very well. My first book, before the DS, before the prompt engineering one, it was self-published. Only 200 people ever bought it and read it. And it was on memetics.

15:57

[SPEAKER_03] I'm not sure. Yeah, yeah, he would be sad to be missing this. I know, right? We need him here. [SPEAKER_04] This is, yeah, yeah it's his time to shine. So what was, what's on your list and did you catch which is what's on his hit list?

15:57

[SPEAKER_03] Yeah, yeah. So on my list, one of the other things I have, I wrote a really terrible book. Well, I say terrible, I think it's pretty good, but it didn't do very well. My first book, before the DS, before the prompt engineering one, it was self-published and only 200 people ever bought it and read it. And it was on memetics, it's how things go viral. I used to work in marketing so it was top of mind. And I remember someone gave me some feedback and he was like, "Bro, you just need to watch all these influencers." He was telling me all the Twitter gurus of how they try and make everything go viral. "You need to watch this, this, and this person," and then just take all the core material and make it more digestible for people. And I'm like, yeah, I'd love to do that, but I just spent a long time writing it. I didn't want to ever look at it again. So now I feel like enough time has passed that I just wanted to give that to people and be like, "What can you do with it? Do you make this actually more available to the general public, right? More interesting to the average person?" So that's another thing I had on my list to let rip.

15:57

And another one is I've been working on Rally. This is the AI it's known as the focus group project. I was working on it more full-time last year and now it's just a side project because I have a job. But I have been afraid to commit stuff to that repo because my co-founder Uto, he also works full-time. He is a much better engineer than me, and he set up a bunch of stuff that good engineers set up, but that's made it less intelligible for me. So that's the other thing I wanted to do is take some of the ideas that I had and run it through code review with Fable. I have a bunch of PRs sitting there that he hasn't had time to review, and so I think that's another good piece to increase my confidence that I'm not messing something up with his good setup.

15:57

[SPEAKER_04] Yeah, I mean, if you can just blame it on Fable, like what does he, yeah, exactly. Yeah, well, yeah, I mean if Fable says it was good, then who are we to judge? [SPEAKER_03] Yeah, yeah, I do actually feel that way.

15:57

[SPEAKER_04] Like I just remember when it first came out, there was this palpable change in momentum where Kieran was on this tear where he was on track to spend like 1.5 million dollars in Fable meetings, but his product was getting better second by second. It was crazy, you know? And for me looking at it, I was like that definitely gives me a lot more confidence that I can and should be committing production code to some of these apps we're working on. Which is now giving me ideas like maybe I should try to ship something to production this week while I'm on vacation. [SPEAKER_04] I think we just lost Mike. His hotel internet.

15:57

[SPEAKER_02] What I think is interesting, and I know obviously Anthropic was not in control of this timing around Fable and ideally would not have had it taken offline to then be restored online today, is we're actually getting ready to publish a vibe check with Sonnet 5. How's that going? I'm waiting for the copy to land with me. It's quite extensive. It's going to be a big, thoughtful, typically thorough vibe check. [SPEAKER_04] How do people like it? I've been on vacation so not, it's not that good right, the vibes are off.

15:57

[SPEAKER_02] I've never seen a response like this where it was just not just meh of like, "Oh, it's not better than this." It was like, "It's actually just worse and it's not good at this and it's not good at this. Like, who is this for?" It was actually a genuine question of who is this for? [SPEAKER_04] Interesting. That was the response to Opus 4.8, I think, initially, if you remember. Or maybe it was 4.7. It was one of the, not yeah, this one is pretty, I kind of feel like definitively. I've never seen so uniformly, you know, within our reach test rankings of red, yellow, green, or gold, it's red uniform except for Jack.

15:57

[SPEAKER_02] Except for Jack, it's a green because I had a whole conversation with him about it this afternoon actually. Our senior editor Jack Chang, who's a writer as well as a builder and product guy, he essentially rebuilt Codex on his own, so he's really deep into this stuff. And he said it was green because he is less likely to be giving Sonnet or another model a very long-running task to just set it off and do something the way that Kieran would be like, "Go and do this and I'll come back in two hours and see if it's done." He's working much more collaboratively with it, and he found that in working collaboratively with much shorter chunks of work that it actually was really good. But he said he's the only one who likes it.

15:57

[SPEAKER_04] It's really interesting. It's also interesting because, well, one, the last model that we all didn't really like from Anthropic was I think 4.7, and that was the last model that they dropped without an EAP, which they did for this one too. So we didn't have early access. So I don't know why. Maybe they were distracted with Fable, but also maybe sometimes for some of these models where they're not clearly amazing, they're just dropping it because they're less confident it's great. I'm just going back to look at our coverage for a moment. But also, people often don't like a model and then decide to like it later because models, just because you don't like it initially, what that usually means is the way that you're trying to use it is not good or not better than a previous model, right? But then often what happens is people find new ways of using it with new powers that they didn't realize were there three weeks later or four weeks later or whatever, which we try to solve for by getting access to it early. But we don't have early access to this model, so this is a very, we're just gonna have to give you what we have when we have it. But I'd be really curious to see how it ages in the next five weeks, especially if you're using Fable as an orchestrator for it. If you're not using Sonnet directly, but you're just having Fable use it, right?

15:57

[SPEAKER_02] Yeah, I'm just going back and checking our coverage of 4.7, and we were pretty lukewarm when it first came out and then about a week later published our vibe shift that we've actually had a little bit of a change of heart.

15:57

[SPEAKER_04] Three weeks later or four weeks later or whatever, which we try to solve for by getting access to it early, but we don't have early access to this model. So this is a very—we're just going to have to give you what we have when we have it. But I'd be really curious to see how it ages in the next five weeks, especially if you're using Fable as an orchestrator for it. If you're not using Sonnet directly, but you're just having Fable use it, right?

15:57

[SPEAKER_02] Yeah, I'm just going back and checking our coverage of 4.7, and we were pretty lukewarm when it first came out. And then about a week later, published our vibe shift that we've actually had a little bit of a change of heart on 4.7. The more people are using it and using it for different things. So it's not uncommon, but it was just that if you were in our Slack, it was just a very consistent stream of "why, why?" [SPEAKER_04] Yeah. Um, Kate, I am about to pass away from dehydration. [SPEAKER_02] We don't want that to happen, especially not on camera.

15:57

[SPEAKER_04] I'm sitting in a hot tub and drinking and it's fine. It's not even—it's like one, almost 1:30. So this was fantastic. Thank you for joining. I'm going to be ripping Fable tokens for the rest of the day. If you want to see more, follow me and Ev on X. But more importantly, check out every.to. It's the only subscription you need to stay at the edge of AI. We are going to be publishing a list of Fable prompts, which I think we already have out. Let me just bring that up. We have a Fable prompt library. So if you want to use Fable to the fullest extent, go to every.to/p/cloud-fable-5-prompt-library. Subscribe to Every. We've got ideas, apps, and training. And enjoy the return of Fable on this July 4th weekend. And we'll see you next time.

15:57

[SPEAKER_00] Enjoy Cabo. Thank you. Bye.

16:12

SPEAKER_02

we named variant after him but i think it's true

16:19

SPEAKER_04

all right let's see let's look at their i really love their new uh they haven't they have a new like visual design language for their models that uses all these like butterflies and nature stuff i don't know what it's called when you like i don't know what it's called but it's it's really nice it's beautiful very beautiful okay table five is rolling out it brings fifth generation intelligence to your most ambitious coding and professional work i think that's the first time i've heard them say fifth generation intelligence how is it possible how are we already there uh i don't know uh this okay wait okay on friday i'm gonna read this post

17:02

SPEAKER_04

a little bit so we can talk about it on friday the u.s government applied export controls to our newest cloud paper five and cloud mythos models this required us to restrict access to foreign nationals whether inside or outside the us because the order to affect affect immediately and we had no reliable way to verify national nationality in real time we suspended access to both models for all users as of today the export controls have been lifted it will be available starting tomorrow wednesday july 1st oh this is i think this is the same sorry this is the old announcement let's see yeah this is the same announcement

17:46

SPEAKER_02

so they're not really saying anything new

17:49

SPEAKER_04

yeah oh i didn't really read this yesterday one of the things that they're saying is like they the classifier has a bigger safety margin of what it appears to be potentially harmful so again we're gonna have to see oh let's see uh i just got some results okay so just uh just to set the stage here i just ran fable lutnik's version on our copy editing benchmark which uh basically it's an attempt to get language models to be as good of copy editors as kate is and that has been perfectly extremely extremely difficult and uh basically i use a corpus of 30 000 of her historical copy edits and then i've tried a

18:49

SPEAKER_04

bunch of different versions to see if we can get these models to uh given that corpus perform as well as she does on uh on a cop on a copy editing a set of copy editing tasks and then i just set off fable the program to uh go and review all of the copy edit experiments that we ran and fable said the program's engineering and honesty are genuinely excellent excellent the corpus the benchmark discipline and the negative result hygiene are rare and valuable um so it's actually complementing codex a little bit which is interesting to see but it spent five weeks 180 agent threads optimizing

19:30

SPEAKER_04

a target that was never validated strict recall against kate's exact historical edits with a 70 goal set before anyone measured whether even kate can hit 70 against herself the evidence scattered through the repo says she can't so the sealed failures at 47.5 and 50 may actually be near the human ceiling meanwhile the two things that did work banned condition repair for five dollars 29 cents and a live suggestion loop uh dan rated 52 out of 53 useful useful were sidelined and the single cheapest high information action one human spending two hours judging 60 cases has been sitting idle since

20:06

SPEAKER_04

june 23rd the forward path is to split the objective from recall versus history to acceptance precision in the live loop and use that loop as a label factory that i'm always needed really interesting so a it's like saying a bunch of that i don't totally understand but um the i think basically what it's saying is one thing that we're trying to see is how good is a model at making historical edits in the way that uk would make them what we have not um measured is do you make the same edit every single time in the same way that we have to do it and the answer will like change a lot job because if you're not if you yourself are not consistent then

21:00

SPEAKER_02

right what's the expectation should the model should the model still be more consistent than me or no

21:05

SPEAKER_04

i mean i don't i don't know i mean hopefully we we should get to a point where it is where it's uh cake on her best day you know that's what we would that's what we would hope and uh obviously the corpus because it's you know four years of your edits includes kate on many different kinds of days who knows who knows there's been a lot of bad days in there who absolutely knows um

21:36

SPEAKER_02

what would you take from this what would you like say would be a next step based on its

21:41

SPEAKER_04

response well i asked it so it said it it needs yeah would you like plugging me it needs um it needs you to judge the uh the thing that you were at that i would ask you to judge so maybe we could spend a little bit of time on that let me just see um a baseline is there a construction going on i think we need a baseline from you of would you make the same edit

22:31

SPEAKER_01

i'm not talking like i can think about this but it won't hurt to get that from you

22:39

SPEAKER_04

um and uh let's see what else i think it also wants to be in your live loop of editing things which shouldn't be too hard so like okay take a look at the documents that you're doing as you do it and see how how close you are to what you would say is good um and basically i think that what we probably need to do is put it some how how this how should this work so i think we probably need to do do it in a way that is that puts it into your loop so when someone sends you a a document copy edit it it goes through and does a pass first and then you go and accept or reject or and then add things that you think it

23:43

SPEAKER_04

is right do you think that's doable like what i think it's doable i think the question that i have or

23:54

SPEAKER_02

that that and that i've had as we've talked about this ad nauseum over the past months and years is um is that at least at least more recently now that um you know since i have had uh i'm supported by a team and i'm not i'm not doing everything myself often the first time i'm reading something is when i'm getting it for a top edit um and that's you know and that's not every piece there are certainly pieces that i will read drafts and go through with our editors to make sure we're aligned on the direction and on the same page or help troubleshoot problem solve but um typically it's the first time i'm reading it and so i don't i want to

24:34

SPEAKER_02

read it my like there's just something about like i want to read it myself yeah you know um before before and again i don't know how rational that is because i'm not the first person to read it i am a lot of person to read it everyone in the steps in our process but there i like want it before i give it to another set of of to like an ai set of eyes so to speak so but that may be that's just my desire it's not actually necessarily um like it's not something that like couldn't be done in the process

25:12

SPEAKER_04

so well do you want like does it ruin it for you if when you get to a piece it has edits in it that you can accept or reject as

25:25

SPEAKER_02

you're reading it no um because because sometimes an editor will leave like a few pieces that are like a few comments that are like i want kate to chime in on here and so it will be left for me um so that's not totally uncommon um yeah i guess i just have to kind of it's just my own perception of what like an unfiltered reaction is from me um versus knowing that it's you want an unfiltered yeah like for myself but again it's been filtered through my team who i trust very much which is why they are the ones who are doing it um but it's i don't know it's a really good question at the same time as you know again

26:08

SPEAKER_02

we've talked about like if an ai can catch that like missing oxford comma like great better that it does it than i have to do that um yeah it's hard to say i mean as you know like i am ultimately i read everything that i read and i'm responsible for every single thing we publish and so you know i take that

26:28

SPEAKER_04

seriously i have to think about that okay it was like at least for now i guess every time you get a draft we could set it off and it can do stuff and you can

26:49

SPEAKER_02

do your thing and then we compare at the end that's actually a really that's actually a good idea to have it running concurrently yeah and then um yeah i like that idea actually okay let me just ask table

27:04

SPEAKER_04

what it thinks okay while you're doing that one thing that kate wants is she likes to look at a draft um without seeing a bunch of copy edits from a model because she wants to get her own take on it first so is it possible for us to have a workflow where when she gets a draft she sets off the model to like do its own version and then she does her version and then we compare and then we improve that way is that possible or how would the loop work you can do it also you can do a production deploy and make sure it's like has a database like we don't

28:08

SPEAKER_04

want it to get lost um yeah okay we'll see oh we've got mike taylor mike hello hey hey just uh just jumped out of the conference to test all right all right please introduce yourself and tell people which conference you are at and tell us your reaction to uh they will be back yeah so uh first of all

28:35

SPEAKER_03

yeah i didn't realize you'd actually be in the hot tub that's pretty funny we don't miss we don't mess around yeah yeah nothing nothing gets in the way of people um yeah i uh yeah i was at the ai engineering conference in san francisco uh i think this is the fourth fourth one fourth annual one um i i went to the first one uh a while ago uh a while ago and uh it was crazy packed like dense with every um like weird nerdy ai engineering thing that i was doing at the time and uh like i learned a lot from that so i was like i wanted to see how it how it differed and now it's at moscow and center it's huge it's you know

29:15

SPEAKER_03

like pretty much everyone in ai is here it seems it feels like uh apart from the people in the hot tub um uh but uh but yeah it's uh it's been really fun and and like as soon as fable came out of course like it was a stream of people yeah did it empty out like what happened no no i mean there's still yeah it's still a lot of uh people there that's so funny yeah but i was like yeah i need to i need to go test this thing uh so yeah i didn't introduce myself but uh i i'm uh mike taylor i i work as uh the head of ai uh tech consulting so uh work with natalia if you've seen her podcast recently as she was on

29:53

SPEAKER_03

and uh yeah mostly doing a lot of uh ai enablement stuff so helping people uh figure out how to get more ai pilled and be a little bit more like uh like us weird people who uh test fable and hot tubs

30:12

SPEAKER_04

uh amazing awesome and uh what what's the first thing you're gonna do with table uh so i i'm uh

30:18

SPEAKER_03

i'm publishing a book on dspy at the minute it's my my second book first was on prompt engineering um and uh it's all drafted so it's like ready like i'm not allowed to use ai uh to to write it um uh and actually ai wasn't like that good at dspy anyway when i started writing it so um it wasn't a big loss but uh but yeah it's uh the the main thing i want to do is like figure out if there's anything i'm missing i think one thing fable was good at when we were testing it was uh figuring out from a lot of context what was meaningful and i i think with previous models i haven't really seen it like make

30:57

SPEAKER_03

good judgments um and fable still i would still prefer like human judgment over fable but i would say that it could at least supplement and say like hey mike you kind of missed this uh major thing so that that was like that was one of the first things i wanted to try cool um do you think how how much

31:15

SPEAKER_04

have you taken a look at the kate bench stuff and do you think that fable will help there yeah that's

31:20

SPEAKER_03

another good one um so if people don't know uh kate bench we're trying to automate kate we just went

31:27

SPEAKER_02

through we did go through it we just didn't call it kate bench but yeah yeah we were calling it kate

31:32

SPEAKER_03

kate bench yeah there we go you're famous um first person to be benchmarked um uh yeah so um yeah so so the the stuff we're working on there um generally speaking uh the the edit corpus that we have it's like thousands thousands of edits right but um i think it's only from 14 documents is that is that right dan

31:59

SPEAKER_04

from 14 documents no it's like from three four thousand documents today

32:04

SPEAKER_03

oh okay interesting yeah um yeah so so like that was one of the things validated edits or something from 14 but there's a lot of like that oh yeah okay for the validate okay yeah cool so yeah so so one of the problems i was running into the models it just it just wasn't like learning uh general um uh like rules or like general patterns um uh it was basically like over indexing on on a few small things um and then uh and just like not really predicting what the right text would be so um i think the the main thing that fable can help with here is actually maybe just like

32:42

SPEAKER_03

redesigning the experiment to some degree uh because uh the experiment we're running is like kind of complex it it has um like the evals uh basically happen all at once like across the whole corpus uh so so um when when you're testing whether it did a good job of finding the same edits as kate one of the ways it tests it as i understand it is that um it it decides like which ones uh uh kate uh like which which edits were the most meaningful and because it has like a ranking and uh that's like one of the big things that's tripping it up um and and when i've done optimization in the past i've

33:18

SPEAKER_03

done it more row level where uh it's like um you you you uh you can test each individual edit in isolation from all the other edits um so i think it'll be interesting to see what fable thinks about the architecture of the experiment and whether we can make it like make it the right shape for something like dspy jepa which is like a prompt optimization tool um or fine tuning um so that

33:42

SPEAKER_04

that's like that's something i want to ask him that's that makes sense so i actually did ask it and it um let me tell you what it said um it's pretty interesting uh it said uh the programs engineering and honesty are genuinely excellent blah blah blah um however it spent five weeks and 680 agent threads optimizing a target that has never been validated a strict recall against kate's exact historical edits with a 70 goal except for anyone measured whether even kate can hit 70 against herself the evidence is scattered through the repro says she can't that the sealed failures at 47 and a half 50 percent may actually be near the human ceiling

34:27

SPEAKER_04

meanwhile the two things that did work which is span conditional repair and a live suggested suggestion loop were sidelined and the single cheapest high information action which is once human spending two hours judging 60 cases has been sitting idle since june 23rd um so uh it thinks that repair at a known site is learnable cheaply so that means like if we've identified that a particular site is something that there's a site that kate would fix it knows what the we it'll know what the fix would be um but detection and selection is unsolved so uh it's been really hard for us to

35:16

SPEAKER_04

detect which parts of the of a paragraph for example kate would notice and that doesn't seem like something that we can uh fix easily um and let's see where is it how is it determining that i would not

35:37

SPEAKER_02

be that i would be 50 accurate against myself not 70 70 percent well that's why we're getting that from

35:44

SPEAKER_04

the like do the one that i've been yeah okay be like do you make the same edit twice right but i've

35:52

SPEAKER_02

even only done that a handful of times is that what it's taking from just even though uh oh i think it's

35:59

SPEAKER_04

just saying it doesn't know it doesn't know yet oh okay okay not that you're not accurate against yourself because we don't got it yeah so okay so if you just got here we are ripping fable tokens from the jacuzzi in cabo um we've got a fable prompt library so if you want to start ripping fable tokens like we are you should go to every dot to slash p slash claw dash fable dash five dash prompt dash library it is but i promise it's worthwhile it's got all the prompts that we at every are using to rip fail fable tokens and it also has prompts from people like mike krieger who's on my podcast

36:39

SPEAKER_04

when people first came out talking about it i think it was was that for people i don't know it was like around the fable time period um and uh uh it's pretty cool you should check it out it was yeah if you don't know every is the only subscription you need to say at the edge of ai we do ideas apps and training on the idea side we have a daily newsletter whenever something happens in ai like today we uh we publish five checks we publish our own commentary and our own results in using these models to figure out what works what doesn't work uh what they can be used for and uh what's good for what use case uh we also

37:18

SPEAKER_04

have apps uh we have a suite of apps that we build internally with these models that help us work at the edge of ai everything from quora which is an eight-day big email inbox to um sparkle which is a file organizer we have a bunch more we got monologue which is a whisper flow speech to text mark application app which is all time you should check it out and then we do training so we do live streams like this one we do courses we do um uh consulting with large companies that's that's part of what mark mike does anyway if you don't know ever you should you should check it out um normally i'm not in a hot tub

37:52

SPEAKER_04

uh this is a this is a special so i don't want to promise more hot tub time uh this is more of like like a special thing for people but uh in general we have a lot of a lot of really really cool stuff that we do um now where were we uh does anyone remember where i think i think getting to my head

38:14

SPEAKER_03

because i don't remember where we're talking about uh uh cape bench and um and like how to measure like we actually one thing i wanted to point out here is that because i did a bunch of work with ai personas last year uh for my project i was working on um and uh we found that we we basically could never reconstruct what a certain person actually did uh the accuracy was always pretty low about 60 percent uh where we got much better results was um in uh predicting something that the person could plausibly do as in like it is like a uh something that that person would agree yeah this this is something i

38:53

SPEAKER_03

might do uh because um the world you know like you edit the same thing twice and there's going to be a different result in some cases depending on what mood you're in if you had lunch today and you're like there's like a million like there's a million studies like this out yeah like say with judges you know like they my one of my favorite studies is like they found that uh judges are much stricter if they if they uh hadn't eaten lunch yet um you know things like that there's there's all sorts of variables that you're trying to capture and they're not necessarily important variables to capture um so so i would say

39:25

SPEAKER_03

that like the important thing is like do like would kate agree with the edit um once fable has made it i think like that's like maybe that's what we can design study around or maybe fable can help us uh

39:36

SPEAKER_04

organize that i agree yeah i think that the strict exact thing is probably worse than does kate agree if this is a good idea or not i have i just hit our first lutnik guardrail which is i asked fable to put together a report on all of my token spending because i've been being hounded we're sort of like a real company now and so i can't just like rip tokens unaccountably anymore anymore unfortunately so i've been asked for a report of all my token spend and i asked fable to do it and it got triggered on a command that was like basically killing a process on my computer and trying to do some work with chrome and that silently like i would say like 20 minutes ago or so

40:45

SPEAKER_04

trip tripped it up and now it's back to opus 4.8 what i'm really curious about actually is if i flip the model back if it'll keep going and be okay keep going um so like i don't i don't know if this would have tripped previously but it's a very innocuous it's a very innocuous um

41:20

SPEAKER_02

command so your mileage may vary and what do you what do you think that's what is that sorry about the noise in my background um but what is that saying in your mind about like the kinds of space bars

41:35

SPEAKER_04

it's now putting up um it looks like i mean the thing that i got tripped up on is ending a uh ending a process uh and doing something in my browser so there's something about maybe controlling your computer that makes it like a little bit more sensitive i would guess but i to be honest i don't know we're gonna have to do some more do some more uh thinking and talking um so luke face he says lol i flipped past this because i only read the title and it looked like an indie hacker bro building in public slot fest push-ups and tech on bs hahaha but then he said but you dudes in hot tubs

42:18

SPEAKER_04

is awesome which i thoroughly agree with um and uh and he was requesting techno but then said lawnmowers the same thing so thank you luke face and i agree rich in paradise um who is probably i think that rich paradise is probably howard lutnik's um handle only hackers and processes so that seems like something that howard lutnik would say um howard if you would like to come on the show and talk to us about um these new classifications we're uh we're very interested in talking about it it does seem so i asked fable to continue after it got flipped back to opus i asked to continue and it does seem that it is

43:07

SPEAKER_04

back on the same thread so if you're trying to use this and it gets um if you're trying to use this and it gets uh flipped to opus you can flip it back i i would guess depending on the um depending on the conversation so in this conversation i think it's like it was a pretty low risk conversation i'm just asking it to tally up my token spend and there was one command that flipped it and

43:35

SPEAKER_03

then i just restarted and it seems to work so that's interesting one thing people saw uh from from the last um uh time we had fable was uh if they had something in memory that would trip it quite often as well so there's a case of this guy on twitter was saying that because he is a biologist uh it was like remembering biology stuff and then tripping uh some of those safeguards so like yeah i like you know dan i'm like famously against uh having memory turned on i wonder if that that's gonna be more more normal for people to turn it off yeah to turn it off just to control the context when yeah

44:16

SPEAKER_03

also like i you know the tokens are so expensive i don't know like i want to be like really uh judicious about what tokens go into fable you know when i'm paying for usage

44:31

SPEAKER_04

all right what else mike what else is on your mind we haven't talked about

44:39

SPEAKER_03

bringing up my hit list i made a made a little note actually it's funny me and me and kieran were at the conference yesterday together he gave a talk on compound engineering and uh when uh they said they were going to release fable he like bolted up and he was like let's go and i was like wait hold on hold on it's not out yet uh and then he was like he showed me his hit list and i showed him mine we both had like uh you know here's what we're going to test where is what was on yeah i think he may be

45:05

SPEAKER_04

traveling yeah he's traveling back yeah he's traveling oh is he like on a plane or something or in a car

45:11

SPEAKER_03

i'm not sure yeah yeah he was uh he'd be sad to be missing this i know right like we need him here

45:18

SPEAKER_04

this is like yeah like yeah it's it's his time to shine so what was what's on your list and did you

45:26

SPEAKER_03

catch which is what's on his hit list yeah yeah yeah so um so yeah on my on my list um one of the other things i have i wrote a really terrible book uh well i say terrible i think it's pretty good but it didn't do very well um my first book like before the ds before the prompt engineering one it was self-published and only 200 people ever bought it and read it um uh and it was on memetics it's like you know the measuring how things go viral you know uh i used to work in marketing so it was top of mind um and uh i remember someone gave me some feedback and he was like bro you just need to

46:04

SPEAKER_03

watch all these influencers he was telling me all that you know the all the twitter gurus of like yeah they try and make everything go viral you need to watch this this and this person and then and then just take all the core material and just make it um more like digestible for people and i'm like yeah i'd love to do that but i just like spent a like a long time writing it i didn't want to ever look at it again so now i feel like enough time has passed that i'm like i just wanted to give that to people and just be like what can you do with it do you make this actually like more uh available to

46:36

SPEAKER_03

the general public right like more interesting to the average person uh so that's like another thing i had on my list to let rip the general public right um yeah and another one is uh i just um i i've been um i've been working on uh rally like uh the past year uh this is uh the ai it's known as like uh focus group uh project uh i was working on it uh more full-time last year and now it's just a side project um uh because i have a job uh but uh uh but uh i have been afraid to like commit stuff to that uh to that repo because uh my um my co-founder's uto like he also works full-time um he is uh like a much better engineer than me

47:22

SPEAKER_03

uh and and he set up a bunch of stuff that like good engineers set up uh but that's made it like less intelligible for me uh so um like that's the other thing i wanted to do is kind of take some of the ideas that i i had and just like run it through like code review with fable um i have like a bunch of pr sitting there that he hasn't had time to review and so i think like that's another good piece case just like increase my confidence and that i'm not messing something up with his good setup yeah i mean if you

47:49

SPEAKER_04

can like if you can just blame it on fable like what does he yeah exactly yeah yeah well yeah i mean if

47:57

SPEAKER_03

fable says it was good then like who you know who are we to judge yeah yeah i do actually feel that way

48:06

SPEAKER_04

like i just remember when it first came out there was this palpable change in momentum where kieran was on this tear where he was on track to spend like 1.5 million dollars in fable meetings but his product was like getting better second by second it was like kind of crazy you know and and for me looking at it i was like that definitely gives me a lot more confidence that i can and should be committing production code to some of these some of these apps we're working on um which is now giving me ideas like maybe i should try to try to ship something to production this week while i'm on vacation i think we just lost mike oh his hotel internet what i think is interesting

48:51

SPEAKER_02

also and i know obviously anthropic was not in control of this timing around fable and ideally would not have had it taken offline to then be restored online today is we're actually getting ready to publish a vibe check with sonnet 5. um how's that going i'm i'm waiting for the copy to land with me it's quite extensive it's going to be a big uh thoughtful you know typically thorough vibe check

49:15

SPEAKER_04

um how do people like it i've been on vacation so not it's not that good right the vibes are off

49:24

SPEAKER_02

i've kind of never seen a uh a response like this where it was just kind of not just meh of like oh it's not better than this it was like it's actually just worse and it's not good at this and it's not good at this like who is this for it was actually a genuine question of who is this for

49:41

SPEAKER_04

interesting that was like a i mean that was the response to opus 4.8 i think initially if you remember or maybe yeah maybe it was 4.7 it was one of the like not yeah this one is pretty like i kind

50:01

SPEAKER_02

of feel like definitively um i've sort of never seen so uniformly like you know within our reach test rankings of you know red yellow green or gold it's like red uniform except for yes except for jack for jack it's a green because and i had a whole conversation with him about it this afternoon actually our senior editor jack chang who's a writer as well as a builder uh and product guy um he uh he's essentially like rebuilt codex on his own so he he he's really deep into this stuff um and he uh he said it was green because it was like he he is less likely to be giving sonnet or another model like that a

50:42

SPEAKER_02

very long-running task to just set it off and do something the way that kieran would be like go and do this and i'll come back in two hours and see if it's done he's working much more collaboratively with it and he found that in working collaboratively with much shorter chunks of work that it actually um was really good but he said he's the only one who likes it it's really interesting it's it's also

51:04

SPEAKER_04

interesting because well one the last model that we all didn't really like from anthropic was i think of four seven and that was the last model that they dropped without an eap which they did for this one too so we didn't have um so i don't know why maybe they were distracted with fable but also maybe sometimes for some of these models where they're like it's not clearly amazing they just sort of they're just dropping it because they're you know uh uh without an eap because they're you know uh less confident it's great i'm just going back to look at our coverage for a moment but also like people often don't like a model and then decide to like it later because

51:57

SPEAKER_04

models just because you don't like it initially what it what that usually means is the way that you're trying to use it is not good or not better than a previous model right but then often what happens is people find new ways of using it with new powers that like they didn't realize were there three weeks later or four weeks later or whatever which we try to solve for by getting access to it early but we don't have early access to this model so this is a very like we're just gonna have to give you what we have when we have it but i'd be really curious to see how it ages in the next

52:28

SPEAKER_04

you know five weeks especially if you're using fable as a orchestrator for it if you're not using sonnet directly but you're just having fable use it right yeah i'm just going back and checking

52:43

SPEAKER_02

our coverage of four seven and we were we were we were pretty lukewarm when it first came out and then about a week later published our vibe shift that we've actually had a little bit of a change of heart on four seven the more people are using it um and using it for different for different things so uh it's not it's not uncommon but it was just that it was if if you were in our slack it was just a sort of

53:10

SPEAKER_04

very consistent stream of why why yeah um kate i am about to pass away from dehydration um

53:23

SPEAKER_02

we don't want that to happen especially not on camera

53:27

SPEAKER_04

i'm sitting in a hot tub and drinking and it's okay not even it's like one almost 1 30 so this was fantastic thank you for joining i'm going to be ripping fable tokens for the rest of the day if you want to see more follow me and every on hex but more importantly check out every.to it's the only subscription you need to say at the edge of ai we are going to be publishing um we are going to be publishing a list of fable prompts uh which i think we already have out uh let me just let me bring that up uh we have a fable prompt library so if you want to use fable to the fullest extent go to every.to

54:08

SPEAKER_04

slash p slash cloud fable five prompt library with dashes in between there it is subscribe to every we've got ideas apps and training and enjoy the return of fable on this july 4th weekend uh and we'll see you next time

54:32

SPEAKER_00

enjoy cabo thank you bye ! name ! name name name name !

55:08

Thank you.

55:38

Thank you.

56:08

Thank you.

56:38

Thank you.

57:08

Thank you.

57:38

Thank you.

58:08

Thank you.

58:38

Thank you.

59:08

Thank you.

59:40

Thank you.

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note