AI Engineer

How long can your skills be before your agent forgets what you told it? — Laurie Voss, Arize AI

2105 summary words 9 min summary Watch video

Start with the signal

9 min read

Summary

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: Frontier models' raw capacity to follow simultaneous named instructions has improved roughly 10x in a year—from failure around 200-300 constraints to roughly 2,000-5,000 on the tested leaders—but reliable agent design now depends more on verification, model-specific failure handling, cost, and latency than prompt compression.
  • Why it matters: This directly challenges the old architecture assumption that long skills files must be aggressively sharded across specialist agents, while underscoring that apparent compliance can mask partial completion, refusals, context degradation, or reasoning failures.
  • Best use: Use this as an input to redesign OpenClaw/agent skill-loading policies and to define a model-specific eval suite that tests instruction coverage, completion integrity, ordering sensitivity, and cost-latency trade-offs.

Executive Summary

Laurie Voss replicates the IfScale benchmark, which asks a model to write a business report containing a growing list of exact required words. The benchmark treats each required word as a proxy for a discrete instruction such as including a pricing section, honoring a legal disclaimer, or avoiding a prohibited phrase. His replication of year-old frontier models found the original result was real: performance began breaking down around 200-300 simultaneous constraints, with substantial loss by 500.

Using the same test against then-current models, Voss found that GPT-5.5, Claude Opus 4.7, Gemini 3.1 Pro, and DeepSeek V4 Pro initially achieved 100% accuracy through the benchmark's original 500-word maximum. After expanding the test to 10,000 required words, he found the new practical boundary varied sharply by model: DeepSeek began forgetting around 750 rules, while leading models could track roughly 2,000 constraints and the strongest results extended to about 5,000. GPT-5.5 reportedly sustained 99% accuracy through 5,000 rules.

The central operational warning is that aggregate accuracy conceals qualitatively different failures. DeepSeek silently drops constraints; Claude may refuse due to safety classifiers triggered by combinations of otherwise random words; Gemini may consume its thinking-token budget and return little useful output; and GPT-5.5 may begin a polished answer but abandon it partway through. Therefore, a successful API call or credible-looking opening is not evidence that an agent completed its assigned work.

Voss's conclusion is not that giant prompts are solved. The prior compression problem has shifted into a verification and economics problem: long skill files are increasingly feasible, but they add cost and latency and do not prove robust reasoning through conflicting or reordered rules. Supporting research cited in the talk reports 30-50% long-context accuracy degradation before context-window limits and major sensitivity to phrasing and ordering, making production evals and output monitoring essential.

Key Takeaways

  • Claim: The old approximately 200-instruction ceiling was a real empirical limitation for prior frontier models, not merely prompting folklore. | Evidence: Voss reran IfScale on the surviving previously tested APIs—GPT-4.1, Claude Sonnet 4, and Gemini 2.5 Pro—and reports results matching the original paper within noise boundaries; around 200-300 rules models began to degrade, and by 500 they lost roughly 30-50% of required constraints. | Implication: Any agent architecture built around sub-200-rule prompt budgets was rational at the time, but that budget should no longer be treated as a permanent system constraint. | Caveat: IfScale uses exact required words in a synthetic report, so it measures a favorable proxy for discrete instruction retention rather than full real-world reasoning over a complex skill file.
  • Claim: Current frontier instruction capacity improved by about an order of magnitude, allowing many more named constraints in a single prompt. | Evidence: The same benchmark was expanded from 500 to 1,000, 2,000, 5,000, and ultimately 10,000 required words because GPT-5.5, Claude Opus 4.7, Gemini 3.1 Pro, and DeepSeek V4 Pro all aced the original 500-rule cap. Voss reports usable boundaries near 2,000 for some models and near 5,000 for the best; GPT-5.5 reached 99% accuracy through 5,000 rules. | Implication: A complete style guide, compliance rules, brand requirements, and operational constraints can increasingly live together in one agent skill or prompt instead of being fragmented solely to avoid instruction loss. | Caveat: Capacity differed substantially across models, from DeepSeek failing materially near 750 rules to other models reaching 5,000 or more; these figures should not be generalized across providers or model versions.
  • Claim: Model choice must account for failure mode, not just an instruction-following score. | Evidence: DeepSeek V4 Pro silently forgot rules and was dropping nearly half by 2,000; Claude Opus 4.7 issued API-level refusals when random word combinations looked dangerous; Gemini 3.1 Pro remained strong to 5,000 but spent its thinking budget at high density and returned little usable output; GPT-5.5 sometimes began the report and then explicitly stopped partway through because the task was unreasonable. | Implication: The control plane should classify and monitor distinct failure classes—silent omission, safety refusal, token-budget exhaustion, and premature/partial completion—rather than treating all failures as generic low quality. | Caveat: Claude's reported behavior was partly benchmark-induced: random vocabularies included combinations such as anthrax and cyanide, requiring Voss to filter words through OpenAI's safety filter before Claude could complete runs.
  • Claim: Longer prompts have shifted the primary engineering decision from a hard capacity limit to a cost-and-latency trade-off. | Evidence: Voss argues that teams previously compressed skills below about 200 rules and created chains of sub-skills or specialized agents; with thousands of constraints now feasible, the limiting question is whether the additional input tokens are worth the slower and more expensive call. | Implication: Consolidate instructions when it removes orchestration and handoff risk, but retain modular loading where dynamic relevance, response time, or token economics justify it. | Caveat: A model accepting a very large prompt does not establish that it reasons correctly over it, resolves conflicting rules, or executes all requirements coherently.
  • Claim: A polished output cannot be assumed complete, particularly with models that fail late rather than refuse early. | Evidence: Voss calls GPT-5.5's partial-report behavior more dangerous than Claude's refusal because it can produce hundreds of words of plausible output before stating it will not continue, leaving many requested constraints unmet. | Implication: Production agents need completion and constraint-coverage checks on outputs, with explicit detection for truncation, abandonment language, required-field absence, and task-specific success criteria. | Caveat: Manual review can catch this in low-volume workflows, but it does not scale and may miss silent omissions such as DeepSeek's failure pattern.
  • Claim: Raw constraint-retention results should not be mistaken for reliable long-context reasoning or stable instruction following. | Evidence: Voss cites Chroma's context-rot work across 18 models, reporting 30-50% accuracy declines on long inputs before the formal context limit; he also cites 'Revisiting the Reliability of Language Models in Instruction Following,' covering 46 models, which found that rewording or reordering equivalent instructions can radically change compliance. | Implication: Evaluate real skills with conflicting constraints, natural document structure, alternative phrasings, and order permutations; do not certify a prompt design using one fixed wording or one synthetic benchmark. | Caveat: The cited finding that coherent, well-structured text may degrade more than randomly shuffled input was noted by Voss as surprising and not yet explained in the talk.
  • Claim: Running bespoke instruction-following evaluations is affordable enough to be a routine engineering practice. | Evidence: Voss reports spending $209 for 2,300 calls across seven models to run this research and points to additional emerging benchmarks including Firebench, CCRbench, and Guidebench. | Implication: Ken can fund model- and workflow-specific regression tests rather than relying on vendor claims or static prompt heuristics established six months earlier. | Caveat: Token cost for a production workload may be materially different from this benchmark, particularly when prompts, tool traces, outputs, and reasoning budgets are larger.

Detailed Brief

Benchmark interpretation and what it does not establish

  • Claims: IfScale measures whether a model can retain and execute many independent, named constraints, making it directionally useful for assessing skills files.; The metric is a ceiling rather than a proof of real-task capability: failing to include simple specified words implies likely difficulty with more complex instructions, but passing does not guarantee sound reasoning.; The relevant research question has broadened from token-window size to robustness under large sets of messy, interacting requirements.
  • Evidence: The benchmark asks for a business report that contains each item from a specified vocabulary exactly, then scores the percentage of words included.; Voss names Firebench, CCRbench, and Guidebench as newer efforts intended to assess instruction following under more realistic multi-constraint conditions.; The talk's test code and data were said to be published via a GitHub link, although the URL is not captured in the supplied transcript.
  • Caveats: Exact-word inclusion does not test prioritization, semantic interpretation, cross-rule conflict resolution, factual correctness, tool use, or multi-step agent execution.; Model releases are fast enough that a model-specific chart becomes stale quickly; Voss notes Opus 4.8 arrived shortly after he tested Opus 4.7.
  • Implications: Treat benchmark results as a trigger to update assumptions and run local tests, not as a universal maximum prompt-size policy.; Version-pin evaluation results and re-run them on model upgrades before changing skill composition or routing decisions.

Design consequence: simplify only the fragmentation caused by obsolete limits

  • Claims: The prior need to route users through a 'Byzantine' network of sub-skills was largely an adaptation to weak constraint retention.; Consolidation can reduce specialized-agent handoff failures because more of the governing policy can be present at once.
  • Evidence: Voss characterizes 2,000 named constraints as enough to contain an entire style guide, including brand rules and legal disclaimers.; The talk frames the former approach as sharding instructions across many agents and hoping their handoffs remained clean.
  • Caveats: Consolidating prompts does not eliminate the need for modularity based on context relevance, tool boundaries, permissions, security isolation, or economical retrieval.; Large inputs can still produce high latency and high cost even when the model follows them accurately.
  • Implications: Separate architectural modularity that exists for authority, retrieval, and workflow clarity from modularity that existed only to fit an obsolete instruction-count limit.; Use selective instruction loading as an optimization policy, not as an untested belief that models cannot handle a comprehensive skill.

Notable Concepts & Terms

  • IfScale: A benchmark that tests simultaneous instruction following by requiring a model to include a large list of specified words in a generated report and scoring constraint coverage.
  • Instruction density (n): The number of rules or named constraints presented simultaneously; this is the horizontal scaling variable in the benchmark.
  • Skills file: A file or prompt containing agent operating rules, edge cases, formatting requirements, tone guidance, and other constraints; the talk argues these may now be materially longer than older design norms assumed.
  • Context rot: Accuracy degradation on long inputs before reaching the advertised context-window limit; cited research found drops of 30-50% across evaluated models.
  • Thinking-token exhaustion: Gemini's observed failure pattern at extreme instruction density: it allocates its output budget to internal reasoning and leaves too little capacity for a useful final response.
  • Completion integrity: Whether an output actually finishes the assigned work rather than merely appearing credible at the beginning; especially relevant to GPT-5.5's observed mid-response abandonment.
  • LLM-as-judge eval: Using another model to monitor or score outputs for compliance and failure conditions when manual review is impractical and failures are not exposed by an API error.
  • Prompt-order sensitivity: The cited phenomenon that equivalent instructions can be followed very differently when their wording or ordering changes, undermining confidence from a single successful prompt formulation.

Operator Notes / Why Ken Should Care

  • Replace any inherited fixed rule-count limit for skills files with per-model, per-workflow measurements; test the currently pinned production model rather than relying on six-month-old prompt heuristics.
  • Build a regression suite that includes required-constraint coverage, conflicting rules, natural structured instructions, shuffled/reordered instructions, paraphrased instructions, and representative task outputs.
  • Instrument completion integrity: detect early exits, truncated deliverables, refusal-like language embedded in otherwise valid output, missing required sections, and deviations from output schemas.
  • Add separate handling policies for safety/API refusals, silent rule omissions, excessive reasoning-token consumption, and partial completions; each needs a different retry, fallback, or escalation path.
  • Compare a consolidated-skill design against selective modular retrieval using quality, latency, token cost, and handoff-error metrics; do not preserve agent fragmentation solely because older models had a 200-rule ceiling.
  • Re-run the suite on model-version changes and routing changes, and retain model-specific failure signatures as part of deployment criteria.

Source/Metadata

  • Title: How long can your skills be before your agent forgets what you told it? — Laurie Voss, Arize AI
  • Transcript words: 6558
  • Duration seconds: 1345
  • Timestamp note: No timestamps or chapters were present in the supplied transcript; the transcript also contains a substantial duplicated closing segment.
Full transcript 4063 words · 29 min read
0:00

All right, hello everybody.

0:12

Thank you for coming to this delightfully nerdy talk. This talk has a really long title, so let me give you the short version up front. You write skills files and stuff them full of instructions. At some point the model stops keeping track of all of them. The question is, where is that point? At what point have you put too many instructions in your skills files? And the answer has changed a lot in the last year. I'm Laurie, I'm head of developer relations at Arise AI. In a former life I co-founded NPM Inc., so some of you may know me from the days of JavaScript. These days I spend a lot of time thinking about AI and how to test it.

0:40

A few months ago I was at AI Engineer in Miami, which was a good conference. And I was watching a talk by Dexter Horthy. It was a good talk, it was not about this topic at all. But while he was giving that talk, he mentioned as an aside that an agent can follow up to about 200 instructions before it starts forgetting those instructions. And then he moved on in his talk, and it was entirely an aside. And he mentioned that that figure is from 2025, so things might be better now. And I stopped listening for a second because I was thinking, 200 instructions is not very many instructions at all, right? A decent skills file blows past 200 instructions almost immediately.

0:59

If the user says X, do Y, always include a section on Z, never use the phrase W—every one of those is a separate instruction. And if the model quietly stops tracking them after 200, that's a really hard ceiling on the complexity of what you can build. So I wanted to know where he got that number first, and I wanted to know if it was true. So you know the feeling that I'm talking about. You write this big beautiful skills file, pages of rules, edge cases, tone, formatting. You hand it to the agent, it does the thing, and you look at the output and go, did it actually pay attention?

1:15

Did it actually follow all of these rules, or did it just do what it felt and give me a close simulacrum of what I was expecting? You can't really tell. Or can you? More on that later. And so you live with this low-grade anxiety every time you hit run, and that feeling is what this research is about and what we're trying to find out if we can avoid. So here's my promise for your next 18 minutes. I'm going to show you where that 200 number came from, whether it's still true, and what the real number is today, because it moved by an order of magnitude.

1:30

And then we're going to talk about what that means for you to take away—how long your skills and prompts can actually be and what changes you should make to your workflow as a result. So the 200 number isn't folklore. It comes from a real benchmark called IfScale, from a paper by this person whose name I'm going to mess up, Jeroslawicz, and co-authors last year. And the test is beautifully simple. Here's how IfScale works. You ask the model to write a business report, and you give it a list of specific words that it has to include exactly in the report. Include the exact word customer, include the exact word revenue, and so on for as many words as you want.

1:45

Each of those is an instruction that it has to follow. And then you count how many of those exact words showed up. So because the test is so simple, you only have to keep two numbers in your head. One is density, which we call n. That is how many rules we're talking about at once. And the second is accuracy, which is the percentage of those rules that it was able to actually follow. Now you might say that including random words in a report is not the same as following real instructions, and fair enough, and we're going to talk about that. But the key words are a proxy. Include the word revenue is the same shape of task as include a section on pricing, right?

2:03

Or never use this phrase. It is a discrete named constraint that you've told the agent that it has to follow. If a model can't track 200 words in one prompt, it's definitely going to struggle with 200 more complicated instructions. So if anything, it's going to do worse. So this number is a ceiling. This number is as high as you can go. If you give it more complicated instructions, the number is probably going to get lower. And 200 is a really low ceiling. So before chasing new models, you have to do good science, which means that you have to replicate the old result and make sure that the 200 ceiling is real. So I reran the original benchmark.

2:24

The original paper tested a whole batch of models, and models live and die really fast. So by the time I got around to doing this testing, only three of the models in the original set of 10 models that they used were still available via any kind of API. So they were GPT 4.1, Claude Sonnet 4, and Gemini 2.5 Pro. Those were models that were available 12 months ago that are still available now. And that is why we tested those three, because they were what was left. And since I first published this research a couple of weeks ago, one of those three models has been retired. So this was the last possible time that I could have run this test.

2:41

So of that lineup, we're already down to two. So don't get attached to your models. Here are the results that we got replicating the original IfScale finding. That is accuracy on the vertical axis. So it starts at 100% and begins to fall off. And then the number of rules going up along the bottom on log scale. So every time it gets halfway across, it has doubled the number of rules that it's dealing with. So by 500 rules, you're losing 30, 40, 50% of them. Our curves matched the results in the original paper within the noise boundaries, so the finding was real. A year ago, somewhere around 200 to 300 rules, frontier models started falling apart.

3:12

That is a really low ceiling. So that is our baseline. And now comes the fun part where we took the exact same test and pointed it at the current frontier, or rather what the current frontier was when I ran this test. So I ran GPT 5.5, Claude Opus 4.7, because 4.8 came out a week after I ran this test, Gemini 3.1 Pro, and Deep Seek V4 Pro. So I gave them the same prompt, the same words, the same everything, and I immediately ran into a problem, which is that they aced it. They all scored 100% immediately on this test. Absolutely no bugs.

3:27

So we built a test to find the ceiling, and the models had walked straight through the ceiling without noticing that the ceiling was there.

3:32

And that was a problem, because the benchmark was written to top out at 500 words, so I had to change the benchmark in order to be able to find the new ceiling. So I moved the goalposts. I gave it more words to include. I doubled it from 500 to 1,000. I doubled it again from 1,000 to 2,000, and I kept doing that until I hit a 10,000-word vocabulary, and that is where I began to find the ceiling of what models can do these days. So let me put up this—this is the money slide. This is the results. Remember, log scale on the x-axis there, so it's going from 500 to 1,000 to 5,000 to 10,000.

3:50

So it looks like that scale is falling off a cliff, and it's actually happening over 1,000 numbers. But look how far to the right these new curves get before they bend. A year ago, they were falling over at 200 to 300 instructions, and now, depending on the model, the boundary is closer to 2,000, and for the best of them, it is up to 5,000 instructions before they begin to fall off a cliff. So in about 12 months, frontier models got close to 10 times better at following instructions simultaneously. That is the headline finding, and there is a lot of nuance that we need to get into. The capacity to track 2,000 named constraints in a single prompt is there.

4:04

And that's really interesting because I think—I don't know if everybody else feels this way—but it felt to me like the jump from GPT 5.1 to GPT 5.5 was incremental, right? It didn't feel like we'd gotten 10 times better, but this is a test that really matters to a very practical thing like how long can my skills file be? And in the course of a year, we got 10 times better. And the thing that gets me is that this benchmark is barely a year old. A year later, 500 is a rounding error, and this keeps moving under my feet.

4:18

I tested 4.7, Opus 4.8 is even better. So this chart is a little out of date already, which is the whole point. If you set your engineering assumptions about how skills files should work,

4:23

in a single prompt is there. And that's really interesting because I think I don't know if everybody else feels this way, but it felt to me like the jump from GPT 5.1 to GPT 5.5 was incremental, right? It didn't feel like we'd got 10 times better, but this is a test that really matters to a very practical thing, like how long can my skills file be? And in the course of a year, we got 10 times better. And the thing that gets me is that this benchmark is barely a year old. A year later, 500 is a rounding error, and this keeps moving under my feet. I tested 4.7, Opus 4.8 is even better. So this chart is a little out of date already, which is the whole point. If you set your engineering assumptions about how skills files should work, about how long your prompt can be, and you did that more than about six months ago, you are incorrect now, and you should probably be re-engineering how you do stuff.

4:25

But there is more to this story because the way that the models failed changed dramatically, and the way that they failed is very important. This part was a completely unexpected finding when I started running the experiment, and it totally messed up my test to start with because the old failure mode was boring. They would just forget instructions, and I could measure how many instructions they had remembered or forgotten, but the new ones fall apart in their own weird, extremely on-brand way. So let me introduce you to how these four models fail.

4:30

DeepSeek 4 is a traditional model. It just forgets things. It doesn't have any drama. It starts forgetting instructions around 750 rules, and by 2000, it's dropping nearly half of them. So it just forgets, which frankly is the failure mode that I trust most because it's predictable, it's very easy to measure, and the other models were not nearly as cooperative.

4:32

Opus 4.7 would decide repeatedly that the test was dangerous, and what it would do is it would refuse at the API level to complete the test. I didn't know that there was an API response that you could get from Claude where it was like, no, I could do this, but I'm not going to. But that's absolutely an API-level response that Claude supports because they care so much about safety, and I started getting those all of the time, and the reason that was happening is because Claude has a very sensitive safety classifier, and if you put in certain combinations of words, like say anthrax and cyanide, it decides that the whole request is dangerous, and it bails out, and if you remember what my test does, my test is throwing five to 10,000 random words into an instruction file, and so my randomly selected words contained all sorts of things that looked dangerous in combination to the safety filter, and so it kept bailing, saying that I was asking it to make a bomb or something.

4:33

So we had to, to get Claude to cooperate, I had to take all of my words and run them through OpenAI's safety filter and filter out all of the naughty-looking words so that it could get anywhere. Once I'd given it that, Claude did really well. But the failure mode is that Claude is more likely to decide what you're doing is dangerous very early on, at even two or three hundred instructions, if what you're doing is something to do with medical advice, because medical things often are dual purpose, they can be dangerous, they can be safe.

4:36

So the third failure mode was Gemini 3.1 Pro. Gemini is rock solid all the way out to 5,000 instructions. It does extremely well. Genuinely one of the best on the chart. And then past that, it gets weird. It doesn't forget the instructions. It gets overwhelmed by the instructions. What it tries to do is it uses thinking tokens to make sure that it is following all of the instructions at once. And when the number of instructions gets really high, it uses all of its thinking tokens. It uses its entire token budget thinking, and then it doesn't give any output. It gets to 9,500 tokens worth of thinking, and then give you a 500 word response, which doesn't contain any of the tokens. So it thinks itself into a corner and runs out of room to actually answer, which is very expensive and totally unhelpful, which is on brand, isn't it?

4:39

And finally comes the winner, which is GPT 5.5. GPT 5.5 is the best of the lot. 99% accuracy all the way out to 5,000 rules. But if you push it far enough, it is by far the weirdest of the bunch. Because it doesn't refuse outright, it doesn't silently forget. Instead, what it does is it gets frustrated and tells you that the test is stupid. It starts the report, it gets a few things in, and it doesn't start out just saying no. It starts the report, it starts writing the report, and 500 words into the report, it's like, no, this is dumb. I'm not going to do this. And then it politely tells you this is dumb. I'm not going to do this anymore. That is the actual response that it gave me.

4:42

But that was 5,000 words into this business report that I told it to generate. So it's not wrong, right? I was asking for a coherent business report on no particular subject that contains 5,000 random words. You're right, GPT. This is a stupid thing to ask for. Which is a deeply unreasonable request, and GPT called this out on it. But it still counts as a failure in the test, because the half-finished report that it gives you is missing most of the keywords, and it is also the hardest one to detect, because Claude bails immediately. Claude says, no, I'm not going to do this. DeepSeek does its best. But GPT does what looks like a good job, unless you read all the way to the end of the report, where it says, no, actually, I'm going to bail, because this is stupid.

4:45

So if you step back and look at the four together, DeepSeek quietly forgets, Claude gets scared and refuses, Gemini overthinks itself into silence, and GPT 5.5 finishes half of the job and tells you that the rest of it is beneath it. And the point isn't which one of these is funniest, although it is genuinely a little funny. The point is that "did it follow my instructions" no longer has one failure mode. It has four different ways that it can fail, and you can't recognize that failure unless you know which model you're dealing with and what its pattern of failure is going to be.

4:50

So the models got 10x better. They fail in funny ways. Why should you care when you get back to your desk? Because three things have changed to your workflow. The first is that a year ago, the smart move was to keep every skills file very, very short, under 200 instructions, then point off to sub-skills, and a whole Byzantine labyrinth of additional skills files and sub-files and things like that. And you are compressing your instructions to fit into a very small available space, and you don't need to do that anymore. Your skills files can be very long.

4:53

Number two is that if your use case needs 100 specific rules or 300, you can just put them all in the prompt. You don't have to wonder which ones the model silently ignored. And if you've been thinking about your own lived experience of using models, you probably recognize this. You've discovered that you've got less worried about how long your prompt is going to get because the models have genuinely got 10x better at following your prompts. 2,000 named constraints is an entire style guide, right? Like, it's every brand rule, every legal disclaimer. A year ago, you'd have had to shard that across a dozen specialized agents and hope that your specialized agents are handing off to each other cleanly, but now you can ignore that.

4:56

But the third thing is the big one. The question used to be, can the model even do this? And the answer is now firmly yes. Well, reasonably firmly. Is it worth the cost is the new question because you can include 10,000 words of instructions into your prompt, but that is going to be an enormous prompt. It's going to be a very expensive prompt. It's going to be a very slow prompt. So what used to be a hard wall that you would run against has now become a soft trade-off of, is it worth me adding all of these extra instructions if it's going to give me more cost and more latency?

5:01

And now some caveats to head off the Q&A. First and most important, I mentioned this earlier, this is a proxy task. Including random words in a fake business report is evidence that long skills files work. It is not the same as proof that a long skills file works. Also, the models hit the wall at wildly different points. And the answer is now firmly yes. Well, reasonably firmly. Is it worth the cost is the new question because you can include 10,000 words of, sorry, 10,000 different instructions into your prompt, but that is going to be an enormous prompt. It's going to be a very expensive prompt. It's going to be a very slow prompt.

5:18

So what used to be a hard wall that you would run against has now become a soft trade-off of, is it worth me adding all of these extra instructions if it's going to give me more cost and more latency? And now some caveats to head off the Q&A. First and most important, I mentioned this earlier, this is a proxy task, including random words in a fake business report, is evidence that long skills file works. It is not the same as proof that a long skills file works. Also, the models hit the wall at wildly different points, anywhere from 750 to 9,000 plus, so you have to pick your model very carefully. What our test doesn't do is measure

5:50

whether the model reasoned clearly over a giant prompt. So the good news is since I did my research several weeks ago, a whole bunch of people have piled in on this, and now there's good research. Actual scientists have got involved and done, Chroma has done context rot work across 18 models, showing that accuracy on long inputs can fall 30 to 50 percent well before you hit the context window limit. And the weird part of their finding was that coherent, well-structured text is more likely to hit that failure mode than if you just put your instructions into a random order and shuffle them in. I don't know why that's the case. I'll have to read their report.

6:25

So the model can track 2,000, 5,000, possibly 10,000 instructions, but it's not necessarily going to reason clearly over them. It's not necessarily if those instructions conflict, if there is tension between them, it's not necessarily going to get that right. And then there's the other one I mentioned briefly. Claude's refusals are annoying, but they are loud. You get an error, you know it failed. GPT's polite half-finished report is much more dangerous because it looks like a real answer. You have to read the whole thing to notice that it gave up quietly halfway, which means that you can't trust the output. It means you have to read the output every single time

6:54

to make sure whether or not it's working. So the model will accept your 2,000 rules and it will hand you back something that looks, at least to begin with, confident and polished, but could be bailing out halfway through. So as an aside, people always ask me, well, how much did all of this cost me? It cost me $209 to run all of these queries. 2,300 calls across seven models came to $209.

7:23

It turns out novel research doesn't cost very much. And this is the part of the talk where I was saying that you have to check this stuff in production because you can't trust that your model isn't going to silently fail. So you knew I was going to mention evals eventually because I work at Arise and this is where I do that, but there are plenty of plugs for Arise. So I'm just going to say one true thing, which is that if you are building a real AI application and you are giving it genuinely tricky tasks, you are going to run into one or more of these failure modes with a frontier model. And unless it's Claude telling you to fuck off

7:49

at the API level, the only way to know that something went wrong is monitoring your outputs with another LLM. That is an eval and that is what Arise does and I'll leave it at that. I already mentioned that there's been new research since we did our own. Here's another important one. A paper landed testing 46 models called Revisiting the Reliability of Language Models in Instruction Following, which you can bet made my ears perk up after I did that research myself. And they found something uncomfortable, which is that a model can ace a benchmark like ours and still be wildly unreliable. Because if you reword the same instruction

8:34

in a slightly different way, it can make a radical difference to how well it follows those instructions. So the model can follow 2,000 instructions and it can do it really well, but if you put the same instructions, the same 2,000 instructions in a different order, it can suddenly make the model much worse at following those instructions. And how exactly to do that, what is the correct order of instructions to give your model such that it follows them perfectly as opposed to getting confused, is still research that is being done. So capacity went up, but reliability is still a problem. And then this is just a brag, because I was happy about it,

9:10

I'm not a scientist, I did some research, and then a whole bunch of other actual scientists piled in and did real science on the same question. There's now a whole bunch of benchmarks that have shown up to measure this same question. Firebench, CCRbench, Guidebench, are all trying to measure the same thing. How well models follow a lot of real, messy constraints at once. And now the whole field is looking at it, so if you want better science than my 10,000 random words, the real science exists now. So that gets me to where I will leave you. A year ago, the hard part of writing a skill was fitting everything in without the model losing the plot.

9:42

That was a compression problem, and the compression problem is gone. The model will hold your 2,000 instructions just fine. The new hard part is knowing whether it actually did what you said, and that is a verification problem. A verification problem doesn't get solved by writing a better prompt. It gets solved by checking the output every time, the same way that you would test any other code, which is to say an eval. The ceiling moved by 10x in one year, so go back and check the assumptions that you made six months ago about how big your prompts should be, how big your instructions can get, because they might already be wrong. So that is the talk.

10:23

If you want all of the code and all of the data, it is at this GitHub URL. And this other QR code is something marketing made me insert. We are having a World Cup watch party tonight at 5 p.m. You can come to our party. That link is to the luma that will get you into the party. I hope this talk has given you some novel information, or at least a couple of laughs.

10:49

And thank you so much for your time and attention. I had to take all of my words and run them through OpenAI's safety filter and filter out all of the naughty-looking words so that it could get to anywhere. Once I'd given it that, Claude did really well. But the failure mode is that Claude is more likely to decide what you're doing is dangerous very early on, at even two or three hundred instructions, if what you're doing is, you know, contains anything to do with medical advice, because medical things often are dual purpose, they can be dangerous, they can be safe. So the third failure mode was Gemini 3.1 Pro. Gemini is rock solid all the way out to 5,000 instructions.

11:28

It does extremely well. Genuinely one of the best on the chart. And then past that, it gets weird. It doesn't forget the instructions. It gets overwhelmed by the instructions. What it tries to do is it uses thinking tokens to make sure that it is following all of the instructions at once. And when the number of instructions gets really high, it uses all of its thinking tokens. It uses its entire token budget thinking, and then it doesn't give any output. It gets to like nine, you know, if you've given it 10,000 tokens worth, it'll get 9,500 tokens worth of thinking, and then give you a 500 word response, which doesn't contain any of the tokens.

12:07

So it thinks itself into a corner and runs out of room to actually answer, which is very expensive and totally unhelpful, which is kind of on brand, isn't it?

12:19

Which, you know, I would never say that out loud. And finally comes the winner, which is GPT 5.5. GPT 5.5 is the best of the lot. 99% accuracy all the way out to 5,000 rules. But if you push it far enough, it is by far the weirdest of the bunch. Because it doesn't refuse outright, it doesn't silently forget. Instead, what it does is it gets frustrated and tells you that the test is stupid.

12:45

It starts the report, it gets a few, like, that's the thing, it doesn't start out just saying no. It starts the report, it starts writing the report, and like 500 words into the report, it's like, no, this is dumb. I'm not gonna do this. And then it politely tells you this is dumb. I'm not going to do this anymore. That is the actual response that it gave me. But that was like 5,000 words into this business report that I told it to generate. So it's not wrong, right? I was asking for a coherent business report that on no particular subject that contains 5,000 random words. You're right, GPT. This is a stupid thing to ask for. Which is a deeply unreasonable request,

13:25

and GPT called this out on it. But it still counts as a failure in the test, because the half-finished report that it gives you is missing most of the keywords, and it is also the hardest one to detect, because Claude bails immediately. Claude says, no, I'm not going to do this. Deep Seek does its best. But GPT does what looks like a good job, unless you read all the way to the end of the report, where it says, no, actually, I'm going to bail, because this is stupid. So if you step back and look at the four together, Deep Seek quietly forgets, Claude gets scared and refuses, Gemini overthinks itself into silence, and GPT 5.5 finishes half of the job

14:00

and tells you that the rest of it is beneath it. And the point isn't which one of these is funniest, although it is genuinely a little funny. The point is that did it follow my instructions no longer has one failure mode. It has four different ways that it can fail, and you can't recognize that failure unless you know which model you're dealing with and what its pattern of failure is going to be. So the models got 10x better. They fail in funny ways. Why should you care when you get back to your desk? Because three things have changed to your workflow. The first is that a year ago, the smart move was to keep every skills file very, very short, under 200 instructions,

14:42

then point off to sub-skills, and a whole like, you know, Byzantine labyrinth of additional skills files and sub-files and things like that. And you are compressing your instructions to fit into a very small available space, and you don't need to do that anymore. Your skills files can be very long. Number two is that if your use case needs 100 specific rules or 300, you can just put them all in the prompt. You don't have to lie awake wondering which ones the model silently ignored. And if you've been thinking about your own lived experience of using models, you probably recognize this. You've discovered that you've got less worried

15:23

about how long your prompt is going to get because the models have genuinely got 10x better at following your prompts. 2,000 named constraints is an entire style guide, right? Like, it's every brand rule, every legal disclaimer. A year ago, you'd have had to shard that across a dozen specialized agents and hope that your specialized agents are handing off to each other cleanly, but now you can ignore that. But the third thing is the big one. The question used to be, can the model even do this? And the answer is now firmly yes. Well, reasonably firmly. Is it worth the cost is the new question because you can include 10,000 words of,

16:04

sorry, 10,000 different instructions into your prompt, but that is going to be an enormous prompt. It's going to be a very expensive prompt. It's going to be a very slow prompt. So what used to be a hard wall that you would run against has now become a soft trade-off of, is it worth me adding all of these extra instructions if it's going to give me more cost and more latency? And now some caveats to head off the Q&A. First and most important, I mentioned this earlier, this is a proxy task, including random words in a fake business report, is evidence that long skills file works. It is not the same as proof that a long skills file works.

16:42

Also, the models hit the wall at wildly different points, anywhere from 750 to 9,000 plus, so you have to pick your model very carefully. What our test doesn't do is measure whether the model reasoned clearly over a giant prompt. So the good news is since I did my research several weeks ago, a whole bunch of people have piled in on this, and now there's good research. Actual scientists have got involved and done, Chroma has done context rot work across 18 models, showing that accuracy on long inputs can fall 30 to 50 percent well before you hit the context window limit. And the weird part of their finding was that

17:25

coherent, well-structured text is more likely to hit that failure mode than if you just put your instructions into a random order and shuffle them in. I don't know why that's the case. I'll have to read their report. So the model can track 2,000, 5,000, possibly 10,000 instructions, but it's not necessarily going to reason clearly over them. It's not necessarily if those instructions conflict, if there is tension between them, it's not necessarily going to get that right. And then there's the other one I mentioned briefly. Claude's refusals are annoying, but they are loud. You get an error, you know it failed. GPT's polite half-finished report is much more dangerous

18:05

because it looks like a real answer. You have to read the whole thing to notice that it gave up quietly halfway, which means that you can't trust the output. It means you have to read the output every single time to make sure whether or not it's working. So the model will accept your 2,000 rules and it will hand you back something that looks, at least to begin with, confident and polished, but could be bailing out halfway through.

18:31

So as an aside, people always ask me, well, how much did all of this cost me? It cost me $209 to run all of these queries. 2,300 calls across seven models came to $209. It turns out novel research doesn't cost very much. And this is the part of the talk where I was saying that you have to check this stuff in production because you can't trust that your model isn't going to silently fail. So you knew I was going to mention evals eventually because I work at Arise and this is where I do that, but there are plenty of plugs for Arise. So I'm just going to say one true thing, which is that if you are building a real AI application and you are giving it genuinely tricky tasks,

19:10

you are going to run into one or more of these failure modes with a frontier model. And unless it's Claude telling you just to fuck off at the API level, the only way to know that something went wrong is monitoring your outputs with another LLM. That is an eval and that is what Arise does and I'll leave it at that. I already mentioned that there's been new research since we did our own. Here's another important one. A paper landed testing 46 models called Revisiting the Reliability of Language Models in Instruction Following, which you can bet made my ears perk up after I did that research myself. And they found something uncomfortable,

19:43

which is that a model can ace a benchmark like ours and still be wildly unreliable. Because if you reword the same instruction in a slightly different way, it can make a radical difference to how well it follows those instructions. So the model can follow 2,000 instructions and it can do it really well, but if you put the same instructions, the same 2,000 instructions in a different order, it can suddenly make the model much worse at following those instructions. And how exactly to do that, what is the correct order of instructions to give your model such that it follows them perfectly as opposed to getting confused, is still research that is being done.

20:20

So capacity went up, but reliability is still a problem. And then this is just a little brag, because I was happy about it, like I'm not a scientist, I did some research, and then a whole bunch of other actual scientists piled in and did real science on the same question. There's now a whole bunch of benchmarks that have shown up to measure this same question. Firebench, CCRbench, Guidebench, are all trying to measure the same thing. How well models follow a lot of real, messy constraints at once. And now the whole field is looking at it, so if you want better science than my 10,000 random words, the real science exists now. So that gets me to where I will leave you.

21:00

A year ago, the hard part of writing a skill was fitting everything in without the model losing the plot. That was a compression problem, and the compression problem is gone. The model will hold your 2,000 instructions just fine. The new hard part is knowing whether it actually did what you said, and that is a verification problem. A verification problem doesn't get solved by writing a better prompt. It gets solved by checking the output every time, the same way that you would test any other code, which is to say an eval. The ceiling moved by 10x in one year, so go back and check the assumptions that you made six months ago about how big your prompts should be,

21:35

how big your instructions can get, because they might already be wrong. So that is the talk. If you want all of the code and all of the data, it is at this GitHub URL. And this other QR code is something marketing made me insert. We are having a World Cup watch party tonight at 5 p.m. You can come to our party. That link is to the luma that will get you into the party. I hope this talk has given you some novel information, or at least a couple of laughs. And thank you so much for your time and attention.

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note