Beating RL With Reflection: GEPA and Optimize Anything — Lakshya A. Agrawal, GEPA
Description
After one round of reflection on just three examples, GEPA doubled the gains that the RL algorithm GRPO reached after 25,000 rollouts. Lakshya A. Agrawal, creator of GEPA and a PhD student at UC Berkeley's Sky Computing Lab, explains why. RL squeezes a whole rollout down to a single score. GEPA instead has a model read the full trace, including chains of thought, tool calls and error messages, and write a better prompt. A Pareto pool of candidates keeps the search out of local optima. The same idea works on anything you can write as text and score, which is where Optimize Anything comes in. Agrawal shows it taking a four-line program to a six-step agent that lifts Gemini Flash on ARC-AGI from 32.5% to 89.5%. It pushed an AMD NPU coding agent from 4% to 30%, and improved a GPT-5 mini coding agent on Go issues from 24% to 93%, using learned skills that also carried over to Claude Sonnet. He also covers Databricks tuning an open model to beat Claude Opus at 90x lower cost, learning evals from production traces, and co-optimizing prompts and weights. Speaker info: X/Twitter: @LakshyAAAgrawal (https://x.com/LakshyAAAgrawal) LinkedIn: https://www.linkedin.com/in/lakshyaaagrawal/ Website: https://lakshyaaagrawal.github.io/ Related links: GEPA on GitHub: https://github.com/gepa-ai/gepa Timestamps: 0:00 Intro: reflective optimization 0:30 How we teach AI new tasks 1:00 The sample-efficiency bottleneck 2:05 What RL throws away 3:00 Reflecting in text space 4:25 GEPA 4:50 GEPA vs. GRPO: 3 examples vs. 25,000 rollouts 5:40 What GEPA learns 6:55 GPT-4.1 mini beats GPT-4.1 7:15 A new AMD NPU: 4% to 30% 8:25 How GEPA works 9:00 Why a Pareto pool beats a simple loop 10:10 Results across benchmarks 10:35 Beyond prompts: Optimize Anything 12:05 The API 12:55 A 3D unicorn 13:35 Discovering agent harnesses: ARC-AGI from 32.5% to 89.5% 14:40 MATH-500 15:10 Optimizing agent skills: 24% to 93% 16:40 Three optimization modes 17:30 In production: OCR and a 90x cheaper agent 18:40 Why be
Summary
Generated by gpt-5.6-terraAt-a-Glance
- Verdict: Watch fully
- Core thesis: GEPA reframes AI-system improvement as sample-efficient evolutionary search in text space: use rich rollout feedback to reflectively optimize prompts, agent harnesses, skills, code, and other scoreable text artifacts rather than relying only on reward-driven weight updates.
- Why it matters: For agent systems with expensive rollouts, sparse proprietary data, and domain-specific tool traces, the approach offers a practical control-plane pattern for extracting more learning per execution without retraining a model.
- Best use: Use this as an architecture and experimentation reference for building self-improving agents: define a rigorous evaluator, preserve diagnostic side information, search a Pareto pool of candidate artifacts, and validate generalization before deployment.
Executive Summary
Lakshya A. Agrawal argues that conventional fine-tuning and reinforcement learning are often impractical for real-world agents because they require large datasets or hundreds of thousands of expensive rollouts. In an agent workflow, a binary reward discards the most useful information: reasoning traces, tool calls, tool outputs, compiler errors, retrieval results, and other diagnostics. GEPA instead lets an LLM inspect those artifacts, infer why a candidate succeeded or failed, and write a materially improved textual instruction or program.
The system is presented as an evolutionary optimization loop. It runs an existing pipeline on a small set of examples, passes the score plus actionable domain feedback to a reflector, generates revised text candidates, and retains a Pareto pool rather than only the globally best candidate. The stated reason is search diversity: candidates that are strong on different examples may enable later improvements, whereas repeatedly improving one incumbent tends to settle in local optima.
The broader OptimizeAnything abstraction is the more strategically relevant portion: if an artifact can be represented as text and evaluated, it can theoretically be optimized. That includes prompts, agent harness code, coding-agent skills, CUDA kernels, cloud scheduling policies, and even serialized numeric parameters. The important implementation principle is not the optimizer alone; it is exposing high-quality, domain-specific side information through the evaluator.
The speaker reports large benchmark and production-adjacent gains, including a four-line reasoning program evolved into a six-step agent, code-repository skills transferred from a cheap model to a stronger one, and lower-cost model deployments. These are compelling patterns but mostly speaker-reported results; the video gives limited detail on holdout design, optimization cost, safety constraints, or failure modes under reward hacking. The operational takeaway is to treat GEPA as a high-leverage evaluation-and-search layer, not as a substitute for robust evaluators and deployment controls.
Key Takeaways
- Claim: Reflective optimization can be substantially more sample-efficient than reinforcement learning because it learns from the full rollout trace rather than compressing every run into a scalar reward. | Evidence: The speaker contrasts RL with verified rewards, where GRPO converts terminal rewards into gradients, against GEPA, which reads chains of thought, tool calls, environment responses, and errors. He reports that one GEPA reflection round using three data points produced roughly twice the performance gain of GRPO after 25,000 rollouts on a cited comparison. | Implication: For long-running or tool-heavy agents, collect and structure execution diagnostics before considering expensive online RL; a small number of richly annotated failures may be more useful than a large number of binary-scored rollouts. | Caveat: The comparison is presented without task setup, confidence intervals, compute accounting, or independent replication in the transcript, so the magnitude should be treated as a directional claim rather than a planning estimate.
- Claim: Natural-language prompt updates can cause large behavioral changes quickly, making prompt space a useful optimization surface alongside model weights. | Evidence: The talk uses the example of changing 'generate a one line summary' to 'generate a ten line summary' and argues that a small textual edit can produce a large policy change that would require many small gradient updates in weight space. GEPA-generated prompts are described as problem specifications containing task context, input interpretation, and lessons from observed cases. | Implication: Treat prompts as first-class, versioned operational artifacts with automated evals, rather than manually maintained strings tuned by intuition. | Caveat: Prompt changes can also introduce regressions, brittleness, or overfitting to the examples used during search; the transcript does not describe a formal regression-testing or prompt-versioning process.
- Claim: Maintaining a Pareto pool of candidates is central to GEPA's search advantage because it preserves diverse partial solutions instead of repeatedly improving only one top-scoring prompt. | Evidence: GEPA retains every candidate that wins on at least one training example. The speaker says a single incumbent loop gets trapped at a local optimum, while Pareto selection explores a more balanced search tree; across four benchmarks, he reports that this design accounted for more than half of GEPA's gains and yielded nearly twice the gain of a simple LLM-in-a-loop baseline. | Implication: When optimizing agents across heterogeneous tasks, retain non-dominated candidates and test them across a representative suite instead of selecting solely by aggregate score after each iteration. | Caveat: A Pareto pool increases evaluation and candidate-management complexity, and its benefit depends on having multiple cases whose performance tradeoffs are meaningful.
- Claim: The same reflective loop can optimize an entire agent harness or other executable text, allowing architecture discovery rather than merely prompt improvement. | Evidence: For an RKGI reasoning task, the speaker says OptimizeAnything started from a four-line Python program that directly invoked a model and, after 16 reflection rounds, discovered a six-step agent that performed rule-hypothesis induction, code synthesis, execution tracing, debugging, and final answer generation. Reported Gemini Flash accuracy rose from 32.5% to 89.5%. On MAT500, GPT-4.1 nano reportedly gained 20% accuracy through an automatically created two-step agent. | Implication: Agent orchestration should be considered a searchable design space. Start with constrained components and explicit tool policies, then optimize against quality, latency, cost, and safety criteria jointly. | Caveat: Automatically generated harnesses require security review, execution sandboxing, tool-permission controls, cost limits, and held-out evaluation; optimizing for a benchmark can produce unsafe or non-generalizable control flow.
- Claim: Actionable side information is the key interface contract for universal optimization: the evaluator should return not only a score but every useful diagnostic artifact. | Evidence: OptimizeAnything accepts a fitness/evaluator function over a set of problems and an open-ended dictionary of feedback. Examples include expert comments, compiler and profiler output, tool-call errors, documentation, job traces, and SLA violations. The speaker applies this pattern to CUDA code, cloud scheduling policies, numerical optimization, and agent harnesses. | Implication: Design evaluators as observability products: emit structured, redacted, task-relevant diagnostics and make the optimization objective explicitly multi-dimensional where cost, latency, and safety matter. | Caveat: More feedback is not automatically better: sensitive production traces may contain credentials, customer data, or proprietary logic, and unfiltered feedback can expand the attack surface of the optimizer.
- Claim: Cheaply learned repository-specific skills can transfer to stronger coding models and improve both issue-resolution accuracy and execution cost. | Evidence: Using a MiniSWE-style agent with GPT-5 Mini, the speaker says skill optimization increased Go repository issue-resolution performance from 24% to 93%. He then reports applying the resulting skills to Claude Sonnet 4.5, reaching 100% issue resolution on the cited evaluation and cutting resolution time by nearly 50%. The skills encode repository layout, test invocation, feature locations, and build-system knowledge. | Implication: Build skills from successful and failed coding trajectories, but attach them to repository revisions and automatically revalidate them after meaningful codebase or tooling changes. | Caveat: The reported 100% result is task-suite-specific and should not be interpreted as universal repository-solving performance. Repository skills can also stale quickly as codebases, tests, and build systems change.
- Claim: Subjective agent tasks can be brought into the optimization loop by first learning an LLM-as-judge from a small set of richly annotated production traces. | Evidence: The proposed process is to collect production trajectories, have humans provide detailed feedback on roughly 50 of them, optimize a judge prompt using those annotations, then use that judge to optimize the production agent. The speaker describes this as a continuing data flywheel. | Implication: For quality dimensions without deterministic ground truth, invest in a calibrated judge-evaluation layer before attempting self-improvement; do not optimize directly against vague or unvalidated proxy metrics. | Caveat: A learned judge can encode annotator preferences, be gamed by the optimized agent, and amplify mistaken criteria; periodic human calibration and adversarial checks remain necessary.
Detailed Brief
Reported applications and strategic claims
- Claims: GEPA is presented as model- and framework-agnostic, with no hard dependencies, so it can sit outside an existing agent framework rather than requiring a model-training stack.; The speaker argues that stronger models increase, rather than reduce, the value of prompt optimization because capable models follow precise domain instructions more effectively.; The approach is described as extending beyond text-only agents to vision-language tasks and operational-policy optimization.
- Evidence: On AMD's NPU-XDNA2 programming task, where online information was scarce and GPT-4o reportedly struggled, the speaker says applying GEPA to an existing agent raised performance from 4.25% to 30.52%. A discovered instruction reportedly warned against using AMD's ADF.Edge library for the targeted hardware generation.; Reported downstream results include optimizing GPT-4.1 Mini to exceed GPT-4.1 on a math task, cutting OCR errors by nearly 35% in an externally validated report, and reducing cloud-scheduling cost by almost 40% versus expert heuristics.; The speaker cites Databricks as having tuned GPT-OSS 120B to outperform Claude Opus in a deployed agent while being 90 times cheaper, and refers to production use by companies including Dropbox and Shopify.
- Caveats: Most application results are asserted in a conference-style presentation, not substantiated here with datasets, baselines, reproducible configurations, or cost of the optimization run.; Claims about named organizations and external deployments should be independently checked before they inform vendor selection, investment conclusions, or performance forecasts.
- Implications: The highest-value opportunities are likely domains where error messages, traces, expert notes, and measurable operational outcomes already exist but labeled training data does not.; A capable evaluator and clean feedback pipeline may become a more durable advantage than hand-authored prompts or a one-time choice of foundation model.
Optimization modes and longer-term model-learning direction
- Claims: OptimizeAnything offers a single-problem mode, a multitask search mode for related tasks that may share information, and a skill-building/generalization mode for deployment on unseen cases.; The speaker briefly introduces 'fast-slow learning,' a proposed scheme to co-optimize model weights and prompt harnesses for continual learning.
- Evidence: Examples used to distinguish modes include optimizing one matrix-multiplication kernel, jointly searching for matrix multiplication and dot-product kernels, and learning a reusable math prompt or agent architecture from a training set for future unseen queries.; The cited paper is titled 'Learning Fast and Slow'; no technical details or empirical results for weight-and-harness co-optimization are provided in the transcript.
- Caveats: The continual-learning proposal cannot be evaluated from this talk alone because its methodology, stability properties, retention behavior, and safety controls are not explained.
- Implications: Separate immediate operational use of text-space optimization from longer-horizon research into joint prompt, harness, and weight adaptation; they have different infrastructure, risk, and validation requirements.
Notable Concepts & Terms
- GEPA: The reflective optimization method discussed in the talk; it uses LLM-generated textual reflections and evolutionary search to improve prompts and other scoreable text artifacts. The transcript repeatedly pronounces it as 'JEPA,' while the title names it GEPA.
- Reflective optimization in text space: Replacing or supplementing weight updates with natural-language or code edits derived from analysis of execution feedback.
- Actionable side information: Diagnostic context returned by an evaluator beyond a scalar score, such as tool traces, compiler errors, profiler results, documentation, expert feedback, or SLA violations.
- Pareto pool: A retained set of candidates that are strong on at least one example or objective, intended to preserve search diversity and avoid local optima.
- OptimizeAnything: The proposed general API that optimizes any candidate expressible as text, provided an evaluator can score it and emit useful feedback.
- Agent harness optimization: Searching over the agent's executable orchestration logic, including tool descriptions, control flow, code execution, and multi-step procedures, rather than only editing prompts.
- G-Skill: An open-source feature described as learning reusable coding-agent skills from repository trajectories, then applying them to similar future issues.
- LLM-as-judge flywheel: Using a small human-annotated trace set to optimize a judge prompt, then using that judge to score and iteratively improve the production agent.
Operator Notes / Why Ken Should Care
- Select one current agent workflow with expensive rollouts and an existing measurable outcome; build a small, representative optimization suite with a genuinely held-out regression set before testing reflective search.
- Modify the workflow evaluator to emit structured and redacted diagnostics alongside scores: tool inputs/outputs, failures, retries, latency, token cost, retrieval evidence, and policy violations.
- Run a constrained pilot that compares manual prompt iteration, best-candidate-only reflection, and Pareto-pool search under the same rollout and model budget.
- For any harness or code optimization experiment, isolate execution in a sandbox, constrain tool permissions and network access, cap spend and iteration count, and require human approval before production promotion.
- Use multi-objective acceptance gates that include task quality, latency, token cost, security compliance, and worst-case regression; do not promote a candidate solely because it improves mean benchmark score.
- If deploying learned skills or judge prompts, version them against repository/schema/policy revisions and schedule recalibration with fresh human annotations to detect staleness or reward hacking.
Source/Metadata
- Title: Beating RL With Reflection: GEPA and Optimize Anything — Lakshya A. Agrawal, GEPA
- Transcript words: 5658
- Duration seconds: 1287
- Timestamp note: No timestamps or chapters were provided. The transcript includes substantial repeated material in its latter half.
Transcript
Lakshya A. Agrawal The problem of reflective optimization or how can we self-improve prompts, agents, and models from textual feedback. The question we start with is how can we improve prompts, agents, and models from textual feedback. The first step is how can we teach AI to perform new tasks. The standard way has been to perform weight updates with gradient descent either during pre-training, supervised fine-tuning, or reinforcement learning. This has proven to be extremely effective, but it requires a huge number of examples. Trillions of tokens for pre-training, tens of thousands of labeled examples for supervised fine-tuning, or hundreds of thousands of rollouts for reinforcement learning in domains like math, coding, etc. However, most teams do not actually have that much data or compute, and the problems that we are trying to tackle with AI now are bottlenecked by sample efficiency. What do we mean by that? Two things. First of all, there is low availability of domain-specific knowledge resources, which means there is not enough data to perform offline algorithms like SFT. Second, the domains that we are trying to apply to AI increasingly are having expensive rollouts, where either the LLM workflow pipeline or agentic rollouts are itself very slow or expensive to do, or the task metric is very slow or expensive to execute. We are seeing that agents can now work for hours on end, and if you were to apply an online learning algorithm to this, it would require hundreds of thousands of rollouts, and it would not be feasible. So we are seeing increasing use of agents for real-world applications where these invoke tools, which can also be long-running, further exacerbating the sample inefficiency issue. The current dominant paradigm is reinforcement learning with verified rewards, where given a model and a task, we perform a number of parallel rollouts and get rewards at the end. Finally, an algorithm like GRPO takes these rewards and converts it into gradients that are applied back to the model. However, as we can see, there was a lot of information in each of these rollouts, but we only learned a zero or one score and propagated that via gradient descent. We can see that there is chains of thought, the tool calls made to the environment, the environment's responses to those tool calls, which could potentially contain error messages, which also provide diagnostic value, and we learned almost nothing from all of that. So the question we ask is, can we make use of this other extremely rich information? Our idea is to perform reflective optimization in text space, where instead of only using the zero or one reward signal, we can have a language model or an agent look at the trace of the entire rollout and reflect on what worked in them, what did not work in them. And this reflection could potentially use all intermediate outputs and potentially even make other tool calls, such as retrieval from your company's knowledge base or some guide textbook and so on. So that's the first key idea. And the second is that instead of only updating weights with small deltas, we can instead update a prompt, where a single natural language update can give a very large behavior change. Let's take a simple example. Let's say you're tasked with writing a text summarization system. And the prompt of that system says, generate a one line summary. If I just go and tweak that prompt to say generate a ten line summary, we can all agree that the behavior of the system would change quite significantly with that just one word change. And making that one word change is quite quick, and we can reflect on our own behavior and identify what needs to change. If we were to achieve a similar kind of behavior update from our AI system, we would have to have thousands of gradient, very tiny gradient updates sequentially. So with that key idea, we propose JEPA, which is a reflective prompt optimization technique for agents. It uses an evolutionary loop along with a novel Pareto-based candidate selection, which I will come to later. It is akin to doing reinforcement learning in text space, where instead of just receiving a reward score, we are actually obtaining score along with textual feedback, which can be very domain specific and learn all about the domain from it. Let's compare JEPA with GRPO, which is one of the leading RL techniques. On the x-axis, we have the number of training steps, also proportional to number of data samples seen. And on the y-axis, we have the performance on our domain that we are training for. And what we can see is that JEPA, in just one round of reflection using just three data points, is already able to get twice the performance gains that GRPO got after 25,000 rollouts. Continuing to run JEPA for a few more steps further increases that gap itself by another 2x. I want to note here that the model Qin3 8b is optimizing itself here. There is no external expert teacher involved whatsoever. And what does JEPA learn? Unlike prior prompt optimizers, someone which would use model idiosyncrasies, like my grandmother will be really angry if you don't generate a good prompt. Here, JEPA is actually giving a very detailed problem specification, which includes how to make sense of the input, what is the purpose and context of this particular part of the pipeline, what are some key observations and lessons from the data. So the prompt we are seeing here is for the second hop of a multi-hop question answering system, where given a question, we need to retrieve some documents that could potentially answer that question, look at those documents, summarize it, and then finally answer the question. And here what we see is JEPA has found out that first hop documents often cover one entity or aspect, and the second hop should actually be recovering documents that are related to it. We have seen that human engineering teams, whenever a new model comes out, spend weeks of their time manually tweaking one word here and there, trying to discover the problem specification. This entire process is fully automated now with JEPA, which takes about half an hour to one hour to run depending on your pipelines. We can also apply JEPA to leading proprietary models. Just for an example, here we were able to optimize GPT 4.1 Mini's performance to outperform GPT 4.1 on a math task. And we can see the kind of information distillation JEPA has done in the prompt space itself. Coming back to the problem of sample efficiency, AMD developed a new hardware accelerator called NPU-XDNA2, which used a completely new API to program, which had almost zero available information on the internet. And because of this, the leading models at the time, which was GPT 4.0, was failing miserably to perform this task. We were able to take an existing agent, which was getting 4.25% on this task, and applied JEPA without any other change to the agent itself. And we got this prompt and pushed this performance 7x to 30.52%. So what this goes to say is there can be lots of domain-specific information, which if you include in your AI systems prompts, the models could actually perform much better, and JEPA can help you fully automatically discover that. I want to highlight the sentence saying avoid including ADF.Edge. Now the interesting thing is AMD actually ships a library called ADF.Edge for programming NPUs, but that did not work with this latest generation of hardware that we were working with, and JEPA was able to discover that in just one step. So how does it work? It's an extremely simple algorithm, which simply takes your AI pipeline written in any agentic framework, or even raw LLM calls that you may have. It simply runs your systems on a few examples and collects domain-specific feedback. Whatever information your environment contains is observed. Second, it runs reflection with an LLM or agent that reads the feedback and proposes a better prompt. Finally, and most importantly, it keeps a Pareto pool where it keeps every single candidate that wins on even one training example and not just the top scorer. The question is, but why keep a Pareto pool? And we kept getting asked this question a lot, that is JEPA really better than running the model in a loop? So we went and tested it out, and what happens is a loop keeps only the best and gets stuck in a local optima. So on the left-hand side, you see a search tree that was generated by using an LLM in a loop. Starting from a seed prompt at the top left, we asked the LLM to improve the prompt. It improved the prompt and it generated a prompt that gave us the middle node. However, this prompt got stuck in a local optima, and once again when we asked the LLM to try and improve it, it proposed something but that was not actually better. So it went back and it again tried to improve it. And it kept doing this and it exhausted all of the search budget. On the other hand, with JEPA's Pareto-based candidate selection strategy on the right, we can see that it maintains a much more balanced search process, eventually converging to a much higher score. Across four benchmarks, we saw that more than half of the gains seen with JEPA actually account for this, and it gets almost twice the performance gains that you would get with just applying the model in a loop. JEPA can perform really well across diverse benchmarks. Here we see results on question answering, instruction following, claim verification, as well as math, which all the leading frontier model companies are already optimizing their models a lot for. However, this prompt got stuck in a local optima, and once again when we asked the LLM to try and improve it, it proposed something but that was not actually better. So it went back and it again tried to improve it. And it kept doing this and it exhausted all of the search budget. On the other hand, with JEPA's Pareto-based candidate selection strategy on the right, we can see that it maintains a much more balanced search process, eventually converging to a much higher score. Across four benchmarks, we saw that more than half of the gains seen with JEPA actually account for this, and it gets almost twice the performance gains that you would get with just applying the model in a loop. JEPA can perform really well across diverse benchmarks. Here we see results on question answering, instruction following, claim verification, as well as math, which all the leading frontier model companies are already optimizing their models a lot for, and we are still able to get plus 10% just by optimizing the prompt on it. So we have so far seen JEPA only optimizing the prompts, but JEPA goes far beyond prompts. And because prompts are just text artifacts that determine AI system behavior, the same algorithm can improve anything that you can express as a piece of text and you can score. For example, your entire agent harness is eventually just a Python or a JavaScript file, and we can apply the same kind of reflective optimization process to that entire file and we can work with it. So if you can write it as text and score it, JEPA can optimize it. So with that insight in mind, we propose OptimizeAnything, which is a universal API for optimizing any text parameter, given any domain like code optimization where, let's say, you want to optimize the CUDA kernel code. The input is just that CUDA kernel code where an evaluator looks at this piece of code, maybe compiles it, profiles it, generates a bunch of related information that we call as actionable side information, which is then provided to an LLM, which proposes a better candidate maintaining this Pareto pool, and it keeps repeating this process till we get convergence. The same thing can be applied to numeric optimization, where your numbers can actually be serialized as text, or harness optimization where an entire harness can be serialized as text, or even cloud scheduling policy optimization where the scheduling policy or heuristic algorithm can be expressed as a piece of text, and the evaluator can be something like the negative of cost or some function measuring accuracy, efficiency, and the actionable side information can be something like job traces, SLA violations, and so on. The API is dead simple to use. All it requires is you give us the set of problems that you care to be solved, along with an evaluator function or a fitness function that returns a score along with any available domain-specific side information. If your domain produces expert feedback, return that. If your domain produces compiler error messages, profiler messages, tool call error messages, return that. If you have maybe a written up documentation, return that. Any kind of—it's a very open-ended dictionary—you can return literally anything, and all you do is you call optimize anything with this fitness function and the set of problems that you have, and optimize anything will take care of it and give you an optimized solution. Let's see some application. Let's say you were tasked with generating a 3D unicorn. This is all the code that you would write or your agent can now write it because we have seen that optimize anything is a very easy to use API for leading agents like CloudCode. So all you do is write this code which says optimize a Python program to generate a 3D unicorn, and the candidate is a Python script that produces a PNG rendering, whatever. And here is the result. On the left-hand side, we can see Cloud Opus 4.6. If you gave it this task, this is what it generated. And on the right-hand side, the unicorn that we get with optimize anything. This is just for fun, but let's say you were tasked with writing an agent to solve a specific task. Typically, teams spend lots and lots of time tweaking their agents, building tools for it, writing tool descriptions, carefully orchestrating the control flow, and so on. Here we started with a simple four-line Python program that was simply calling a model's chain of thought to solve an RKGI problem. Within just 16 rounds of reflection, JEPA within OptimizeAnything was able to find this sophisticated six-step agent that took RKGI accuracy of Gemini Flash from 32.5% to 89.5%. And we can see that this agent is automatically, by itself, doing rule hypothesis induction, code synthesis. It executes and traces the code, automatically debugs this code, goes back and proposes new versions of that code, and finally it runs it on the actual test inputs and returns the output. This is a runnable example. You can go to this QR code and you can run this example right now. So applying the same approach of discovering agent harnesses to MAT500, we are able to push its accuracy of GPT 4.1 nano by 20%, by simply creating a two-step agent. And again, I want to emphasize that all we did is we asked OptimizeAnything to optimize an agent file, and it was automatically discovering the sophisticated agent architecture, and we did not have to do anything other than specifying the objective and the task. Finally, every single one of us is using some coding agent like Cloud Code or Codex or maybe your favorite agent, and agent skills has become a very leading part of the ecosystem where almost all coding agents understand skills. Let's say you want to optimize skills for your specific repository. This is the code that you write, which says, learn a skill from the trajectory. When the coding agent is presented with a similar problem, the skill should be helpful. We just give it this natural language behavior, and what we see is we started with MiniSui agent with GPT 5 Mini, because we were very budget constrained, and we were able to take its performance from 24% to 93%, almost a 3x jump on Go repository issue resolution, but more importantly, the skills that were optimized very cheaply on a GPT 5 Mini agent, we were able to take that and apply it to the latest Cloud Sonnet. This was done about a few months back, but we applied it to Cloud Sonnet 4.5, pushing its accuracy to 100% issue resolution, while more importantly, cutting down the execution time or issue resolution time by almost 50%. We cut it down in half, which also means it spent less tokens, because skills contain information about how the repository is organized, how to invoke the test cases, where a particular feature is implemented, what are the build systems used by this repository, and so on. This is a feature called G-Skill. You can find it in the JEPA repository, and it's fully open source as well. So OptimizeAnything is a single interface that provides three optimization modes. If you have just a single problem, like there is a single matrix multiplication kernel that you want to optimize, you can use it that way. If you have n number of related problems, like you want to optimize a matrix multiplication kernel, along with a dot product kernel, and you know there might be some information transfer between these two, you can use what we call as the multitask search mode. And finally, build a skill, which is, if you want to optimize on a set number of problems, but your deployment can actually come up with many new problems. So in case of math prompt optimization, we are training on some examples, but when we deploy it, we can receive a completely new kind of query. So we care about generalization modes. So there you can do prompt optimization, agent architecture optimization, and so on. So OptimizeAnything can be used for a broad set of domains, including cloud scheduling policy optimization, where we are able to cut costs by almost 40% compared to expert heuristics, write custom solvers to match and exceed Optina even in black box mathematical optimization, create agent skills, prompt optimization, and so on. It is so easy to use that within just 20 hours of releasing it, people at Snorkel had already improved some of their internal benchmarks with it and were tweeting about it. So JEPA also improves multimodal VLM models performance. Here, we are able to cut OCR error rates for leading models by almost 35%, and this is an externally validated report. Similarly, Databricks actually achieved 90x cost reduction in their deployed agent's performance, and here they were able to tune GPT OSS 120B to outperform Claude Opus while being 90x cheaper. More importantly, the performance delta improvement that you see on top of Claude Opus is actually bigger than the one you see on open source models. Some people have asked me that, oh, as models get better, the importance of prompt optimization will go down. I argue the opposite, which is, as models get better, they will get better at instruction following, and the more precise instruction about your task that you have to give to a very smart model, the better that model will be at solving your task. And this is exactly what we see happening here. The better the instruction was, Claude Opus actually jumped much higher. Some people have this question of what if we have subjective tasks which are very hard to evaluate. JEPA can actually learn evals for your task from production traces. and here they were able to tune GPT OSS 120B to outperform Claude Opus while being 90x cheaper. More importantly, the performance delta improvement that you see on top of Claude Opus is actually bigger than the one you see on open source models. Some people have asked me, as models get better, the importance of prompt optimization will go down. I argue the opposite, which is, as models get better, they will get better at instruction following, and the more precise instruction about your task that you have to give to a very smart model, the better that model will be at solving your task. And this is exactly what we see happening here. The better the instruction was, Claude Opus actually jumped much higher. Some people have this question of what if we have subjective tasks which are very hard to evaluate. JEPA can actually learn evals for your task from production traces. The way to do that is you collect a bunch of production traces from your agent, get a human to annotate just about 50 of those trajectories, giving very detailed feedback. This is a long response, this is a short response, this is a good response, this uses this terminology, whatever. And once you get those human annotations, you can use JEPA to optimize an LLM as a judge prompt, and you can use that LLM as a judge prompt then to go back and optimize your agent, and deploy that agent, and this becomes a data flywheel where you can keep improving it, and this is a successful paradigm that some leading teams in production are already using. Then the question we get asked is, can we actually use this reflective optimization to train models? And we recently had this paper called Learning Fast and Slow where we propose fast slow learning, where we can co-optimize model weights and prompt harnesses, and this shows some very strong properties that one would want in a continual learning algorithm. I don't have much time to go over details, but please look at the papers. And since release, JEPA has been used in production by these companies as well as the main methodology in these papers. And here the CEO of Dropbox and Shopify are talking about their use of JEPA, and OpenAI also wrote a blog post about how you can build self-improving AI systems with JEPA. So it's very simple to get started. It can plug into any framework, any model, and it has absolutely zero hard dependencies, so you can deploy it in any kind of setting. So don't be afraid to optimize in the tech space, and many problems can be framed as optimization, so bring actionable site information and surface as much domain-specific information as you can to optimizers, and the optimizers of future will be able to work with them. So please go and check it out. Thank you very much. For example, your entire agent harness is eventually just a Python or a JavaScript file, and we can apply the same kind of reflective optimization process to that entire file and we can work with it. So if you can write it as text and score it, JEPA can optimize it. So with that insight in mind, we propose OptimizeAnything, which is a universal API for optimizing any text parameter, given any domain like code optimization where, let's say, you want to optimize the CUDA kernel code. The input is just that CUDA kernel code where an evaluator looks at this piece of code, maybe compiles it, profiles it, generates a bunch of related information that we call as actionable side information, which is then provided to an LLM, which proposes a better candidate maintaining this Pareto pool, and it keeps repeating this process till we get convergence. The same thing can be applied to numeric optimization, where your numbers can actually be serialized as text, or harness optimization where an entire harness can be serialized as text, or even cloud scheduling policy optimization where the scheduling policy or heuristic algorithm can be expressed as a piece of text, and the evaluator can be something like the negative of cost or some function measuring accuracy, efficiency, and the actionable side information can be something like job traces, SLA violations, and so on. The API is dead simple to use. All it requires is you give us the set of problems that you care to be solved, along with an evaluator function or a fitness function that returns a score along with any available domain-specific side information. If your domain produces expert feedback, return that. If your domain produces compiler error messages, profiler messages, tool call error messages, return that. If you have maybe a written up documentation, return that. Any kind of, it's a very open-ended dictionary, you can return literally anything, and all you do is you call optimize anything with this fitness function and the set of problems that you have, and optimize anything will sort of take care of it and give you an optimized solution. Let's see some application. Let's say you were tasked with generating a 3D unicorn. This is all the code that you would write or your agent can now write it because we have seen that optimize anything is a very easy to use API for leading agents like CloudCode. So all you do is write this code which says optimize a Python program to generate a 3D unicorn, and the candidate is a Python script that produces a PNG rendering, whatever. And here is the result. On the left-hand side, we can see Cloud Opus 4.6. If you gave it this task, this is what it generated. And on the right-hand side, the unicorn that we get with optimize anything. This is just for fun, but let's say you were tasked with writing an agent to solve a specific task. Typically, teams spend lots and lots of time tweaking their agents, building tools for it, writing tool descriptions, carefully orchestrating the control flow, and so on. Here we started with a simple four-line Python program that was simply calling a model's chain of thought to solve an RKGI problem. Within just 16 rounds of reflection, Jepa within OptimizeAnything was able to find this sophisticated six-step agent that took RKGI accuracy of Gemini Flash from 32.5% to 89.5%. And we can see that this agent is automatically, by itself, doing rule hypothesis induction, code synthesis. It executes and traces the code, automatically debugs this code, goes back and proposes new versions of that code, and finally it runs it on the actual test inputs and returns the output. This is a runnable example. You can go to this QR code and you can run this example right now. So applying the same approach of discovering agent harnesses to MAT500, we are able to push its accuracy of GPT 4.1 nano by 20%, by simply creating a two-step agent. And again, I want to emphasize that all we did is we asked OptimizeAnything to optimize an agent file, and it was automatically discovering the sophisticated agent architecture, and we did not have to do anything other than specifying the objective and the task. Finally, every single one of us is using some coding agent like Cloud Code or Codex or maybe your favorite agent, and agent skills has become a very leading part of the ecosystem where almost all coding agents understand skills. Let's say you want to optimize skills for your specific repository. This is the code that you write, which says, learn a skill from the trajectory. When the coding agent is presented with a similar problem, the skill should be helpful. We just give it this natural language behavior, and what we see is we started with MiniSui agent with GPT 5 Mini, because we were very budget constrained, and we were able to take its performance from 24% to 93%, almost 3x jump on Go repository issue resolution, but more importantly, the skills that were optimized very cheaply on a GPT 5 Mini agent, we were able to take that and apply it to the latest Cloud Sonnet. This was done about a few months back, but we applied it to Cloud Sonnet 4.5, pushing its accuracy to 100% issue resolution, while more importantly, cutting down the execution time or issue resolution time by almost 50%. We cut it down into half, which also means it spent less tokens, because skills contain information about how the repository is organized, how to invoke the test cases, where a particular feature is implemented, what are the build systems used by this repository, and so on. This is a feature called G-Skill. You can find it in the JEPA repository, and it's fully open source as well. So, Optimize Anything is a single interface that provides three optimization modes. If you have just a single problem, like there is a single matrix multiplication kernel that you want to optimize, you can use it that way. If you have n number of related problems, like you want to optimize a matrix multiplication kernel, along with a dot product kernel, and you know there might be some information transfer between these two, you can use what we call as the multitask search mode. And finally, build a skill, which is, if you want to optimize on a set number of problems, but your deployment can actually come up with many new problems. So, like, in case of math prompt optimization, we are training on some examples, but when we deploy it, we can receive a completely new kind of query. So, we care about generalization modes. So, there you can do prompt optimization, agent architecture optimization, and so on. So, Optimize Anything can be used for a broad set of domains, including cloud scheduling policy optimization, where we are able to cut costs by almost 40% compared to expert heuristics, write custom solvers to match and exceed Optina even in black box mathematical optimization, create agent skills, prompt optimization, and so on. It is so easy to use that within just 20 hours of releasing it, people at Snorkel had already improved some of their internal benchmarks with it and were tweeting about it. So, and JEPA also improves multimodal VLM models performance. Here, we are able to cut OCR error rates for leading models by almost 35%, and this is an externally validated report. Similarly, Databricks actually achieved 90x cost reduction in their deployed agent's performance, and here they were able to tune GPT OSS 120B to outperform Claude Opus while being 90x cheaper. More importantly, the performance delta improvement that you see on top of Claude Opus is actually bigger than the one you see on open source models. Some people have asked me that, oh, as models get better, the importance of prompt optimization will go down. I argue the opposite, which is, as models get better, they will get better at instruction following, and the more precise instruction about your task that you have to give to a very smart model, the better that model will be at solving your task. And this is exactly what we see happening here. The better the instruction was, Claude Opus actually jumped much higher. Some people have this question of what if we have subjective tasks which are very hard to evaluate. JEPA can actually learn evals for your task from production traces. The way to do that is you collect a bunch of production traces from your agent, get a human to annotate just about 50 of those trajectories, giving very detailed feedback. This is a long response, this is a short response, this is a good response, this uses this terminology, whatever. And once you get those human annotations, you can use JEPA to optimize an LLM as a judge prompt, and you can use that LLM as a judge prompt then to go back and optimize your agent, and deploy that agent, and this becomes a data flywheel where you can keep improving it, and this is a successful paradigm that some leading teams in production are already using. Then the question we get asked is, can we actually use this reflective optimization to train models? And we recently had this paper called Learning Fast and Slow where we propose fast, slow learning, where we can co-optimize model weights and prompt harnesses, and this shows some very strong properties that one would want in a continual learning algorithm. I don't have much time to go over details, but please look at the papers. And since release, JEPA has been used in production by these companies as well as the main methodology in these papers. And here the CEO of Dropbox and Shopify are talking about their use of JEPA, and OpenAI also wrote a blog post about how you can build self-improving AI systems with JEPA. So it's very simple to get started. It can plug into any framework, any model, and it has absolutely zero hard dependencies, so you can deploy it in any kind of setting. So don't be afraid to optimize in the tech space, and many problems can be framed as optimization, so bring actionable site information and surface as much domain-specific information as you can to optimizers, and the optimizers of future will be able to work with them. So please go and check it out. Thank you very much. Thank you very much. Thank you very much. Thank you very much. Thank you very much. Thank you very much. Thank you very much. Thank you very much. Thank you very much. Thank you very much. you