Lakshya A. Agrawal The problem of reflective optimization or how can we self-improve prompts, agents, and models from textual feedback. The question we start with is how can we improve prompts, agents, and models from textual feedback. The first step is how can we teach AI to perform new tasks. The standard way has been to perform weight updates with gradient descent either during pre-training, supervised fine-tuning, or reinforcement learning.
This has proven to be extremely effective, but it requires a huge number of examples. Trillions of tokens for pre-training, tens of thousands of labeled examples for supervised fine-tuning, or hundreds of thousands of rollouts for reinforcement learning in domains like math, coding, etc. However, most teams do not actually have that much data or compute, and the problems that we are trying to tackle with AI now are bottlenecked by sample efficiency. What do we mean by that? Two things. First of all, there is low availability of domain-specific knowledge resources, which means there is not enough data to perform offline algorithms like SFT.
Second, the domains that we are trying to apply to AI increasingly are having expensive rollouts, where either the LLM workflow pipeline or agentic rollouts are itself very slow or expensive to do, or the task metric is very slow or expensive to execute. We are seeing that agents can now work for hours on end, and if you were to apply an online learning algorithm to this, it would require hundreds of thousands of rollouts, and it would not be feasible. So we are seeing increasing use of agents for real-world applications where these invoke tools, which can also be long-running, further exacerbating the sample inefficiency issue.
The current dominant paradigm is reinforcement learning with verified rewards, where given a model and a task, we perform a number of parallel rollouts and get rewards at the end. Finally, an algorithm like GRPO takes these rewards and converts it into gradients that are applied back to the model. However, as we can see, there was a lot of information in each of these rollouts, but we only learned a zero or one score and propagated that via gradient descent.
We can see that there is chains of thought, the tool calls made to the environment, the environment's responses to those tool calls, which could potentially contain error messages, which also provide diagnostic value, and we learned almost nothing from all of that. So the question we ask is, can we make use of this other extremely rich information? Our idea is to perform reflective optimization in text space, where instead of only using the zero or one reward signal, we can have a language model or an agent look at the trace of the entire rollout and reflect on what worked in them, what did not work in them.
And this reflection could potentially use all intermediate outputs and potentially even make other tool calls, such as retrieval from your company's knowledge base or some guide textbook and so on. So that's the first key idea. And the second is that instead of only updating weights with small deltas, we can instead update a prompt, where a single natural language update can give a very large behavior change. Let's take a simple example. Let's say you're tasked with writing a text summarization system. And the prompt of that system says, generate a one line summary.
If I just go and tweak that prompt to say generate a ten line summary, we can all agree that the behavior of the system would change quite significantly with that just one word change. And making that one word change is quite quick, and we can reflect on our own behavior and identify what needs to change. If we were to achieve a similar kind of behavior update from our AI system, we would have to have thousands of gradient, very tiny gradient updates sequentially. So with that key idea, we propose JEPA, which is a reflective prompt optimization technique for agents.
It uses an evolutionary loop along with a novel Pareto-based candidate selection, which I will come to later. It is akin to doing reinforcement learning in text space, where instead of just receiving a reward score, we are actually obtaining score along with textual feedback, which can be very domain specific and learn all about the domain from it. Let's compare JEPA with GRPO, which is one of the leading RL techniques. On the x-axis, we have the number of training steps, also proportional to number of data samples seen. And on the y-axis, we have the performance on our domain that we are training for.
And what we can see is that JEPA, in just one round of reflection using just three data points, is already able to get twice the performance gains that GRPO got after 25,000 rollouts. Continuing to run JEPA for a few more steps further increases that gap itself by another 2x. I want to note here that the model Qin3 8b is optimizing itself here. There is no external expert teacher involved whatsoever. And what does JEPA learn? Unlike prior prompt optimizers, someone which would use model idiosyncrasies, like my grandmother will be really angry if you don't generate a good prompt.
Here, JEPA is actually giving a very detailed problem specification, which includes how to make sense of the input, what is the purpose and context of this particular part of the pipeline, what are some key observations and lessons from the data. So the prompt we are seeing here is for the second hop of a multi-hop question answering system, where given a question, we need to retrieve some documents that could potentially answer that question, look at those documents, summarize it, and then finally answer the question.
And here what we see is JEPA has found out that first hop documents often cover one entity or aspect, and the second hop should actually be recovering documents that are related to it. We have seen that human engineering teams, whenever a new model comes out, spend weeks of their time manually tweaking one word here and there, trying to discover the problem specification. This entire process is fully automated now with JEPA, which takes about half an hour to one hour to run depending on your pipelines. We can also apply JEPA to leading proprietary models. Just for an example, here we were able to optimize GPT 4.1 Mini's performance to outperform GPT 4.1 on a math task.
And we can see the kind of information distillation JEPA has done in the prompt space itself. Coming back to the problem of sample efficiency, AMD developed a new hardware accelerator called NPU-XDNA2, which used a completely new API to program, which had almost zero available information on the internet. And because of this, the leading models at the time, which was GPT 4.0, was failing miserably to perform this task. We were able to take an existing agent, which was getting 4.25% on this task, and applied JEPA without any other change to the agent itself. And we got this prompt and pushed this performance 7x to 30.52%.
So what this goes to say is there can be lots of domain-specific information, which if you include in your AI systems prompts, the models could actually perform much better, and JEPA can help you fully automatically discover that. I want to highlight the sentence saying avoid including ADF.Edge. Now the interesting thing is AMD actually ships a library called ADF.Edge for programming NPUs, but that did not work with this latest generation of hardware that we were working with, and JEPA was able to discover that in just one step. So how does it work?
It's an extremely simple algorithm, which simply takes your AI pipeline written in any agentic framework, or even raw LLM calls that you may have. It simply runs your systems on a few examples and collects domain-specific feedback. Whatever information your environment contains is observed. Second, it runs reflection with an LLM or agent that reads the feedback and proposes a better prompt. Finally, and most importantly, it keeps a Pareto pool where it keeps every single candidate that wins on even one training example and not just the top scorer. The question is, but why keep a Pareto pool?
And we kept getting asked this question a lot, that is JEPA really better than running the model in a loop? So we went and tested it out, and what happens is a loop keeps only the best and gets stuck in a local optima. So on the left-hand side, you see a search tree that was generated by using an LLM in a loop. Starting from a seed prompt at the top left, we asked the LLM to improve the prompt. It improved the prompt and it generated a prompt that gave us the middle node. However, this prompt got stuck in a local optima, and once again when we asked the LLM to try and improve it, it proposed something but that was not actually better.
So it went back and it again tried to improve it. And it kept doing this and it exhausted all of the search budget. On the other hand, with JEPA's Pareto-based candidate selection strategy on the right, we can see that it maintains a much more balanced search process, eventually converging to a much higher score. Across four benchmarks, we saw that more than half of the gains seen with JEPA actually account for this, and it gets almost twice the performance gains that you would get with just applying the model in a loop. JEPA can perform really well across diverse benchmarks.
Here we see results on question answering, instruction following, claim verification, as well as math, which all the leading frontier model companies are already optimizing their models a lot for. However, this prompt got stuck in a local optima, and once again when we asked the LLM to try and improve it, it proposed something but that was not actually better. So it went back and it again tried to improve it. And it kept doing this and it exhausted all of the search budget. On the other hand, with JEPA's Pareto-based candidate selection strategy on the right, we can see that it maintains a much more balanced search process, eventually converging to a much higher score.
Across four benchmarks, we saw that more than half of the gains seen with JEPA actually account for this, and it gets almost twice the performance gains that you would get with just applying the model in a loop. JEPA can perform really well across diverse benchmarks. Here we see results on question answering, instruction following, claim verification, as well as math, which all the leading frontier model companies are already optimizing their models a lot for, and we are still able to get plus 10% just by optimizing the prompt on it.
So we have so far seen JEPA only optimizing the prompts, but JEPA goes far beyond prompts. And because prompts are just text artifacts that determine AI system behavior, the same algorithm can improve anything that you can express as a piece of text and you can score. For example, your entire agent harness is eventually just a Python or a JavaScript file, and we can apply the same kind of reflective optimization process to that entire file and we can work with it. So if you can write it as text and score it, JEPA can optimize it.
So with that insight in mind, we propose OptimizeAnything, which is a universal API for optimizing any text parameter, given any domain like code optimization where, let's say, you want to optimize the CUDA kernel code. The input is just that CUDA kernel code where an evaluator looks at this piece of code, maybe compiles it, profiles it, generates a bunch of related information that we call as actionable side information, which is then provided to an LLM, which proposes a better candidate maintaining this Pareto pool, and it keeps repeating this process till we get convergence.
The same thing can be applied to numeric optimization, where your numbers can actually be serialized as text, or harness optimization where an entire harness can be serialized as text, or even cloud scheduling policy optimization where the scheduling policy or heuristic algorithm can be expressed as a piece of text, and the evaluator can be something like the negative of cost or some function measuring accuracy, efficiency, and the actionable side information can be something like job traces, SLA violations, and so on.
The API is dead simple to use. All it requires is you give us the set of problems that you care to be solved, along with an evaluator function or a fitness function that returns a score along with any available domain-specific side information. If your domain produces expert feedback, return that. If your domain produces compiler error messages, profiler messages, tool call error messages, return that. If you have maybe a written up documentation, return that. Any kind of—it's a very open-ended dictionary—you can return literally anything, and all you do is you call optimize anything with this fitness function and the set of problems that you have, and optimize anything will take care of it and give you an optimized solution.
Let's see some application. Let's say you were tasked with generating a 3D unicorn. This is all the code that you would write or your agent can now write it because we have seen that optimize anything is a very easy to use API for leading agents like CloudCode. So all you do is write this code which says optimize a Python program to generate a 3D unicorn, and the candidate is a Python script that produces a PNG rendering, whatever. And here is the result. On the left-hand side, we can see Cloud Opus 4.6. If you gave it this task, this is what it generated. And on the right-hand side, the unicorn that we get with optimize anything.
This is just for fun, but let's say you were tasked with writing an agent to solve a specific task. Typically, teams spend lots and lots of time tweaking their agents, building tools for it, writing tool descriptions, carefully orchestrating the control flow, and so on. Here we started with a simple four-line Python program that was simply calling a model's chain of thought to solve an RKGI problem. Within just 16 rounds of reflection, JEPA within OptimizeAnything was able to find this sophisticated six-step agent that took RKGI accuracy of Gemini Flash from 32.5% to 89.5%. And we can see that this agent is automatically, by itself, doing rule hypothesis induction, code synthesis. It executes and traces the code, automatically debugs this code, goes back and proposes new versions of that code, and finally it runs it on the actual test inputs and returns the output.
This is a runnable example. You can go to this QR code and you can run this example right now. So applying the same approach of discovering agent harnesses to MAT500, we are able to push its accuracy of GPT 4.1 nano by 20%, by simply creating a two-step agent. And again, I want to emphasize that all we did is we asked OptimizeAnything to optimize an agent file, and it was automatically discovering the sophisticated agent architecture, and we did not have to do anything other than specifying the objective and the task.
Finally, every single one of us is using some coding agent like Cloud Code or Codex or maybe your favorite agent, and agent skills has become a very leading part of the ecosystem where almost all coding agents understand skills. Let's say you want to optimize skills for your specific repository. This is the code that you write, which says, learn a skill from the trajectory. When the coding agent is presented with a similar problem, the skill should be helpful. We just give it this natural language behavior, and what we see is we started with MiniSui agent with GPT 5 Mini, because we were very budget constrained, and we were able to take its performance from 24% to 93%, almost a 3x jump on Go repository issue resolution, but more importantly, the skills that were optimized very cheaply on a GPT 5 Mini agent, we were able to take that and apply it to the latest Cloud Sonnet. This was done about a few months back, but we applied it to Cloud Sonnet 4.5, pushing its accuracy to 100% issue resolution, while more importantly, cutting down the execution time or issue resolution time by almost 50%. We cut it down in half, which also means it spent less tokens, because skills contain information about how the repository is organized, how to invoke the test cases, where a particular feature is implemented, what are the build systems used by this repository, and so on.
This is a feature called G-Skill. You can find it in the JEPA repository, and it's fully open source as well.
So OptimizeAnything is a single interface that provides three optimization modes. If you have just a single problem, like there is a single matrix multiplication kernel that you want to optimize, you can use it that way. If you have n number of related problems, like you want to optimize a matrix multiplication kernel, along with a dot product kernel, and you know there might be some information transfer between these two, you can use what we call as the multitask search mode. And finally, build a skill, which is, if you want to optimize on a set number of problems, but your deployment can actually come up with many new problems. So in case of math prompt optimization, we are training on some examples, but when we deploy it, we can receive a completely new kind of query. So we care about generalization modes. So there you can do prompt optimization, agent architecture optimization, and so on.
So OptimizeAnything can be used for a broad set of domains, including cloud scheduling policy optimization, where we are able to cut costs by almost 40% compared to expert heuristics, write custom solvers to match and exceed Optina even in black box mathematical optimization, create agent skills, prompt optimization, and so on. It is so easy to use that within just 20 hours of releasing it, people at Snorkel had already improved some of their internal benchmarks with it and were tweeting about it.
So JEPA also improves multimodal VLM models performance. Here, we are able to cut OCR error rates for leading models by almost 35%, and this is an externally validated report. Similarly, Databricks actually achieved 90x cost reduction in their deployed agent's performance, and here they were able to tune GPT OSS 120B to outperform Claude Opus while being 90x cheaper. More importantly, the performance delta improvement that you see on top of Claude Opus is actually bigger than the one you see on open source models.
Some people have asked me that, oh, as models get better, the importance of prompt optimization will go down. I argue the opposite, which is, as models get better, they will get better at instruction following, and the more precise instruction about your task that you have to give to a very smart model, the better that model will be at solving your task. And this is exactly what we see happening here. The better the instruction was, Claude Opus actually jumped much higher. Some people have this question of what if we have subjective tasks which are very hard to evaluate. JEPA can actually learn evals for your task from production traces.
and here they were able to tune GPT OSS 120B to outperform Claude Opus while being 90x cheaper. More importantly, the performance delta improvement that you see on top of Claude Opus is actually bigger than the one you see on open source models. Some people have asked me, as models get better, the importance of prompt optimization will go down. I argue the opposite, which is, as models get better, they will get better at instruction following, and the more precise instruction about your task that you have to give to a very smart model, the better that model will be at solving your task. And this is exactly what we see happening here.
The better the instruction was, Claude Opus actually jumped much higher. Some people have this question of what if we have subjective tasks which are very hard to evaluate. JEPA can actually learn evals for your task from production traces. The way to do that is you collect a bunch of production traces from your agent, get a human to annotate just about 50 of those trajectories, giving very detailed feedback. This is a long response, this is a short response, this is a good response, this uses this terminology, whatever. And once you get those human annotations, you can use JEPA to optimize an LLM as a judge prompt,
and you can use that LLM as a judge prompt then to go back and optimize your agent, and deploy that agent, and this becomes a data flywheel where you can keep improving it, and this is a successful paradigm that some leading teams in production are already using. Then the question we get asked is, can we actually use this reflective optimization to train models? And we recently had this paper called Learning Fast and Slow where we propose fast slow learning, where we can co-optimize model weights and prompt harnesses, and this shows some very strong properties that one would want in a continual learning algorithm.
I don't have much time to go over details, but please look at the papers. And since release, JEPA has been used in production by these companies as well as the main methodology in these papers. And here the CEO of Dropbox and Shopify are talking about their use of JEPA, and OpenAI also wrote a blog post about how you can build self-improving AI systems with JEPA. So it's very simple to get started. It can plug into any framework, any model, and it has absolutely zero hard dependencies, so you can deploy it in any kind of setting. So don't be afraid to optimize in the tech space, and many problems can be framed as optimization,
so bring actionable site information and surface as much domain-specific information as you can to optimizers, and the optimizers of future will be able to work with them. So please go and check it out. Thank you very much. For example, your entire agent harness is eventually just a Python or a JavaScript file, and we can apply the same kind of reflective optimization process to that entire file and we can work with it. So if you can write it as text and score it, JEPA can optimize it. So with that insight in mind, we propose OptimizeAnything, which is a universal API for optimizing any text parameter,
given any domain like code optimization where, let's say, you want to optimize the CUDA kernel code. The input is just that CUDA kernel code where an evaluator looks at this piece of code, maybe compiles it, profiles it, generates a bunch of related information that we call as actionable side information, which is then provided to an LLM, which proposes a better candidate maintaining this Pareto pool, and it keeps repeating this process till we get convergence. The same thing can be applied to numeric optimization, where your numbers can actually be serialized as text, or harness optimization where an entire harness can be serialized as text,
or even cloud scheduling policy optimization where the scheduling policy or heuristic algorithm can be expressed as a piece of text, and the evaluator can be something like the negative of cost or some function measuring accuracy, efficiency, and the actionable side information can be something like job traces, SLA violations, and so on. The API is dead simple to use. All it requires is you give us the set of problems that you care to be solved, along with an evaluator function or a fitness function that returns a score along with any available domain-specific side information. If your domain produces expert feedback, return that.
If your domain produces compiler error messages, profiler messages, tool call error messages, return that. If you have maybe a written up documentation, return that. Any kind of, it's a very open-ended dictionary, you can return literally anything, and all you do is you call optimize anything with this fitness function and the set of problems that you have, and optimize anything will sort of take care of it and give you an optimized solution. Let's see some application. Let's say you were tasked with generating a 3D unicorn. This is all the code that you would write or your agent can now write it because we have seen that
optimize anything is a very easy to use API for leading agents like CloudCode. So all you do is write this code which says optimize a Python program to generate a 3D unicorn, and the candidate is a Python script that produces a PNG rendering, whatever. And here is the result. On the left-hand side, we can see Cloud Opus 4.6. If you gave it this task, this is what it generated. And on the right-hand side, the unicorn that we get with optimize anything. This is just for fun, but let's say you were tasked with writing an agent to solve a specific task. Typically, teams spend lots and lots of time tweaking their agents, building tools for it,
writing tool descriptions, carefully orchestrating the control flow, and so on. Here we started with a simple four-line Python program that was simply calling a model's chain of thought to solve an RKGI problem. Within just 16 rounds of reflection, Jepa within OptimizeAnything was able to find this sophisticated six-step agent that took RKGI accuracy of Gemini Flash from 32.5% to 89.5%. And we can see that this agent is automatically, by itself, doing rule hypothesis induction, code synthesis. It executes and traces the code, automatically debugs this code, goes back and proposes new versions of that code,
and finally it runs it on the actual test inputs and returns the output. This is a runnable example. You can go to this QR code and you can run this example right now. So applying the same approach of discovering agent harnesses to MAT500, we are able to push its accuracy of GPT 4.1 nano by 20%, by simply creating a two-step agent. And again, I want to emphasize that all we did is we asked OptimizeAnything to optimize an agent file, and it was automatically discovering the sophisticated agent architecture, and we did not have to do anything other than specifying the objective and the task.
Finally, every single one of us is using some coding agent like Cloud Code or Codex or maybe your favorite agent, and agent skills has become a very leading part of the ecosystem where almost all coding agents understand skills. Let's say you want to optimize skills for your specific repository. This is the code that you write, which says, learn a skill from the trajectory. When the coding agent is presented with a similar problem, the skill should be helpful. We just give it this natural language behavior, and what we see is we started with MiniSui agent with GPT 5 Mini, because we were very budget constrained, and we were able to take its performance from 24% to 93%,
almost 3x jump on Go repository issue resolution, but more importantly, the skills that were optimized very cheaply on a GPT 5 Mini agent, we were able to take that and apply it to the latest Cloud Sonnet. This was done about a few months back, but we applied it to Cloud Sonnet 4.5, pushing its accuracy to 100% issue resolution, while more importantly, cutting down the execution time or issue resolution time by almost 50%. We cut it down into half, which also means it spent less tokens, because skills contain information about how the repository is organized,
how to invoke the test cases, where a particular feature is implemented, what are the build systems used by this repository, and so on. This is a feature called G-Skill. You can find it in the JEPA repository, and it's fully open source as well. So, Optimize Anything is a single interface that provides three optimization modes. If you have just a single problem, like there is a single matrix multiplication kernel that you want to optimize, you can use it that way. If you have n number of related problems, like you want to optimize a matrix multiplication kernel, along with a dot product kernel, and you know there might be some information transfer between these two,
you can use what we call as the multitask search mode. And finally, build a skill, which is, if you want to optimize on a set number of problems, but your deployment can actually come up with many new problems. So, like, in case of math prompt optimization, we are training on some examples, but when we deploy it, we can receive a completely new kind of query. So, we care about generalization modes. So, there you can do prompt optimization, agent architecture optimization, and so on. So, Optimize Anything can be used for a broad set of domains, including cloud scheduling policy optimization, where we are able to cut costs by almost 40% compared to expert heuristics,
write custom solvers to match and exceed Optina even in black box mathematical optimization, create agent skills, prompt optimization, and so on. It is so easy to use that within just 20 hours of releasing it, people at Snorkel had already improved some of their internal benchmarks with it and were tweeting about it. So, and JEPA also improves multimodal VLM models performance. Here, we are able to cut OCR error rates for leading models by almost 35%, and this is an externally validated report. Similarly, Databricks actually achieved 90x cost reduction in their deployed agent's performance,
and here they were able to tune GPT OSS 120B to outperform Claude Opus while being 90x cheaper. More importantly, the performance delta improvement that you see on top of Claude Opus is actually bigger than the one you see on open source models. Some people have asked me that, oh, as models get better, the importance of prompt optimization will go down. I argue the opposite, which is, as models get better, they will get better at instruction following, and the more precise instruction about your task that you have to give to a very smart model, the better that model will be at solving your task. And this is exactly what we see happening here.
The better the instruction was, Claude Opus actually jumped much higher. Some people have this question of what if we have subjective tasks which are very hard to evaluate. JEPA can actually learn evals for your task from production traces. The way to do that is you collect a bunch of production traces from your agent, get a human to annotate just about 50 of those trajectories, giving very detailed feedback. This is a long response, this is a short response, this is a good response, this uses this terminology, whatever. And once you get those human annotations, you can use JEPA to optimize an LLM as a judge prompt,
and you can use that LLM as a judge prompt then to go back and optimize your agent, and deploy that agent, and this becomes a data flywheel where you can keep improving it, and this is a successful paradigm that some leading teams in production are already using. Then the question we get asked is, can we actually use this reflective optimization to train models? And we recently had this paper called Learning Fast and Slow where we propose fast, slow learning, where we can co-optimize model weights and prompt harnesses, and this shows some very strong properties that one would want in a continual learning algorithm.
I don't have much time to go over details, but please look at the papers. And since release, JEPA has been used in production by these companies as well as the main methodology in these papers. And here the CEO of Dropbox and Shopify are talking about their use of JEPA, and OpenAI also wrote a blog post about how you can build self-improving AI systems with JEPA. So it's very simple to get started. It can plug into any framework, any model, and it has absolutely zero hard dependencies, so you can deploy it in any kind of setting. So don't be afraid to optimize in the tech space, and many problems can be framed as optimization,
so bring actionable site information and surface as much domain-specific information as you can to optimizers, and the optimizers of future will be able to work with them. So please go and check it out. Thank you very much. Thank you very much. Thank you very much. Thank you very much. Thank you very much. Thank you very much. Thank you very much. Thank you very much.
Thank you very much. Thank you very much. you