AI Engineer

Einstein Arena: Harnessing Collective Agent Intelligence for Open Science — James Zou, Together AI

1952 summary words 9 min summary Watch video

Start with the signal

9 min read

Summary

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: Rather than rigidly scripting agent workflows, build open, verifier-backed environments with incentives, shared infrastructure, and collaboration/competition mechanisms so collective agent intelligence can produce solutions no individual agent reaches.
  • Why it matters: This is a concrete control-plane pattern for multi-agent systems: deterministic evaluation, public artifacts, real-time feedback, role specialization, and lineage can turn agents from task executors into an iterative problem-solving market.
  • Best use: Use it to inform the design of agent arenas for optimization, research, coding, or operational work where outcomes can be objectively verified and agents can safely reuse one another's work.

Executive Summary

James Zou presents an argument for shifting from agent workflows and harnesses toward agent environments. In his framing, a workflow tells an agent how to work; an environment defines where it works and supplies resources, incentives, guardrails, and feedback. The intended result is more adaptive and creative collective behavior than a fixed sequence of prompts and tools permits.

The flagship example is Einstein Arena, an open, agent-native scientific problem-solving environment. Agents select curated open-ended problems, discuss approaches in a forum, submit solutions, receive real-time deterministic scores, inspect and download others' solutions, and compete on a leaderboard. Zou argues that this mix of transparent collaboration and competition enabled agents to find best-known solutions to 11 problems shortly after launch.

Its headline case is the 11-dimensional kissing-number construction: the stated best-known construction rose from 593 spheres after a 2023 DeepMind advance to 604 through multi-agent refinement in the Arena. Zou emphasizes that no single frontier agent could solve the problem alone; visible solution lineage and discussion-forum knowledge sharing were essential. The same pattern reportedly produced over 2x speedups on some production kernels now used at Together AI.

The second example, DSGEM, applies the environment approach to training and evaluating data-science agents. It combines code execution, parallel Docker-based experimentation, curated tasks, and execution-verified trajectories. Its key benchmark critique is that existing data-science evaluations can be shortcut: agents allegedly solve 20-50% of some tasks without using the underlying data. DSGEM is designed to avoid this and to generate verified training data for small locally runnable open models.

Key Takeaways

  • Claim: The central design move is to replace prescriptive agent workflows with environments that constrain and motivate agents without dictating their exact process. | Evidence: Zou distinguishes workflows—steps, prompts, tools, and instructions that tell agents how to work—from environments that provide incentives, infrastructure, guardrails, and resources defining where they work. | Implication: For complex problems with multiple viable solution paths, Ken should treat the evaluator, reward structure, shared workspace, and resource boundaries as first-class system design choices rather than over-specifying a single orchestration path. | Caveat: The talk makes a directional design argument rather than presenting a controlled comparison across many workflow and environment designs.
  • Claim: A useful multi-agent environment combines deterministic verification with both public collaboration and competition. | Evidence: Einstein Arena gives agents curated problems, a discussion forum, downloadable peer solutions, a real-time scored leaderboard, and deterministic verifiers for each problem; agents must solve a small puzzle to establish that they are agents before participating. | Implication: The strongest early applications are domains such as code optimization, mathematical construction, simulation, testing, and data analysis where submissions can be automatically checked rather than judged only by another model. | Caveat: This architecture depends on tasks having a well-defined, deterministic quality verifier, which excludes many ambiguous or subjective knowledge-work outputs.
  • Claim: Collective agent iteration can outperform isolated frontier agents on open scientific search problems. | Evidence: Zou says Arena agents found new best solutions to 11 curated problems within weeks of its March launch. For the 11-dimensional kissing-number problem, agents reportedly advanced the construction from 593 spheres—the prior 2023 DeepMind result—to 604 in a few days. | Implication: Do not assess an agent system solely by single-agent benchmark performance; measure whether agents can inherit, critique, mutate, and improve shared artifacts across an extended search process. | Caveat: The presentation does not provide independent replication details, exact participating models, compute budgets, or formal proof-review results beyond its assertion that the verifier validates submissions.
  • Claim: Specialized agent personas can improve collective technical optimization by inducing complementary search behavior. | Evidence: In the kernel-optimization arena, Zou describes personas oriented toward profiling, memory consumption, and tensor-computation precision, all working against compile, test, benchmark, and correctness checks. | Implication: For engineering-agent swarms, assign roles based on distinct diagnostic priors and measurable failure surfaces—not generic job titles—and make their outputs composable in a shared evaluation loop. | Caveat: The talk does not isolate how much of the reported gain came from persona specialization versus the shared leaderboard, verifier, or aggregate compute.
  • Claim: The environment pattern can generate commercially useful optimization results, not only scientific demonstrations. | Evidence: Together AI reports up to over 2x speedups for some production kernels, including paged-attention examples across specific shapes, with generalized work across multiple shapes and hardware types; Zou says the agent-designed kernels are already in production. | Implication: A verifier-backed arena is particularly attractive where performance improvements have direct economic value and automated regression checks can protect correctness. | Caveat: The over-2x figure applies to some kernels and configurations, not necessarily end-to-end model-serving throughput or every hardware and shape combination.
  • Claim: Many data-science agent benchmarks overstate capability because agents can solve substantial portions without touching the underlying datasets. | Evidence: Zou reports that on three popular benchmarks, shortcut reasoning or other non-data-based methods can solve roughly 20-50% of tasks. DSGEM instead curates scientific tasks from recent papers with expert review and predictive tasks from recent open Kaggle competitions. | Implication: Ken should require execution-grounded evaluation for analytical agents: inspect dataset access, code execution, artifact provenance, and whether a result changes appropriately when the input data changes. | Caveat: The transcript does not name the three benchmark datasets or provide the measurement protocol behind the 20-50% shortcut estimates.
  • Claim: Execution-verified agent trajectories can serve as a training-data factory for smaller domain agents. | Evidence: DSGEM contains over 1,000 tasks across dozens of domains from biology and physics to economics, supports parallel Docker containers, and verifies trajectories by executing agent code; Zou says fine-tuning on these trajectories produced best-in-class open-source data-science models small enough to run locally. | Implication: Instead of relying only on expensive frontier-model inference, build a closed loop in which successful, reproducible execution traces train cheaper specialized models for recurring internal workloads. | Caveat: “Best-in-class” is not defined in the talk by model size, comparison set, task split, or exact score.

Detailed Brief

Einstein Arena operating model

  • Claims: Problems are selected not merely for difficulty but because an existing human research community cares about them and the submission quality can be deterministically assessed.; The intended mechanism resembles an open scientific community: agents choose problems autonomously, exchange partial findings, observe competitors' artifacts, and repeatedly improve on public work.; The system records a lineage of solution refinements, making it possible to trace how multiple contributions led to a final advance.
  • Evidence: The kissing-number problem asks for the maximum number of non-overlapping equal spheres around a central sphere; it is 2 in one dimension and 6 in two dimensions, but remains difficult and open in higher dimensions.; For the 11-dimensional case, Zou recounts a progression from 440 in the 1980s to 582 in 1980, 592 in 2022, 593 in 2023, and 604 in the Arena.; The talk cites discussion around semidefinite-programming (SDP) approaches as an example of agents sharing attempted methods and findings.
  • Caveats: Open sharing of downloadable solutions is beneficial for cumulative search but creates attribution, credit-allocation, contamination, and potential convergence-on-common-path risks in settings where agents or users have misaligned incentives.; Scientific-value claims rest on the quality of problem curation and verifier implementation; a deterministic verifier can confirm a stated objective without necessarily capturing every broader scientific consideration.
  • Implications: The design is a reusable market-like coordination layer: publish state, make progress legible, reward improvements, and preserve reusable artifacts.; For proprietary deployments, public sharing can be replaced with scoped access controls, but the shared-artifact and independently checked feedback loop should remain.

DSGEM as evaluation and training infrastructure

  • Claims: DSGEM is positioned as one unified execution layer for both evaluating data-science agents and creating training material for them.; Its task mix deliberately separates scientific analysis/discovery from predictive-modeling work, broadening the evaluation beyond a narrow family of notebook tasks.; Frontier models reportedly remain below 50% accuracy on DSGEM, so Zou presents the benchmark as unsaturated rather than a solved leaderboard.
  • Evidence: Agents interact with datasets through a unified interface and code execution, and can launch multiple Docker containers to run experiments in parallel.; Scientific-analysis tasks are derived from recently published papers and reviewed by human scientists and experts.; Predictive-modeling tasks are curated from recent Kaggle competitions that remain open and have high-quality datasets and evaluations.
  • Caveats: The talk does not describe leakage controls for paper-derived tasks, Kaggle-derived tasks, or model pretraining exposure.; An execution-verified trajectory establishes that code ran successfully under the environment; it does not alone establish that an analytical conclusion is causal, robust, or decision-useful.
  • Implications: The durable asset is not just a benchmark score but the corpus of reproducible trajectories, failures, and fixes generated by the environment.; For internal agent training, use held-out task families and perturbation tests so models cannot succeed through template memorization or benchmark-specific shortcuts.

Notable Concepts & Terms

  • Environment design: The proposed successor to rigid workflows: define incentives, infrastructure, guardrails, resources, and an objective rather than a fixed agent procedure.
  • Einstein Arena: Together AI and Stanford's agent-native environment for open scientific problems, combining forums, leaderboards, solution sharing, and deterministic verification.
  • Deterministic verifier: An automatic checker that evaluates solution quality in real time; it is the foundation for trustworthy reward signals and leaderboard competition.
  • Collective agent intelligence: Problem-solving capacity that emerges when multiple agents exchange artifacts and refine one another's approaches, beyond what Zou says any one agent achieved alone.
  • Kissing number problem: A sphere-packing construction problem used as the talk's headline example of multi-agent scientific progress; the claimed 11D construction reaches 604 spheres.
  • Agent personas: Deliberately differentiated roles or priors—such as profiling, memory, and numerical precision—to diversify search in technical optimization.
  • DSGEM: Data Science GEM, a unified environment for evaluating and training data-science agents with real datasets, code execution, and curated tasks.
  • Execution-verified trajectories: Agent solution traces validated by running their code; DSGEM uses them as synthetic fine-tuning data for smaller open models.

Operator Notes / Why Ken Should Care

  • Identify one internal optimization or research workflow with a deterministic scorer and run a bounded arena experiment before attempting broad autonomous-agent orchestration.
  • Design the arena contract explicitly: submission format, sandbox limits, score function, regression suite, artifact visibility, lineage logging, rollback policy, and reward for incremental improvements.
  • Test whether role diversity adds measurable value by running ablations: homogeneous agents versus profiling/memory/correctness/security-specialized agents under the same compute budget.
  • For any analytical-agent benchmark, add a no-data-access baseline and input-perturbation tests to quantify shortcutting before trusting reported accuracy.
  • Treat public solution sharing as a deployment choice, not a default: preserve reusable artifacts internally while applying permissions, provenance, and attribution controls for sensitive work.
  • Request the cited papers and implementation details before relying on the 604-sphere, 11-problem, over-2x kernel, or best-in-class model claims in strategy or investment decisions.

Source/Metadata

  • Title: Einstein Arena: Harnessing Collective Agent Intelligence for Open Science — James Zou, Together AI
  • Transcript words: 4833
  • Duration seconds: 1015
  • Timestamp note: No timestamps or chapters were provided; the transcript includes a substantial repeated passage from the Einstein Arena section onward.
Full transcript 2746 words · 21 min read
0:05

James So All right, I think we'll go ahead and get started with the presentation. My name is James So. I am going to explain some of the work we're doing with Together AI, and it's also in collaboration with Stanford, around designing and optimizing environments for AI agents to enable these agents to make new kinds of scientific discoveries.

0:20

The current paradigm of how people are often using or deploying AI agents often involves designing workflows that tell the agents what to do, or how the agents should work. It's typically done through a series of steps or prompts, tools, and instructions. In contrast, the way we imagine the environment is that the environment should really specify not how the agent should work, but where the agent should work, and the environment then should provide a set of incentives, infrastructure, guardrails, and resources so that the agent can then flexibly work within that environment.

0:25

Our thesis here is that as agents become more and more powerful, if we try to design workflows, it often can limit the capabilities and creativity of the agents. Whereas if we properly design the environment, this can enable a lot more creativity, capabilities, and intelligence for the agents to naturally emerge. This is why I think we're trying to shift away from designing workflows and harnesses toward designing environments.

0:38

What I want to do today is give a few examples of how we design environments for agents, and in particular also show how they are then able, within the right environment, to actually solve some really interesting and innovative problems.

0:46

The first example I want to share is the system environment that we created called the Einstein Arena. It's one of the first environments that enables AI agents to be able to collaborate in the wild and to compete to really solve open-ended scientific problems. We designed the Einstein Arena to be really agent-native. That means that it's very easy for agents to just read the docs on the arena and be able to access the arena. You actually have to solve a little puzzle to prove that you're an AI agent in order to participate in this arena. But any agent in the world can openly and freely participate in the arena.

0:54

Once the agent actually enters into the Einstein Arena, this is what they'll see. They'll actually see a list of curated problems. Each of these problems is actually a problem that we curated, so it's a scientifically interesting problem. And we curated these problems so that, first, there's actually an existing community of human researchers that are interested in these problems. So these are important problems for human scientists. And second, for each of these problems, we can actually create a well-defined and deterministic verifier to assess the quality of the solutions to each of these problems. I'll give some examples in a couple of slides.

1:00

The agents can actually decide which of these problems they're interested in once they log on to the arena. If they enter into a particular problem space, this is what they will see. They'll see some description that precisely explains what the problem is. We have a discussion forum where the agents can communicate. It's almost like a social network where the agents can actually communicate and talk to each other and ask for help or give recommendations. And we also have a leaderboard. This is where the agents can actually see each other's solutions.

1:05

At any time they want, the agents can actually submit a solution to one of these problems. And because we have this verifier, we can actually then determine the quality of that solution and provide a score in real time. So this leaderboard is constantly updated in real time. The agents can also see how other agents are doing on this problem. They can also see other agents' solutions and download those solutions.

1:10

So there's both collaboration dynamics and competition dynamics in this arena. They can collaborate and ask each other questions and help in the discussion forum. But agents are also competing with each other. That's why I think they also simulate how human researchers can compete and also collaborate to solve interesting problems.

1:17

We launched this environment earlier this year, I think in March. Within a few weeks, we were very impressed and very surprised that the agents were actually able to already discover new solutions to 11 problems. These are the best solutions that have ever been found. That means that the solutions discovered by the agents online are certainly better than any previous human solutions or any solutions that you acquired using more specialized AI tools.

1:23

I'll give you an example of one such solution, or one such problem, which is called the kissing number problem. This is actually a very famous problem. It's been around for hundreds of years. For example, Isaac Newton was already working on some versions of this kissing number problem. And it's actually relatively easy to state.

1:30

The kissing number problem asks what is the maximum number of spheres that you can place around a central sphere so that these additional spheres do not overlap each other. For example, in one dimension, around the central sphere, I can place one sphere to the left and one sphere to the right without overlap. So the kissing number in one dimension is easy to compute. It's just two. In two dimensions, it's also easy to show that you can at most place six spheres. So the kissing number in two dimensions is six.

1:36

But it turns out that in higher dimensions, it actually becomes really hard to compute what's the maximum number of non-overlapping spheres. And the kissing number problem in higher dimensions is actually open. It's not clear what the optimal number is. Scientists have been trying to work on this problem for the last several centuries.

1:41

In particular, the kissing number problem in 11 dimensions has attracted a lot of interest for various reasons. This is the progression of the solutions in 11 dimensions. In the 1980s, it was best known that you can place 440 spheres in 11 dimensions without overlap. In 1980, there was a big advance that for the first time showed that you can actually construct 582 spheres in 11 dimensions without overlap. Then it was stuck there for about 40 years, until 2022, when a mathematician was able to publish a new advance, a breakthrough that was able to improve that to 592 spheres. Then there was another breakthrough from DeepMind the following year that advanced that to 593 spheres.

1:46

But on the Einstein Arena, by having these agents able to collaborate actively in the wild, within a few days they were actually able to construct a new solution that shows that, for the first time, you can create 604 spheres in 11 dimensions that do not overlap. And this is not just a problem that's of mathematical interest, because it turns out that the more of these spheres you can place in higher dimensions without overlap, that actually creates better coding systems, including ways of doing error correction codes for information transfer. So by creating these better constructions, that also leads to better engineering algorithms.

1:52

In this case, the collaborations among these agents are really critical for making these advances. This is a problem where not a single agent is able to solve it by itself. GPT 5.5 or Claude models can't really solve the problem by themselves. So the collaboration among multiple agents is really critical. Here we're actually able to show that there's a lineage trace of how the agents are able to collaborate and then take each other's solutions and refine them and further optimize them to arrive at this breakthrough.

1:56

You can also see some of these interactions and discussions on Einstein Arena. Here's an example where one agent was actually asking other agents, have you tried some of these approaches with these SDP approaches? And then the other agents showed that, yes, we have tried these approaches, and here are some of the things that we found. So the information sharing on the forums in the arena is actually really important to help the agents arrive at this solution together.

2:01

In addition to solving these interesting scientific problems, we've also been using platforms like the Einstein Arena to help improve machine learning and AI itself. Here's one example where we actually use these agents to help us create better kernels and speed up those kernels. Here we use the same environment where the agents can compete and also collaborate, and they see this leaderboard. We'll basically change the backend. Instead of trying to verify the solutions to this mathematics problem, here we'll try to compile, benchmark, test, and verify the quality and the speed of the individual kernels. Then we'll provide the feedback to the agents in real time in the form of these leaderboards.

2:06

In this kernel setting, we also found it to be quite useful to have different agents with different personas. These different personas actually correspond to different roles and priors that agents can have. For example, we have one agent that tends to look more at profiling, another agent that tends to look more at memory consumption, and a third agent that looks at the precision of the tensor computations. These agents, across different personas, can collaborate and compete on the arena to speed up the kernels.

2:14

In this case, the agents were also able to collaborate and lead to really quite substantial speedups, including sometimes over 2x speedups in some of these production kernels. Here I'm just showing you a few examples for things like page detention. These are for specific shapes, but we also have generalized this to many different shapes and different hardware types, where we're actually seeing that we're getting up to sometimes over 2x speedup in these kernels compared to the previous state-of-the-art kernels for these problems. These improved kernels, created and designed by the agents, are actually already used in production at Together AI.

2:22

In the last few minutes, I want to show a second example of a kind of environment that we created as a way to train and create better data scientist agents. We call this DSGEM, which stands for Data Science GEM, which is a unified environment that we created both for evaluating and training data science agents to solve complex data science problems.

2:34

Here in this DSGEM environment, we also curated and created a unified list of different data sets and tasks. These data sets span many different settings. The agents are then able to interact with these different data sets that we have through a unified interface and through code execution. In the DSGEM environment, we also provide a unified infrastructure for the agents. For example, the agents can actually spin up many different Docker containers to test their data science algorithms and run them in parallel.

2:47

In the process of actually creating the data sets and tasks for the DSGEM environment, we initially wanted to incorporate some of the existing data science benchmarks that have been used to evaluate agents. But we quickly realized that many of the existing widely used benchmarks actually have many problems. One big problem is that they're actually very vulnerable to shortcuts.

2:54

By shortcut, I mean that I'm showing three different common popular data science benchmarks. In green here shows the performance of the agents on these benchmarks. But the red bar also shows what fraction of the benchmark the agents can actually solve without actually using the data sets themselves, just by reasoning or by doing other shortcuts without actually working with the underlying data sets. Across many of these different benchmarks, sometimes up to 20 to 50 percent of the tasks can be solved without actually looking at any of the underlying data, which I think is really a significant problem with many of the existing benchmarks.

2:59

To address that, we actually carefully curated our own benchmarks, both for scientific analysis and also for predictive modeling. For scientific analysis and discovery, the way we did this is that we actually went through recently published papers and then carefully curated data and tasks from those papers. Then we also had human scientists and experts review each of those tasks.

3:06

For predictive modeling, the way we did this was to go through all the different Kaggle competitions to look for some of the recent Kaggle competitions that are still open, where you also have high-quality data sets and high-quality evaluations. Then we curated those into DSGEM as a kind of task for evaluating how well model agents can actually build predictive models.

3:11

Altogether in DSGEM, we actually created over a thousand different tasks. They span dozens of different scientific domains, ranging from biology to physics to economics. It also involves many different data types and data modalities. This actually makes it very easy for us to evaluate different models, both open- and closed-source models. One thing we found is that the existing models, even the frontier models, often still achieve less than 50% accuracy performance on the DSGEM tasks. So these are definitely not saturated benchmarks.

3:20

We can also use DSGEM as a training factory to improve these open-source models. One thing we did here is that in DSGEM, the system itself will actually create all these execution-verified trajectories, which means that these are trajectories generated by agents that have been verified through actually executing the code from the agents. By generating these execution-verified trajectories, then we are able to fine-tune small open-source models that now achieve best-in-class open-source model performance in terms of solving these kinds of data science tasks. These models are small enough that you can actually run them locally on your laptops and your computers.

3:25

So just to summarize this part on Data Science GEM, with DSGEM we created this unified execution layer so people can actually run all these different tasks across many different domains. We have carefully verified that there are no shortcuts in these tasks, which has been a common challenge with existing data science benchmarks. And we also enable in DSGEM a way to generate synthetic data so you can easily use that to improve and train your own data science agents.

3:30

Just to summarize the presentation, I think the main takeaway here is that we're seeing this interesting progression in terms of how we build different AI systems. In the past, people have been building these AI systems mostly by designing individual models or individual tools. Currently, there's a lot of focus on designing agents or harnesses and workflows around agents. But what our research shows is that I think we're really moving toward the next stage, where rather than trying to design workflows or specific agents, what we really want to do is design environments, which are a set of infrastructure and incentives that motivate the agents to actually solve more and more challenging problems.

3:34

With appropriate designs, these environments can actually unlock much more creativity and collective intelligence from the agents than is possible with existing workflows. Here are some of the references for the papers that we published that describe these in more detail. Thank you very much. We have a discussion forum where the agents can communicate. It's almost like a social network where the agents can actually communicate and talk to each other and ask for help or give recommendations. And we also have a leaderboard. This is where the agent can actually see each other's solutions.

3:55

Right? So in any time they want, the agents can actually submit a solution to one of these problems. And because we have this verifier, we can actually then determine what is the quality of that solution and provide a score in real time. So this leaderboard that is constantly updated in real time. And the agents can also see how other agents are doing on this problem. And they can also see other agents' solutions and download those solutions. So there's both the collaboration dynamics and also competition dynamics in this arena, right? They can collaborate and ask each other questions and help in the discussion forum. But agents are also competing with each other.

4:31

And that's why I think they also sort of simulate how human researchers can compete and also collaborate to solve interesting problems. So we launched this on the inside of the environment earlier this year, I think in March. And within a few weeks, it's actually, we're very impressed and very surprised that the agents are actually able to already discover new solutions to 11 problems. That's sort of the best solutions that have ever been found. Right? So that means that the solutions that are discovered by the agents online are certainly not more better than any previous human solutions or any solutions that you acquired using more specialized AI tools.

5:12

So I'll just give you an example of one such solution or one such problem, which is called the kissing number problem. So this is actually a very famous problem. It's been around for hundreds of years. So, for example, Isaac Newton was already working on some versions of this kissing number problem. And it's actually relatively easy to state. Right? So the kissing number problem basically asks that what is the maximum number of spheres that you can place around the central sphere so that these additional spheres do not overlap each other?

5:40

So, for example, in one dimensions, right, so around the central sphere, I can place one sphere to the left, one sphere to the right, without overlap. So the kissing number in one dimension is easy to compute. It's just two. In two dimensions, it's also easy to show that you can at most place six spheres. Right? So the kissing number in two dimensions is six. But it turns out that in higher dimensions, it actually becomes really hard to compute what's the maximum number of non-overlapping spheres. And the kissing number problem in higher dimensions is actually open. Right? It's not been, it's not clear what is the optimal number.

6:11

And so scientists have been trying to work on this problem for the last several centuries. And in particular, right, so the kissing number problem in 11 dimensions has attracted a lot of interest for various reasons. So this is actually sort of the progression of the solutions in 11 dimensions. So in the 1980s, right, so it's best known that you can place 440 spheres, right, in 11 dimensions without overlap.

6:46

So yeah, so in 1980, there was a big advance that for the first time showed that you can actually deconstruct was 582 spheres in 11 dimensions without overlap. And then it's sort of stuck there for about 40 years, right, until 2022, where a mathematician is able to publish a new advance, right, a breakthrough that's able to improve that to 592 spheres. And then there's another breakthrough from DeepMind the following year that advances that to 593 spheres. But with on the Einstein arena, by having these agents able to collaborate actively, right, in the wild,

7:24

within a few days, you're actually able to construct a new solution that shows that for the first time, you can create 604 spheres in 11 dimensions that do not overlap. And this is not just a problem that's of mathematical interest, because it turns out that the more of these spheres you can place in higher dimensions of that overlap, that actually creates the better coding systems, including ways of like doing error correction codes for information transfer. So this actually is by creating these better constructions, that's all these two is better engineering algorithms.

7:56

And in this case, actually, the collaborations among these agents is really critical for making these events, right? So this is a problem where not a single agent is able to solve by itself, right? Not, you know, GPT 5.5 or cloud models. They can't really solve the problem by itself. So the collaboration among multiple agents is really critical. And here we're actually able to show that there's this sort of a lineage trace of how the agents are able to collaborate, and then basically take each other's solutions and refine that and further optimize it to arrive at this breakthrough. And you can also see some of these interactions and discussions on Iceland Arena, right?

8:32

Where here's an example where one agent actually was asking other agents, have you tried some of these approaches with these SDP approaches? And then the other agents show that, yes, we have tried these approaches, and here are some of the things that we found. So the information sharing on the forums on the Arena is actually really important to help the agents to arrive at this solution together. So in addition to solving these interesting scientific problems, but we've also been using platforms like the Einstein Arena to help to improve machine learning and AI itself. So here's one example where we actually use these agents to basically help us to create better

9:13

kernels for and speed up those kernels. And here we use the same environment where the agents can compete and also can collaborate, and they see this leaderboard. And we'll basically change the backend instead of trying to verify the solutions to this mathematics problem. Here we'll basically try to compile and benchmark and test and verify the quality and the speed of the individual kernels. And then we'll provide the feedback to the agents in real time in the form of these leaderboards. In this kernel settings, we also found it to be quite useful to have different agents with different personas.

9:48

But these different personas actually corresponds to different roles and priors that agents can actually have. So for example, we have one agent that tends to look at more of the profiling, another agent that tends to look at more of the memory consumptions, and a third agent that looks at the precision, the tensor computations. And these agents can, across different personas, they can be able to collaborate and compete on the arena to speed up the kernels. And in this case, right here, the agents were also able to collaborate and lead to really quite substantial speed ups, including sometimes over 2x, two-fold speed ups in some of these production kernels.

10:26

So here I'm just showing you a few examples for things like page detention. And these are sort of for specific shapes, but we also have generalized this to many different shapes and different hardware types. Right? Where we're actually seeing that we're getting up to sometimes over 2x speed up in these kernels and compared to the previous state-of-the-art kernels for these problems. And these improved kernels created, designed by the agents are actually already used in production at Together AI.

10:56

So in the last few minutes, I want to show a second example of a kind of environment that we created as a way to train and to create better data scientist agents. Right? So we call this DSGEM, which stands for Data Science GEM, which is sort of like a unified environment that we created for both for evaluating and sort of training data science agents to solve complex data science problems. So here in this DSGEM environment, we also curated and created a unified list of different data sets and tasks. Right? So these data sets can combine spans across many different settings. And the agents are then able to interact with these different data sets that we have

11:37

through a unified interface and through code execution. In the DSGEM environment, we also provide a unified infrastructure for the agents. So for example, the agents can actually spin up many different Docker containers to test their data science algorithms and to run them in parallel.

11:57

So in the process of actually creating the data sets and tasks for the DSGEM environment, so we initially actually wanted to incorporate some of the existing data science benchmarks that have been used to evaluate agents. But we actually quickly realized that many of the existing widely used benchmarks actually have many problems. And one big problem is that they're actually very vulnerable to shortcuts. So by shortcut, I mean here is that here I'm showing sort of three different common popular data science benchmarks. Right? And in green here basically shows like the performance of the agents on these benchmarks.

12:31

But the red bar also shows how well they're able to, what fraction of the benchmark the agents can actually solve without actually using the data sets themselves. Right? So just by reasoning or by doing other shortcuts without actually working with the underlying data sets. And across many of these different benchmarks, right, sometimes up to 20 to 50 percent of the tasks can be solved without actually looking at any of the underlying data. Which I think is really a significant problem with many of the existing benchmarks. So to address that, we actually carefully curated our own benchmarks, right, for both for scientific analysis and also for predictive modeling.

13:12

So for scientific analysis and discovery, the way we did this is that we actually went through recently published papers and then carefully curated data and also tasks from those papers. And then we also had human scientists and experts to review each of those tasks. And for predictive modeling, the way we did this is actually go through all the different Kaggle competitions to look for some of the recent Kaggle competitions that are still open. And where also you have high quality data sets and also high quality evaluations. Then we curated those into DSGIM as a kind of task for evaluating how well models agents can actually build predictive models.

13:48

So altogether in DSGIM, we actually have created over a thousand different tasks. They span across dozens of different scientific domains, ranging from biology to physics to economics. It also involves many different data types and data modalities.

14:06

So this actually makes it very easy for us to evaluate different models, both open and closed source models. And one thing we found is that the existing models, even the frontier models, often only still achieves less than 50% accuracy performance on the DSGIM tasks. So these are definitely not saturated benchmarks. We can also use the DSGIM as sort of like a training factory to improve these open source models. So one thing we did here is that you generate in DSGIM, actually the GM itself will actually create all these execution verified trajectories, which means that these are trajectories generated by agents that have been verified through

14:44

actually executing the code from the agents. So by generating these execution verified trajectories, then we are able to like fine tune sort of small open source models that actually now achieve sort of the best in class open source models in terms of solving these kind of data science tasks. And these models are small enough that you can actually run them locally on your laptops and your computers. So just to summarize, this part was a data science gem. Right? So we, with DSGIM, we created this unified execution layer so people can actually run all these different tasks across dozens of different tasks across many different domains.

15:21

We have carefully verified that there are no shortcuts in these tasks, which has been sort of a common challenge with existing data science benchmarks. And we also enable in the DSGIM a way to generate synthetic data so you can easily use that to improve and to train your own data science agents. So just to summarize the presentation, I think the main takeaway here is that I think we're seeing this interesting progression in terms of how we build different AI systems. So in the past, people have been building these AI systems mostly by designing individual models or individual tools. And currently, there's a lot of focus on creating designing agents or harnesses and

16:00

workflows around agents. But what our research shows is that I think we're really moving towards the next stage, where rather than trying to design workflows or specific agents, what we really want to do is to design environments, which is a set of infrastructure and incentives that motivates the agents to actually solve more and more challenging problems. And with appropriate designs, these environments can actually unlock much more creativity and collective intelligence from the agents that's limited by the existing workflows. And here are some of the references for the papers that we published that describe these in more detail. So thank you very much.

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note