An AI Agent Became the #1 Contributor in OpenAI's Hiring Challenge — Zhengyao Jiang, Weco
Description
Earlier this year, OpenAI ran Parameter Golf, a model-training competition that doubled as a hiring filter. Over 1,000 researchers competed to train the best small language model under a 16MB cap. The top contributor was the one candidate OpenAI couldn't hire. Our autonomous research agent Aiden finished with 7 merged records, more than twice as many as any other contributor, and ended up the most-cited participant in the community. This talk is about what those 22 days showed. I'll cover on high level how does it works and which of its ideas produced the records. But the part worth more than the leaderboard is the collaboration itself, the community and AI agent building on each other's work, the largest natural experiment in human-AI collaboration I've seen run in public. I'll close with what it tells us about where humans and autonomous research each still matter for the foreseeable future. 1:57 PM # An AI Agent Became the #1 Contributor in OpenAI's Hiring Challenge **Location:** Main Stage **When:** Day 3 - July 1, 2026 · 1:55pm-2:15pm ## Speakers ### Zhengyao Jiang CEO & Cofounder · Weco AI [X/Twitter](https://x.com/zhengyaojiang) · [LinkedIn](https://www.linkedin.com/in/zhengyao-jiang-387b44145/) · [Website](https://zhengyaojiang.github.io/) Cofounder & CEO @WecoAI - automated hill climbing with LLMs. Previously: PhD in ML at UCL Timestamps 0:00 Introduction to Parameter Golf and the Aiden agent 1:06 Defining the challenge: Auto-research vs. human community 1:47 About Weco AI and the development of Aiden 3:07 Evaluating Aiden's impact and H-index in the community 4:01 Why autonomous AI is powerful: Throughput and efficiency 5:21 Human-AI collaboration: How ideas move the frontier 6:32 Case study: Combining research, architecture, and tokenization 7:41 Summary of auto-research strengths: Execution and search 9:06 The role of human design in competition 10:04 The Andrej Karpathy metaphor: Gradient descent and coding 11:19 Auto-research as training a model
Summary
Generated by gpt-5.6-solAt-a-Glance
- Verdict: Watch fully
- Core thesis: Autonomous research agents can already outperform individual ML engineers at sustained implementation and experimentation, but their value depends heavily on human-supplied ideas, well-designed evaluations, and code abstractions that constrain the search.
- Why it matters: Aden offers a concrete operating model for agentic R&D: agents search and execute at high throughput, while humans move up the stack to define objectives, environments, abstractions, and quality controls.
- Best use: Use this as an architecture and operating-model reference for building research or coding agents that produce mergeable work rather than merely optimizing isolated benchmark scores.
Executive Summary
Weco sent its autonomous research system, Aden, into OpenAI's 22-day Parameter Golf competition, where roughly 1,000 participants filed 2,000 submissions and only 47 passed review. Aden produced seven accepted leaderboard records—more than twice the best human contribution—and achieved an H-index of 10 across pull requests versus 7 for the next participant, suggesting that other competitors reused its work rather than merely observing its scores.
The result was not driven by brute-force compute or autonomous scientific originality. Aden ran approximately 1,300 experiments on one H100 node, consumed at most 4% of the competition's total compute, and produced about 15% of its records. Most of its successful ideas originated in human papers, community pull requests, or informal notes; Aden's advantage was finding those ideas, implementing abandoned or difficult variants, resolving constraints, and testing combinations efficiently.
The larger strategic argument is that auto research makes evaluation and codebase design more valuable, not less. The evaluation acts like a loss function or reinforcement-learning environment, while the code abstraction behaves like a model architecture that biases what the system can discover. Humans therefore move from manual hill-climbing toward designing the hill, controlling information boundaries, defining quality, and supplying creative judgment.
Key Takeaways
- Claim: Aden became the competition's most prolific accepted contributor and produced work that other participants built upon. | Evidence: Of approximately 2,000 submissions from 1,000 participants, only 47 passed OpenAI's review; seven accepted leaderboard records came from Aden, while the best human produced three. Aden's pull-request H-index was 10 versus 7 for the next human contributor. | Implication: Agent performance should be assessed through acceptance, reuse, and downstream influence—not only the benchmark score the agent directly optimizes. | Caveat: The H-index is borrowed from academic citation analysis and used here as a proxy for community influence; it does not by itself establish scientific novelty or long-term production value.
- Claim: Aden's performance came from efficient, high-quality search rather than overwhelming the competition with compute. | Evidence: Across 22 days, it ran about 1,300 experiments on a single H100 node, used at most 4% of the competition's total compute, generated roughly 15% of the leaderboard records, and placed 28% of its submissions on the leaderboard—about six times the community-average hit rate. | Implication: For agent systems, experiment selection, filtering, and quality gates may provide more leverage than simply maximizing parallel execution. | Caveat: These figures describe one bounded ML competition and do not demonstrate equivalent efficiency in open-ended research or production environments.
- Claim: The agent's primary advantage was execution and recombination, not independent generation of fundamentally new research ideas. | Evidence: Almost all ideas behind Aden's record pull requests were traced to research papers, Parameter Golf participants, nanoGPT, or informal comments about abandoned approaches. Only a small fraction were described as original ideas generated by Aden. | Implication: Near-term research agents should be designed as high-throughput implementers and synthesis engines with strong retrieval and provenance, rather than treated as replacements for human agenda-setting and conceptual invention. | Caveat: This makes Aden dependent on a rich human knowledge ecosystem and limits what the result proves about autonomous scientific creativity.
- Claim: Aden created value by recognizing synergy among individually weak or incomplete ideas. | Evidence: It implemented Gated Attention from a Qwen paper, then devised quantization when the added parameters violated the 16 MB file limit. Those changes alone barely improved the score, but combining them with another participant's tokenizer improvement produced a large performance jump and a leaderboard record. | Implication: Agent workflows should preserve failed and neutral experiments as reusable ingredients, because the valuable result may emerge from later combinations rather than from isolated interventions.
- Claim: Evaluation design is the control surface that determines what an auto-research system optimizes and can become a durable vertical advantage. | Evidence: Jiang compares an agent's evaluation to the loss function and data used in model training, or to the environment in reinforcement learning. He argues that proprietary evaluation data and domain-specific knowledge of what matters can form a vertical moat. | Implication: Investment in domain-specific eval suites, private feedback data, and objective design should be treated as core product infrastructure rather than as final-stage testing. | Caveat: A strong evaluation still captures only what its designers can specify and measure; an incomplete target can amplify the wrong behavior as agent capability increases.
- Claim: Codebase abstraction strongly biases an agent's search and can prevent apparently successful but invalid solutions. | Evidence: In a fraud-detection experiment, a loose API let one function process both training and test data, producing strong scores contaminated by test-set leakage. A stricter API that prevented test data from reaching training reduced the leakage rate to zero. | Implication: Agent-ready repositories need capability boundaries, typed interfaces, data isolation, and invariant checks designed before autonomous optimization begins. | Caveat: The speaker acknowledges that an agent may still reward-hack even under a better abstraction, so structural constraints reduce risk without eliminating it.
- Claim: Auto research shifts human labor toward higher-level system design rather than eliminating the engineer. | Evidence: Jiang compares the transition to deep learning replacing handwritten rules without eliminating software engineering. He argues that humans will increasingly design evaluations, abstractions, and competitions while agents perform repetitive search and implementation. | Implication: The scarce skill becomes the ability to construct a productive search environment and exercise judgment over objectives, constraints, and accepted outputs.
Detailed Brief
Aden's research-to-PR operating loop
- Claims: Aden is described as a multi-agent, self-improving system that reads public research material, conducts experiments, and publishes its findings as pull requests.; Publication is conditional on findings passing an internal quality gate, making the public artifact—not merely the internal experiment—the unit of output.
- Evidence: Its source material includes research papers, other contributors' pull requests, and informal community notes.; Weco positioned this as the next step beyond its earlier machine-learning-engineering agent evaluated in OpenAI's MLE-bench paper.
- Caveats: The talk does not disclose the agent topology, model choices, quality-gate criteria, cost, human review load, or degree of operational intervention during the 22-day run.
- Implications: The most relevant replication target is an end-to-end evidence-to-change pipeline with reviewable artifacts, rather than a general-purpose agent that merely returns recommendations.
Why competition and repository design retain human leverage
- Claims: The person or organization defining the task can have more leverage than participants conducting individual experiments because a poorly designed challenge can render all downstream optimization useless.; Repository architecture establishes the agent's practical hypothesis space by making some modifications easy, difficult, or inaccessible.
- Evidence: Jiang describes the competition's evaluation design as tremendously important even though it does not appear as a leaderboard contribution.; He maps codebase abstraction to neural-network architecture: multiple architectures may theoretically represent the same function, but each creates different optimization biases.
- Caveats: The analogy to model training is useful but imperfect because software agents can inspect, reinterpret, or circumvent interfaces in ways ordinary gradient descent cannot.
- Implications: Control-plane design, repository structure, and task specification are likely to become strategic engineering disciplines as autonomous search becomes commoditized.
Notable Concepts & Terms
- Parameter Golf: OpenAI's constrained model-training challenge in which Aden's submissions were externally reviewed and compared with a large human participant pool.
- Aden: Weco's multi-agent auto-research system that reads public knowledge, executes experiments, applies a quality gate, and submits pull requests.
- PR H-index: An adaptation of the academic H-index used to estimate how often a contributor's pull requests influenced or were reused by other competition entries.
- Evaluation as loss function: The framing that an agent's eval determines its optimization target in the same way that data and a loss function shape model training.
- Codebase abstraction as architecture: The idea that APIs and repository structure bias an agent's search toward certain classes of solutions, just as a neural architecture shapes what a model learns easily.
- Reward hacking: An agent satisfying the measured objective through invalid shortcuts, illustrated by test-set leakage in the fraud-detection pipeline.
- Gated Attention: An idea sourced from a Qwen paper that Aden implemented and later combined with quantization and a tokenizer improvement.
- Designing the hill: Jiang's description of the emerging human craft: defining the environment, constraints, interfaces, and objective that an automated system will climb.
Operator Notes / Why Ken Should Care
- Require agent-generated changes to carry machine-readable provenance linking each hypothesis to papers, prior pull requests, comments, and experiments.
- Add a promotion gate that separates experiment success from permission to publish or merge, with reproducibility, contamination, regression, and policy checks.
- Maintain a searchable registry of failed and neutral experiments so future agents can test combinations instead of repeatedly rediscovering isolated dead ends.
- Red-team repository interfaces for leakage and objective shortcuts before allowing agents to run long autonomous optimization loops.
- Track accepted-output rate, downstream reuse, rollback frequency, and compute per accepted improvement alongside raw task scores.
- Do not infer general scientific autonomy from the leaderboard result until Weco discloses human intervention, operating cost, model stack, and quality-gate details.
Source/Metadata
- Title: An AI Agent Became the #1 Contributor in OpenAI's Hiring Challenge — Zhengyao Jiang, Weco
- Transcript words: 1798
- Duration seconds: 976
- Timestamp note: The supplied transcript contains no timestamps or chapter markers; the reported video duration is 16:16.
Transcript
[SPEAKER_00] This April, OpenAI ran a hiring challenge, a competition called Parameter Golf. The top contributor was one candidate that they couldn't hire. It wasn't a person. It's an agent we built called Aden. In Parameter Golf, the goal is to train the best language model you can under size and computation constraints. About 1,000 machine learning engineers and researchers participated. They filed 2,000 submissions. Only 47 passed OpenAI's review and made it onto the leaderboard. Seven of those are actually Aden's, more than twice what any human contributed. You've seen a lot of auto research today. Agents are here, climbing benchmarks. Those are really impressive results. The question I want to ask is a bit different here. Can the auto research agent produce work that a human community actually recognizes? Beyond a good score the agent is optimizing for, something that other engineers can merge, fork, and build on. So instead of having an agent just hill-climbing locally, we built one that publishes its own work, and that's Aden. Quick context on us. Weco is an auto research company that was founded about two and a half years ago. I'm co-founder and CEO Zheng Yao. I got my PhD at UCL in reinforcement learning. About two years ago, we built Aden, the top auto research agent independently evaluated by OpenAI in their MLE-bench paper, even though back then there was no such name as auto research. People called it a machine learning engineering agent. Aden is the next step in an experimental prototype. It's a multi-agent, self-improving system that can read public information, such as research papers and other PRs, run its own experiments, and submit a PR once the findings pass a quality gate. We sent Aden to the Parameter Golf competition, and it ran for about 22 days. By the end, Aden had set seven leaderboard records. Each one was the new best for the competition, stamped by OpenAI. And the best human only made three. Passing the host review is one signal of quality. A second, maybe more important one, is whether other participants would build on your work. And it turns out Aden's work had the highest impact within the whole community. Here we are using an influence measure that is widely used in academia. It's called an H-index. Roughly, if you have X papers that get cited X times, then your H-index is X. Computed over PRs, Aden's was 10, and the next human's was 7. The whole community was building on Aden's work, including many of the other leaderboard entries. To break it down a little bit, why can an autonomous AI system be so powerful? One obvious reason is that it's an AI; it can run tirelessly. Over 22 days, it ran about 1,300 experiments on a single H100 node. But throughput isn't the whole picture. A well-tuned AI system can also keep its output quality high. On the compute side, it used at most 4% of the competition's total compute. And it made about 15% of the records. Also, 28% of its submissions made the leaderboard, roughly a six-times-higher hit rate than the community average. So Aden actually lifted the signal-to-noise ratio within the whole community's public communication channel, which is PRs. It didn't win through massive parallelization, even though auto research has tons of potential for parallelization. By those numbers, it might feel as though auto research already dominates human experts in ML engineering and research, but that's not the full story I want to tell. Humans and AI actually contribute in very different ways. When we trace the ideas in Aden's record PRs, almost all of them come from humans: research papers, other participants in Parameter Golf, or similar communities such as nanoGPT. Those ideas are not necessarily in a merged PR. Sometimes it's a note—a human researcher said, "Oh, I gave up this idea because of some implementation difficulty." And the agent is good at finding them and actually implementing them. There are also a very small fraction of original ideas that Aden came up with by itself, which emerged from its efforts to navigate the file-size constraints. Here's a concrete example that traces the patterns I just talked about. Aden picked up an idea from a Qwen paper called Gated Attention, and it worked. But it introduced more parameters, and it broke the 16-megabyte file-size limit. So it figured out a quantization mechanism to bring the file size down. But with those two primitives combined, the score barely moved. Then another contributor posted a tokenizer improvement. Aden recognized the idea, combined it with the architectural work, and it worked for five days or so. And after this combination, the three ideas turned out to have a huge synergy. That led to a big jump in performance, and they became one of Aden's leaderboard records. So to sum up how I interpret Aden's effectiveness and, in general, the effectiveness of auto research systems, it's very strong at finding and implementing ideas. In the case we just saw, it brought an idea from a recent paper into an actual implementation in the competition. And it's good at finding promising ingredients from the Parameter Golf community, even though the public channel is actually very noisy information-wise. It can also come up with logically straightforward ideas. For example, in this case, once you add the parameters and it breaks the file-size limit, one obvious next move is just quantization. And it's really fast and efficient at finding the right combinations across a huge search space. Okay, maybe none of those sounds very sexy. Most of them are just good execution. But in reality, execution is mostly the bottleneck. What moves the frontier is usually exactly some belief in existing ideas and tons of good execution. So the state of human-AI collaboration is that humans collectively provide a lot of creative ideas, and the agent does the execution to solve a concrete challenge. What we are looking at is a large group of humans and one AI system. Does it mean a single human engineer's contribution marginally gets smaller? I'd say even for that, not really. In the Parameter Golf competition, it's easy to focus only on engineers who are actually doing hill-climbing. But the design behind the competition itself is tremendously important. A bad design can make the whole community's effort useless. And their eval design work will have huge leverage in the auto research era. I really like one tweet from Andrej Karpathy about ten years ago, where he said, "Gradient descent can write code better than you. I'm sorry." For context, about ten years ago, deep learning was starting to eat up a lot of software engineering, such as conventional coding work. And his tweet was arguing against those people who thought they could handwrite better code than a trained model. Okay, now obviously no one is seriously trying to handwrite code to beat a model. However, software engineering as a job still exists. And so many people's jobs are just training those models, and those are some of the most well-paid jobs today. I think how gradient descent changes coding is a great metaphor for how auto research will change research in ML engineering. It accommodates certain execution skills. At the same time, it makes some higher-level skills far more valuable. So actually doing auto research is a lot like training a model. Your codebase abstraction is essentially the architecture. It sets the constraints and priorities for what the agent can explore. Your eval is the loss function and the data. It sets what the agent optimizes for. Take the eval first. The eval is the signal you use to train a model. In this case, it's training your code. It plays the same role as data and the loss function in model training. Or, in a reinforcement learning setting, it's like an environment in which the agent is training. Nowadays, no one would argue that data or environments don't matter. And this is where a vertical moat can also be built. You might have proprietary data for evaluation or a unique understanding of a particular field—of what matters and how to measure it. And a good evaluation will be amplified more and more as auto research gets stronger. The other one I think is really underrated is codebase abstraction. The abstraction provides the framework that auto research can iterate on. And that starting point hugely biases the whole search direction. The abstraction is a lot like architecture design in neural networks. Different architectures, in theory, can represent the same function. But the architecture systematically makes some of the functions easier to learn. And a good architecture biases the optimization toward solutions that generalize better and perform better, even when the training loss might look the same. That's exactly the same for auto research. Here's an example. We ran auto research for a fraud detection pipeline. And we ran auto research for a lot of data processing. First, we gave it a loose API where the same function processed both the training and testing data. And the score looked great. But the solution was polluted because certain test-set information got leaked into the training information. We then tightened the abstraction to a stricter API where the test data couldn't reach the training. And the data leakage rate dropped to zero. In this case, a good abstraction leads to better solutions, even though, if the agent really wants, it can still reward-hack. So my point is that using auto research is a new craft. It's about designing a hill for an agent to climb. And we are still very early in it. I think that makes this an extremely exciting time to be an AI engineer. Auto research will change what skills matter most: creativity and the judgment to design a good eval or an abstraction. Those will soon get exponentially more important. Driving those systems itself will be a new skill. And that one barely existed one or two years ago. So the search is automated. The human will just move up the stack, not out of it. So the user is a new product. Again, Weco is an auto research product research lab. We keep sharing what we are learning as we build on our blog. And I will also post some of my thinking on X. If you think some of this is useful to you, feel free to follow me on X. Thank you. I want to invite us from now. I want to invite us from now. I want us from now. Thank you.