AI Engineer

Multiplayer agentic engineering — Arjun Singh, Superconductor

2075 summary words 9 min summary Watch video

Start with the signal

9 min read

Summary

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: Effective agentic engineering is a multiplayer operating model: run agents in isolated cloud environments, preserve a shared session across the team’s existing interfaces, convert external signals into executable work, and continuously benchmark models on the actual codebase.
  • Why it matters: This is a concrete control-plane pattern for scaling coding agents beyond individual developers without locking the organization into one model, losing human review, or giving autonomous agents unsafe access to developer machines.
  • Best use: Use it as an architecture and operating-model reference for designing collaborative coding-agent workflows, especially cross-functional intake, sandboxing, model routing, and evaluation.

Executive Summary

Arjun Singh argues against the laptop-centric view of coding agents. His central premise is that agents should augment a team’s existing collaborative system rather than become isolated personal tools: a single agent session should be addressable from Slack, an engineering app, GitHub, and other relevant surfaces, with shared context, visible artifacts, and clear visibility into which humans have participated or reviewed the work.

Superconductor’s most distinctive workflow is turning external organizational signals into code. Its meeting bot can join long Meet, Zoom, or Teams calls, recognize a product request or issue, match it against existing work, create a ticket, and begin implementation. Singh’s example was a four-hour expo-booth meeting in which a request for explicit acceptance criteria was turned into a ticket, a live modification to the ticket form, and a screenshot for review. The intended outcome is not automatic shipment of every idea, but rapid conversion of scattered feedback into concrete, inspectable prototypes and PRs.

The enabling layer is an isolated cloud development environment. Singh frames this as both a productivity requirement—people can close their laptops while agents continue—and a security boundary. Instead of relying on permissive local-agent modes that may find unintended credentials or files, agents should receive only scoped credentials and network access, with requests for new destinations surfaced for approval. This also lets support and growth staff initiate real engineering work without needing local development environments.

Finally, Singh recommends model and harness agnosticism backed by codebase-specific benchmarks rather than public benchmarks. His team measures quality against cost and elapsed time using representative PRs, then changes defaults as the frontier changes. Superconductor reports that 99.9% of its PRs are heavily agent-generated but still human-reviewed; it used 10.5 billion tokens in one month. The talk is product-led and the performance results are self-reported, but the operational lessons—shared state, sandboxing, local evaluation, and review gates—are broadly reusable.

Key Takeaways

  • Claim: A coding agent should be a persistent shared session across the team’s work surfaces, not a task trapped on one developer laptop or in one chat tool. | Evidence: Singh describes starting a session in Slack, continuing it in a desktop or mobile engineering environment, and finishing in GitHub while retaining the same session context. Team members can ask the agent why an implementation choice was made instead of waiting for the original developer to answer. | Implication: A team agent platform needs durable session identity, cross-channel state synchronization, permissions, participant visibility, and artifact links—not merely separate Slack and GitHub bots.
  • Claim: Making agent work observable and collaborative is essential when work can be initiated by non-engineers. | Evidence: The product view shown tracks who created, viewed, and is notified about a work session; Singh uses this to determine whether a support-generated task has been vetted by engineering. Agents also publish screenshots, videos, and other artifacts so work can be inspected from any interface. | Implication: Cross-functional agent intake should carry provenance, current reviewers, generated evidence, and a clear human owner so that fast initiation does not produce unaccountable code changes. | Caveat: Visibility improves coordination but does not substitute for a defined approval and merge process.
  • Claim: External signals such as calls, support conversations, bug reports, and feature requests should be ingested, deduplicated, prioritized, and turned into executable code rather than manually copied into a backlog. | Evidence: A meeting bot listened to a four-hour Google Meet at Superconductor’s expo booth, linked ideas to existing work where applicable, and turned a request for acceptance criteria into a ticket and a working UI modification. Singh says onboarding, customer, and internal calls commonly yield dozens of prototypes and at least some shippable PRs with minimal intervention. | Implication: The high-leverage workflow is to make qualitative customer and internal feedback executable quickly, while retaining product judgment through preview, evaluation criteria, and review before release. | Caveat: Singh explicitly says he would probably not ship the demonstrated acceptance-criteria change exactly as generated; conversion to code creates a reviewable starting point, not an automatic product decision.
  • Claim: Cloud-isolated environments are a prerequisite for both durable multi-user agent workflows and safe autonomy. | Evidence: Singh contrasts cloud execution with agents running on laptops that may contain unintended files, tokens, and production credentials. He gives a failure scenario in which an agent instructed to wipe staging discovers a token and mistakenly accesses production. His proposed control is scoped network allowlisting with per-ticket or project-level approval for new destinations. | Implication: Treat agent execution as an isolated workload with least-privilege credentials, explicit egress controls, and auditable access escalation—not as a trusted extension of every developer’s workstation. | Caveat: Sandboxing reduces exposure but requires accurate environment setup, credential scoping, and network policy maintenance; it is not a complete security program by itself.
  • Claim: Public agent benchmarks are insufficient for selecting models and harnesses; teams should evaluate them against representative work in their own repositories. | Evidence: Singh notes that SWE-bench is Python-centric while Superconductor uses Ruby on Rails. His team selects pull requests representing strong engineering work—human, agent, or hybrid—and compares agent quality against cost and time on its own codebase. | Implication: Build an internal evaluation suite from real tasks and use it to govern defaults, procurement, and task routing; vendor leaderboard performance should be treated as a hypothesis, not a deployment decision. | Caveat: The comparative results Singh presents—such as Anthropic being more expensive and Codex being cheaper for his team—are explicitly specific to Superconductor’s codebase and should not be generalized.
  • Claim: Model and harness agnosticism is an operating advantage because the cost-quality-speed frontier shifts rapidly. | Evidence: Superconductor changed its default to Codex after its internal measurements, temporarily switched to Fable when it appeared, then reverted when Fable became unavailable. Singh also reports satisfaction with GLM 5.2 as a cheaper open-weight option. In one month, the team ran 3,300 Claude Code jobs with roughly $10,000 in token value, while Codex had four times as many sessions and was cheaper overall. | Implication: Avoid workflow designs coupled to one vendor’s interface or session format; preserve a common execution and evaluation layer so a newly available model can be tested and adopted without disrupting delivery. | Caveat: The stated token costs, session counts, and model comparisons are self-reported and lack task-level distributions or independent validation.
  • Claim: Very high agent-generated code volume can coexist with human accountability if review remains mandatory. | Evidence: Singh says 99.9% of Superconductor pull requests are heavily agent-generated and that the team used 10.5 billion tokens in the prior month, but states that every change is still human-reviewed because quality, reliability, and security matter. | Implication: Measure autonomy by the speed and quality of reviewed, merged changes—not raw agent runs or token consumption—and preserve human merge authority even as generation becomes pervasive. | Caveat: The talk does not specify review-depth standards, automated test gates, rollback practices, or defect outcomes, so the claimed production reliability model is incomplete.

Detailed Brief

Workflow design: from asynchronous ideas to inspectable implementation

  • Claims: The bottleneck in agent adoption is often coordination rather than raw model capability: information remains fragmented across systems, and humans still have to decide which email, ticket, or conversation an agent should act on.; Artifacts are a cross-surface collaboration primitive because they let stakeholders inspect progress without locating the original application or reading an entire agent transcript.; The speaker views nontechnical teams as meaningful product contributors when they can request and inspect a concrete change directly, rather than merely file an item for later triage.
  • Evidence: Examples of external signals include Slack conversations, customer onboarding and sales calls, internal meetings, Sentry alerts, bug trackers, customer reports, emails, and feature requests.; The demonstrated flow produced a screenshot of a modified ticket form containing new acceptance-criteria fields, allowing the presenter to test the concept in a live preview.; Singh contrasts the direct-request workflow with the traditional path of putting an issue in Linear, awaiting PM triage, and eventually assigning it to engineering.
  • Caveats: A system that automatically creates implementation work from every conversation needs strong deduplication, priority controls, and ownership rules; the talk says these exist but does not explain their policy or accuracy.; Meeting transcription and automated work extraction may introduce privacy, consent, retention, and false-positive concerns that the talk does not address.
  • Implications: Design the intake system around evidence and reversible prototypes, not only text tickets.; Use generated artifacts as the review interface for product, support, growth, and engineering stakeholders with different technical depth.

Benchmarking as the basis for task-aware routing

  • Claims: The next step beyond repository-level benchmarking is automatic model routing by task type, using measured performance on the team’s own code rather than a third party’s generic routing assumptions.; The value of benchmarking is not just choosing a permanent winner; it eliminates the recurring cost of manually testing each newly hyped model.
  • Evidence: Singh describes separate quality-versus-cost and quality-versus-time views across multiple harnesses.; His qualitative observations were that Anthropic agents had improved consistently but were not faster, Codex and Cursor were fast and strong, and open models were improving but slower on Superconductor’s workload.
  • Caveats: Task-aware routing requires a sufficiently broad, current, and representative evaluation corpus; PRs selected as 'great engineering work' can bias results if they omit routine, ambiguous, security-sensitive, or operational tasks.
  • Implications: Maintain benchmark freshness as frameworks, repositories, and models evolve; otherwise routing will optimize for obsolete work.; Record task attributes and outcome data now if future automatic routing is a goal.

Notable Concepts & Terms

  • Multiplayer agentic engineering: Singh’s term for organizing humans and coding agents as a shared team workflow rather than individual developers using private agent sessions.
  • Model- and harness-agnostic: The ability to change both underlying models and the agent execution interface without disrupting the team’s delivery workflow.
  • Same agent session: Persistent context and identity that can be accessed from Slack, GitHub, and dedicated applications, avoiding context copying and agent amnesia.
  • External signal to code: A pipeline that turns conversations, customer feedback, operational alerts, and requests into linked tickets, prototypes, and potentially pull requests.
  • Meeting bot: A bot invited to video meetings that identifies actionable ideas, links duplicates to existing work, and triggers implementation workflows.
  • Lit anxiety: The concern of needing to keep a laptop open and connected while a local coding agent runs; cloud execution removes this dependency.
  • Configurable network sandbox: An egress-control mechanism that allows agents to access only approved destinations, with scoped approvals for additional access.
  • Codebase-specific benchmarking: Evaluating models and harnesses on representative pull requests from the organization’s own stack to compare quality, cost, and time.

Operator Notes / Why Ken Should Care

  • Audit whether current coding-agent sessions can be resumed and inspected across the interfaces where engineering, support, product, and GTM teams already work; prioritize persistent session IDs and artifact links over adding another standalone bot.
  • Require an agent execution baseline of ephemeral cloud environments, least-privilege credentials, controlled egress, and per-task access escalation before allowing broad autonomous execution.
  • Create a repository-native evaluation set from completed PRs, segmented by task type and risk, then score candidate model/harness pairs on completion quality, cost, and latency.
  • Define a conversion policy for customer calls and support signals: deduplication threshold, product owner, required acceptance criteria, preview/review gates, and conditions under which an agent may open versus merge a PR.
  • Keep human merge authority and instrument downstream quality outcomes—test failures, reverts, incidents, and review time—so agent adoption is measured by production results rather than token volume.

Source/Metadata

  • Title: Multiplayer agentic engineering — Arjun Singh, Superconductor
  • Transcript words: 4991
  • Duration seconds: 1124
  • Timestamp note: No usable timestamps or chapter markers were present in the supplied transcript; the latter portion contains duplicated benchmark and closing content.
Full transcript 3672 words · 22 min read
0:00

Music Alright, hey everyone, I'm Orjan Singh. Today I'm going to talk to you about Multiplayer Agenting Engineering, or how to enable your whole team and your best agents to work together. If you go to the talks or go around the expo, you're going to see that a lot of people are talking about putting the agents at the center of everything. It makes sense. They're really powerful, they're really cool. But you don't see a lot of people talking about the people. This is all for us, to make our lives better, more productive, whatever. And so we're going to really focus on how the people fit into these agentic workflows.

0:21

So just a little bit about us first. Our team has worked together building software for over a decade. My co-founder Sergey and I met in the PhD program at Berkeley. I worked on robotics, we worked on computer vision. And during that, we co-founded a company called Gradescope. Some of you may have used it. It's used by millions of students worldwide at thousands of universities. It helped instructors grade their students' work. And pretty much the entire team working on Superconductor used to work together on Gradescope. And so we've had a team that's worked together productively from first user to acquisition, working on something new together again. And I think it's an interesting experiment because, over the past year, we've very aggressively integrated agents into our workflows. And we've surfaced all the different bottlenecks and friction points that come up, and how to do that productively and keep collaborating the way we used to, but with the new power of agents.

0:23

So today I'm going to talk to you about the lessons we learned from solving those friction points and bottlenecks. In the talk description, I mentioned five lessons. And I'm going to be an engineer and start from zero and add a sixth one in there. The first one I'm going to start with is just to be model- and harness-agnostic. So there's a few reasons for that. The best model and harness can change weekly. It could change because a new one comes out. It could change because the best one got taken away. Things happen. And you don't want that to disrupt your entire team's flow.

0:27

The other thing is that open weight models are actually pretty good now. We've been really happy with GLM 5.2. They're much cheaper, and you want to be able to explore with them and integrate them without, again, having to change your entire workflow.

0:29

The last thing I'll mention on this is that the incentives of the people selling you tokens aren't really aligned with yours. You're here for a reason. You're working on things for a reason. You're trying to make your customers' lives better, make your product better, delight your customers. And they want to sell you more tokens. And you might be happy to pay for as many tokens as it takes, but you don't want to pay for more than that. And so, again, being able to switch between things lets you stay in control of all of that.

0:32

As I go through the talk, I'll mention a couple of places where our product makes it easy for us. But whether you use us or not, I'm just going to leave things with you that I think are really important for you to be able to work collaboratively and effectively.

0:36

So the next one is to turn every human interface into an agent and human interface. Typically, when people are working with coding agents, they're on their laptop, stuck on that laptop. Nobody else can talk to that agent. So the first place people go to expose more interfaces for them is Slack. Cloud has a Slack bot. Codex has a Slack bot. We have a Slack bot. It's really cool. You can say, add superconductor, do X, Y, Z. It does it. Somebody else can talk to it. But it's not enough. Because now we've taken it from trapped on somebody's laptop to trapped in Slack. And a lot of work happens in Slack, so that's better than nothing. But certainly not all work happens in Slack.

0:41

So what we really wanted was to be able to work with the same session from every relevant interface. It could be Slack. It could be our app. It could be GitHub. It could be elsewhere. And so one possible flow is you start and collaborate on a session in Slack, and then maybe you continue in a more engineer-focused environment in the desktop app or the mobile app, and then you can finish it up in GitHub. And the important thing here is the exact same agent session. So it's like the agent didn't forget what you did in one place in Slack when you go and talk to it from GitHub. It's the same session. It's got the same context.

0:44

And the second lesson builds on top of that, which is to make the agent work visible and collaborative across the team. And so obviously Slack makes it more collaborative. But here we've got that app view, and Sergey made this ticket. I've been talking to the same ticket. A growth person hopped in as well. And so you can see at the top here all the different people that interacted with it. So I can see who's getting notified about this session, who's seen it. That's especially important when you have work triggered by non-technical people. So it's like our customer support person created a ticket. It's working really well. I want to understand, has this been vetted by an engineer or not, and see who's involved really easily.

0:48

And then if I'm reviewing something, I can just pop in and say, hey, why did you do it this way? And again, because it's the same agent session, I don't need to wait for Sergey to get my notification on GitHub and respond to me. The answer to the question is almost certainly in this thread. I also don't want to read the entire thread, so I can just ask the agent.

0:51

Or how we most often make the work visible is with artifacts. So it doesn't matter where the work started or where it's finishing. The agent can show you the work it's doing as a screenshot or video or other, and you can see it from everywhere. So again, you don't have to worry about, oh, where is that thing? I got to go to GitHub to see the image or I got to go to Slack to see the image. It's just everywhere. Work is visible everywhere. You can collaborate from anywhere.

0:54

So the third lesson I'm going to talk about here is to turn every external signal into code that your team can quickly evaluate. And I was hoping to show this live, but the Wi-Fi is not quite there. So I'm going to show you something from yesterday.

0:56

But what do I mean by external signal? It could be a Slack conversation. It could be a meeting you have with a customer, an onboarding call, a sales call. It could be an internal team meeting. It could be something from Sentry or a bug tracker, a bug report from a customer, an email, a feature request. And right now what's happening is all that stuff already exists. It's in all those different systems. People hook them together with MCPs. So now your coding agent can check the email or check Notion or whatever it might be. But how does it know what to work on? It's still stuck everywhere. And so some humans are involved in taking stuff from one place and telling it to solve email number 48 or ticket number 6,000. But that's still a lot of coordination.

0:58

And so what we do is we have several different ways to automatically ingest these signals, prioritize what to do with them, and act on them. And my favorite one, the most fun one, is what we call our meeting bot. And so I'm going to switch over to my browser here for a second. And, okay.

1:04

So we've got a booth at the expo, and we had the meeting bot running all day yesterday. So this is a four-hour meeting in a Google Meet. You just invite the bot to Meet or Zoom or Teams or whatever it might be, and it listens all day. And it created all sorts of stuff as it was listening. If it finds existing work, it'll link to it. So it's not going to create new work if it's something you're already working on.

1:09

Some of this is people testing the meeting bot out and telling it to do some weird things or interesting things or just creative ideas. But a lot of it is actually really good ideas that come out of people looking at what we're doing, asking questions, having new ideas on what to do with it. And so it's nice because the last idea that was here was saying, hey, when I work with coding agents, I want to make sure that the agent has clear criteria to evaluate whether it did a good job on the work before it tells me that it's done. And so they had that idea. The bot just picked up on it. None of us did anything manually. And it created all sorts of stuff as it was listening.

1:15

If it finds existing work, it'll link to it. Right? So it's not going to just create new work if it's something you're already working on. Some of this is people testing the meeting bot out and telling you to do some weird things or interesting things or just creative ideas. But a lot of it is actually just really good ideas that come out of people looking at what we're doing, asking questions, having new ideas on what to do with it. And so it's nice because the last idea that was here was saying, hey, when I work with coding agents, I want to make sure that the agent has clear criteria to evaluate whether it did a good job on the work before it tells me that it's done.

1:40

And so they had that idea. The bot just picked up on it. None of us did anything manually. They created this ticket. And started working on it. And then I was able to just say, hey, take a screenshot of what you did. And here's that screenshot. And it modified our ticket form to add these two new fields of acceptance criteria. Now, am I going to ship this one exactly how it is? No, probably not. But it's a new idea. It's concrete. I can play with it. I can go and actually use the live preview and see if this improves performance.

2:24

And so it takes your hundreds or thousands of ideas that are everywhere, and it helps you move with the speed of what your customers are asking you for and what they're thinking. And it's really fun because every time we have an onboarding or customer call or team meeting, we almost always have dozens of new ideas that are prototyped. But more importantly, at least a few shippable PRs with very minimal intervention. So we talk. Stuff comes out. We look at it. We ship it. It's so much fun. Put this back.

3:09

So the next thing I'm going to mention is that these three things that I've talked to you about really rely on having your workflow, your code base, and your project set up to work in an isolated cloud environment. So that way, the agents aren't trapped on an individual's machine. So there are several reasons why this is important. So the first one is to eliminate what some people are calling lit anxiety. You want to be able to close your laptop. You've probably seen people running around the conference with their laptops open while stuff is working, or at the airport.

3:20

Or there are some posts on Twitter or whatever about people having their laptop tethered to their phone in their car as they're driving home. This was actually probably the impetus for me and for a few people on our team to even start working on this. Last year, I started working with Cloud Code a lot. I had, I think at the time, a six-month-old. I didn't want to be tied to my laptop or have that stress. It's like I don't ever want to think about whether I can step away from a laptop or not. And so we moved everything to the cloud. Things are working always. It eliminated that problem for us that people have been talking about for the past year. It's really helpful.

3:43

It's important to me. But I don't think that's the most important reason to do this. I think the most important reason to do this actually was touched on in the previous talk, if you were here for it. I think you should only give your agents access to what they need. Right, so if you think about what's happening, you have a bunch of developers with these agents running on their laptop. Their laptops, unless you have impeccable hygiene, probably have a bunch of stuff on them that you don't want the LLMs or the agents to have access to. And, yeah, everybody's working on these sandboxes and approval flows. But really you're in one of two camps.

4:07

You're either approving a bunch of stuff, or you're hoping that your auto-approval flow or your YOLO mode or whatever is configured properly, and your sandbox is configured properly and doesn't read a bunch of stuff on your laptop that it shouldn't have. And as the previous talk mentioned, these agents are getting more autonomous. They're getting really resourceful. They're trying to please you and do what you said. And so when you say, hey, wipe the staging database, and it finds a token on your laptop that it can use, and it thinks it's working with staging, but actually it's production, and now it just deleted everything.

4:28

I'm not trying to say this is happening constantly, but it still happens. And for us, the peace of mind of just letting anybody run with these experiments and ideas and prototypes and real code without having to worry about this is really worthwhile. To go one step further on that, it's not just, hey, make sure they don't have the credentials that they shouldn't have. It's also making sure they can't exfiltrate your code or your projects or your secrets or your content to somewhere they shouldn't be able to. And so you have a configurable network sandbox and you say, look, these are the places you're allowed to access. These are the ones you can't access.

4:43

And any time it tries to access something that it shouldn't, it just pops up and says, hey, I tried to access something. Do you want to give it access? Maybe you're trying to integrate a new vendor and need documentation. And you can do it on a per-ticket basis or for the whole project. And so again, that peace of mind of people can do things. If they need new access, it's easy to grant it. And we're not going to leak a bunch of important data by running agents in YOLO mode. And the last thing I'll mention about that is that this is the key for allowing your non-technical team members to trigger real work, right?

5:14

Non-technical people don't have development environments set up on their computers. But we've gotten our support people or growth people to actually meaningfully impact the product by just talking to users, seeing bugs, experiencing them themselves, and just going to Slack or the app itself and saying, hey, fix this. They fix it. Screenshots are shown. An engineer gets it. It gets merged. Without that, they'd have to put it in Linear and eventually pick it up, and a PM would triage it or whatever. None of that here.

5:34

You just ask for it and it's done. Now, the reason people didn't do this until somewhat recently is that this was really painful. Getting your full thing set up in this sandbox environment used to be really, really painful. But the agents have gotten better. We have our own environment setup assistant that takes your project and gets it to work in one of these sandboxes. But honestly, whether you use us or not, I highly recommend you get your project working this way. And you can just get Cloud Coder Codex to do this for you. You don't have to use us. But we think it's the best way. And the last lesson is to benchmark agents on your code base.

6:11

So the way we do this is we select pull requests that represent great engineering work. It could be agent-created. It could be human-created. It could be a hybrid. It doesn't matter. You pick the agents you want to use and benchmark. And then you get a quality versus cost and time breakdown in your code base. Now, why do you want to do this? There are many reasons. But one is that if you're going off the public benchmark, sweetbench or terminal bench or other stuff, those tasks may have absolutely nothing to do with your tasks. Sweetbench is all in Python. We're Ruby on Rails. It is not the case that the benchmarks are identical for them. There are trends that do compare.

7:06

But the results can be very, very different. And I'm going to swap over to my browser one more time here. So these are results on our code base of all the different harnesses. This one here is quality versus cost. This one here is quality versus time. I'll start with this one. You can see some trends here. You can see that the Anthropic agents have just been consistently getting better, but not really any faster. The Codex agents and Cursor are actually pretty fast and quite good. The open stuff has been getting better and better over time, but they're kind of slow. This is for our code base again. I'm not trying to make any general claims here.

7:53

By cost, the Anthropic stuff is clearly just so much more expensive for us. And the Codex stuff has been cheaper for us. And so this causes a change of behavior. We still use the different models. There are different use cases for them. We like the variety. We still use all these things. But when we saw these results, they matched our vibe check. We wanted to have hard data too. We switched our default to Codex at that time. Then Fable came out. It was great. I'll start with this one. You can see some trends here. You can see that the Anthropic agents have just been consistently getting better, but not really any faster.

8:40

The Codex agents and Cursor are actually pretty fast and quite good. The open stuff has been getting better and better over time, but they're slow. This is for our code base again.

8:45

I'm not trying to make any general claims here.

8:49

By cost, the Anthropic stuff is clearly just so much more expensive for us. And the Codex stuff has been cheaper for us. And so this causes a change of behavior. We still use the different models. There's different use cases for them. We like the variety. We still use all these things. But when we saw these results, they matched our vibe check. We wanted to have hard data too. We switched our default to Codex at that time. Then Fable came out. It was great. Switched our default to that for the few days we had it. And then it went away, and switched back to Codex.

9:54

But the most important thing is because we're agnostic, none of that had any meaningful disruption on our work. We were able to just switch back and forth really easily. So the next day something new comes out. See if it's good. And go. And the last thing I want to mention around that is, I don't know if this resonates with you all, but a lot of our friends are like, okay, I heard Minimax is good. I heard GLM is good and KMK2 is good. But haven't had the time to try it out. And they keep telling me I need to because it's so much better and faster and cheaper. And you have that anxiety for a little while. And then finally you take the two hours to try it.

10:44

And it's like, oh, actually didn't really work for us. So I just wasted those two hours. This eliminates that. It helps you stay on the cutting edge really seamlessly. Let me go back. So what that turned into for us is essentially 100%, like 99.9%, of our pull requests are heavily agent generated. We know that quality and reliability and security are really important. So we still have humans look at everything. We have agents help with it all, but everything's human reviewed. For our relatively small team, we had 10 and a half billion tokens over the past month. And you can see what we were saying about cloud here. It's a little small, so I apologize.

11:38

But we had 3,300 quad code runs that cost $10,000 in token value. We have planes, so we didn't spend $10,000 on it. And Codex had four times as many sessions, and it was cheaper overall. And so again, the vast majority of our work currently is merged through Codex. We still use the other models. More and more is happening through GLM 5.2. You can invest in that. And one thing that we're really excited to do going forward with this benchmarking is automatically, you've probably heard about people routing tasks to the right models and all that. But how does some third party know what to route for your code base?

12:04

This is a way that you know what's going to work best for which task for your project, and we're going to automatically routing that for you.

12:13

So I'm going to leave you with a few recommendations. So first, get your code base and agents working in a sandbox. It unlocks a lot of different things, a lot of different workflows, everything I've talked about and more. Second, integrate agents into the relevant human interfaces so your team and your agents can work together and don't have to context switch and copy context back and forth. Obviously, we think SuperNectator is the best way to do it, but plenty of people are home rolling things, packing things together. Figure out how to make this happen because if not, the friction is just really high. And lastly, find a way to benchmark and become model agnostic.

12:42

So you're not tied to anybody and you can just constantly stay at that right part on the frontier of cost, speed, quality. So thank you so much. We've got a booth in the expo. Please feel free to come by. You can sign up at SuperNectator.com. Or you can email me with any questions at arjun.supernectator.com. I will be out in the back as well for any questions. Thanks so much. One last thing. At the booth, we're mentioning we're giving away, I'm back with Neo. If you are here, could you sign up for that? Just meet us outside and we will announce the winner. Thank you.

13:25

Thank you. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you. There's many reasons. But one is that if you're kind of going off the public benchmark, sweetbench or terminal bench or other stuff, like, those tasks may have absolutely nothing to do with your tasks. Like, sweetbench is all in Python. We're Ruby on Rails. It is not the case that the benchmarks are identical for them. There's trends that do compare. But the results can be very, very different. And I'm going to swap over to my browser one more time here.

14:11

So, this is, these are results on our code base of all the different harnesses. This one here is quality versus cost. This one here is quality versus time. I'll start with this one. You can see some trends here. You can see that the Anthropic agents have just been consistently getting better, but not really any faster. The Codex agents and Cursor are actually pretty fast and quite good. The open stuff has been getting better and better over time, but they're kind of slow. This is for our code base again. I'm not trying to make any general claims here. By cost, the Anthropic stuff is clearly just so much more expensive for us. And the Codex stuff has been cheaper for us.

14:51

And so this causes a change of behavior. We still use the different models. There's different use cases for them. We like the variety. We still use all these things. But when we saw these results, they kind of matched our vibe check. We wanted to kind of like have hard data too. We switched our default to Codex at that time. Then Fable came out. It was great. Kind of switched our default to that for like the few days we had it. And then went away and switched back to Codex. But the most important thing is like because we're agnostic, like none of that had any meaningful disruption on our work. Like we were able to just kind of switch back and forth really easily.

15:20

So the next day something new comes out. See if it's good. And go. And the last thing I want to mention around that is like, I don't know if this resonates with you all, but a lot of our friends are like, okay, I heard Minimax is good. I heard GLM is good and KMK2 is good. But like haven't had the time to try it out. And we keep telling me I need to because it's so much better and faster and cheaper. And you kind of have that anxiety for a little while. And then like finally you take the two hours to try it. And it's like, oh, actually like didn't really work for us. So like, you know, I just waste those two hours. This kind of eliminates that.

15:51

It helps you kind of stay on the cutting edge really like seamlessly.

15:57

Let me go back. So what that kind of turned into for us is, you know, essentially 100%, like 99.9% of our pull requests are like heavily agent generated. We know that quality and reliability and security are really important. So we still have humans look at everything. We have agents help with it all, but everything's human reviewed. You know, for our relatively small team, we had 10 and a half billion tokens over the past month. And you can kind of see what we were saying about cloud here. It's a little small, so I apologize. But we had 3,300 quad code runs that cost $10,000 in token value. We have planes, so we didn't spend $10,000 on it.

16:34

And Codex had four times as many sessions, and it was cheaper overall. And so again, the vast majority of our work currently is merged through Codex. We still use the other models. More and more is happening through GLM 5.2. You can invest in that. And one thing that we're really excited to do going forward with this benchmarking is automatically, like you've probably heard about people, you know, routing tasks to the right models and all that. But how does like some third party know what to route for your code base? Like this is a way that you know what's going to work best for which task for your project, and we're going to kind of automatically routing that for you.

17:12

So I'm going to leave you with a few recommendations. So first, get your code base and agents working in a sandbox. It unlocks a lot of different things, a lot of different workflows, everything I've talked about and more.

17:26

Second, integrate agents into the relevant human interfaces so your team and your agents can work together and don't have to like context switch and copy context back and forth. Obviously, we think SuperNectator is the best way to do it, but plenty of people are home rolling things, packing things together. Figure out how to make this happen because if not, the friction is just really high. And lastly, find a way to benchmark and become model agnostic. So you're not tied to anybody and you can just constantly stay at that right part on the frontier of cost, speed, quality. So thank you so much. We've got a booth in the expo. Please feel free to come by.

18:02

You can sign up at SuperNectator.com. Or you can email me with any questions at arjun.supernectator.com. I will be out in the back as well for any questions. Thanks so much.

18:16

One last thing. If you know, at the booth we're mentioning we're giving away, I'm back with Neo. If you are here, could you sign up for that? Just meet us outside and we will announce the winner. Thank you. Thank you. Thank you. Thank you.

18:37

Thank you. Thank you. Thank you. Thank you. Thank you. Thank you.

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note