Open Reader

Dark Factory: OpenClaw Ships Faster Than You Can Read the Diff — Vincent Koc, OpenClaw

completed 16:44 Jun 05, 2026 Watch on YouTube

Current Status

completed

Video ID

pmoDeA3RBZY

RAG / Chat

Enabled
Dark Factory: OpenClaw Ships Faster Than You Can Read the Diff — Vincent Koc, OpenClaw
Description

OpenClaw hit 3,000 commits in a single day. Vincent Koc's commit history shows exactly when he goes to sleep and when he wakes up. He and Peter Steinberger ran roughly 60 to 70 agents between them during the great refactor: 2,700 commits, close to a million lines of code changed, 82% of the core codebase touched in one night, plugin architecture shipped by morning. The talk covers how you actually manage this at scale: swim lanes of 15 to 20 parallel coding sessions organized by type, when to nuke a session versus let it run, and what he calls reading the reasoning tokens. The skill is not prompting. It is knowing when an agent is bullshitting you. 2025 was about token maxing. 2026 is about not wasting them. Speaker info: - https://x.com/vincent_koc

Summary

Generated by claude-sonnet-4-5

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: OpenClaw ships 800–3,000 commits/day by treating AI agents as factory workers, not magic—engineering discipline, parallel 'swim lanes,' .skills management, and human factory-manager intuition now trump raw token-burning.
  • Why it matters: This is the first practitioner-level blueprint for running 10–20 AI coding agents in parallel, proving hyper-velocity open-source development is real and will become the norm, forcing a shift from PR review culture to factory-manager taste and process.
  • Best use: Ken should watch to internalize the swim-lane pattern, .skills workflow, and the mindset shift from 'review code' to 'manage agents like staff'—critical for AI-ops GTM, agentic workflow pitches, and understanding how taste/process beat model quality at scale.

Executive Summary

Vincent Koc, core maintainer of OpenClaw (an open-source AI coding tool), describes shipping 800 commits/day as a team of 10–15 part-timers and hitting 3,000 commits/day himself during a 24-hour refactor sprint with Peter at NVIDIA. He argues this velocity—comparable to Anthropic building a C compiler, Spotify claiming no hand-written code, or Steve Yegge pushing 50 PRs/day solo—is not Ralph-style token-burning luck but deliberate 'factory management' engineering. The bottleneck has moved from human hands (weaver's loom analogy) to human taste and process design.

The core method: Vincent runs 5–20 parallel 'swim lanes' (separate Codex sessions, often in Git worktrees), each tackling CI fixes, feature work, bug triage, or P0s. He doesn't use plan/spec modes; he converses with agents, reads their 'reasoning tokens' like a Matrix operator, and kills sessions that waffle or bullshit—exactly as he'd manage human reports. His .skills (prompt templates, workflows) and .files are versioned, improved via Codex log review, and stored publicly. He hit GitHub rate limits, ran 70–80 active Git worktrees, and relies on overfitted unit tests as a green-light harness when refactoring 82% of the codebase (1M LOC, 2,700 commits) overnight.

The 'Great Refactor' story: During a sprint at NVIDIA, OpenClaw moved the entire messaging channel folder structure. Vincent and Peter ran ~15 foreground agents each (60–70 with sub-agents), completed a plugin-architecture refactor touching 82% of code, and barely passed tests at 1 a.m.—saved only by AI-generated overfitted unit tests that acted as a safety net. This proved tokens are cheap; compute, focus, and intuition are the real constraints. Vincent's commit graph literally maps his sleep schedule.

Vincent's thesis for 2026: Stop 'commit maxing' (Ralph looping); start 'bar looping' (opinionated, reward-driven loops). PR review culture breaks at this scale; instead, semantic graphs/vector embeddings de-duplicate 6,000+ PRs, evals run synthetic Slack channels to verify provider integrations, and engineers become vibe maintainers who say no to bloat. Soft skills—knowing when an agent is bullshitting, managing swim lanes, maintaining .skills—now matter more than model choice. He calls this the 'Agent Development Environment' and believes every eng team will adopt factory-manager workflows soon.

Key Takeaways

  • Claim: OpenClaw shipped 800 commits/day at peak (3,000 in Vincent's 24h sprint), touching 1M lines across 82% of the codebase overnight, proving hyper-velocity AI-driven development is production-ready. | Evidence: Vincent's GitHub commit graph shows continuous activity except during sleep; the refactor moved plugins, channels, and MS Teams/Slack integrations in ~2,700 commits; tests barely passed at 1 a.m. but held due to overfitted unit tests AI had generated. | Caveat: Vincent admits he 'vibed too hard' and thought he was Icarus; the process relies on heavy compute (70–80 Git worktrees crashed his machine), high tolerance for jank, and blind faith in the test harness—most teams lack this infrastructure and risk appetite. | Implication: For Ken: This velocity will become the norm, not the exception. Teams must design for agent-scale workflows now—PR review, manual merges, and human bottlenecks won't survive. Consider investing in tooling (Git worktree managers, eval harnesses, .skills frameworks) and training eng leaders to think like factory managers, not code reviewers. | Timestamp: 00:00–03:30
  • Claim: The swim-lane pattern—running 5–20 parallel Codex sessions, each assigned to CI, features, bugs, or P0s—is the core 'factory manager' workflow that scales human attention across dozens of agents. | Evidence: Vincent assigns swim lanes 1–2 to refactoring (low supervision, just commit), lanes 3–4 to Docker/channel features (conversational, investigative), lane 5 to triage new P0s using GitHub/Discord data. He reads agent 'reasoning tokens' like Matrix code to detect waffling/bullshit and kills sessions that feel off. | Caveat: This requires deep intuition built from 'sheer volume of token maxing' over a year; Vincent can 'feel' when an agent is lost, but new users won't have this sixth sense. Also, raw compute and mental bandwidth—not tokens—become the bottleneck. | Implication: Ken: The swim-lane mental model is the unlock for AI-ops at scale. Pitch this pattern to dev-tool startups, AI IDE vendors, and enterprise eng orgs. Invest in products that visualize/manage swim lanes, detect agent drift, and train managers to spot bullshit in agent outputs. Soft skills (managing agents like staff) are the new hard skills. | Timestamp: 08:20–10:15
  • Claim: AI-generated overfitted unit tests, though 'awful,' saved the Great Refactor by providing a green-light harness when 82% of the codebase was ripped apart; tokens are cheap, but harness trust is priceless. | Evidence: At 1 a.m., Vincent feared he'd flown too close to the sun—tests were red. But the overfitted tests eventually passed, proving the refactor worked. He calls this 'blind faith in the harness' and contrasts it with his day job (structured evals, telemetry) converging toward the same idea. | Caveat: Overfitted tests are brittle and don't catch logic bugs—they only confirm structure. If the tests had been sparse or wrong, the refactor would have shipped broken code. This works only if you're willing to fix in production or have downstream monitoring. | Implication: Ken: Harness quality—not model intelligence—is the gating factor for agent-driven dev. Companies building eval frameworks, test-generation tools, or synthetic data harnesses (like OpenClaw's fake Slack) are critical infrastructure bets. Also, this validates the 'vibe maintainer' role: someone who trusts process over paranoia. | Timestamp: 07:00–08:00
  • Claim: The .skills workflow—versioned prompt templates, improved via Codex log analysis, stored as .skills/.files in GitHub—is the 'Agent Development Environment' that replaces traditional dev tooling. | Evidence: Vincent shares .skills publicly (some private); uses tools like .skills.sh and GEPA (a skills gym he contributes to); instructs Codex to 'read the last two weeks of logs, improve this skill'; deploys updated skills into OpenClaw or personal projects. He treats skills like code: versioned, tested, iterated. | Caveat: Vincent doesn't explain how to measure skill quality, handle version conflicts across maintainers, or prevent skill drift when agents misinterpret prompts. This is still artisanal, not industrialized. | Implication: Ken: .skills are the new .dotfiles—expect a Cambrian explosion of skill marketplaces, version managers, and quality benchmarks. Invest in infra that makes skills portable (cross-agent, cross-IDE), testable (skills gym), and discoverable (semantic search, agent stores). This is the supply chain for agent work. | Timestamp: 11:30–12:45
  • Claim: 'Bar looping' (opinionated, reward-driven agent loops) will replace 'Ralph looping' (blind token-burning for hours); 2026 is about token efficiency, not token maxing. | Evidence: Vincent contrasts Ralph looping (give task, burn tokens 8–9 hours, hope) with a smarter approach: structured swim lanes, reward signals (test pass = commit), de-duplicated PR graphs (semantic embeddings, 706 edges/PR), and evals (synthetic Slack channels verify providers). | Caveat: Vincent doesn't define 'bar looping' rigorously—it's a vibe, not a spec. He also admits the current setup is janky (Git worktrees nuke his machine) and that most of the discipline comes from his people-management experience, not AI tooling. | Implication: Ken: The next wave of agent tools must embed reward/feedback loops (not just token budgets), structured task decomposition (swim lanes), and PR/issue de-duplication at the platform level. This is a wedge for YC-style infra startups. Also, 'vibe maintainer' is a real job title—expect demand for training/certification in agent management soft skills. | Timestamp: 05:00–06:30, 14:00–15:00
  • Claim: Most orgs secretly use autonomous agents at ChatGPT-era scale but won't admit it; Anthropic (C compiler), Spotify (no hand-written code), Yegge (50 PRs/day) are early public proof points. | Evidence: Vincent compares current agent secrecy to 2023 ChatGPT denial: 'Everyone in secret was just like, oh my God, what's going on?' Now Anthropic, Spotify, and solo devs openly ship agent-driven work at 10–100x human velocity. | Caveat: Vincent provides no data on adoption rates, failure modes, or how many orgs tried and failed. The named examples are outliers (elite teams, open-source passion projects, or well-resourced labs), not median companies. | Implication: Ken: If ChatGPT adoption (hidden → ubiquitous in 12 months) is the template, agent-driven dev will hit enterprise in 2025–2026. Pitch deck: 'We're in the denial phase; whoever builds infra now wins the next stack.' Also, this is signal for content: interview Steve Yegge, Anthropic compiler team, Spotify eng leads—document the playbooks before they become common knowledge. | Timestamp: 03:30–05:00

Detailed Brief

The Factory Manager Mental Model: Swim Lanes, Not Supervisors

  • Claims: Vincent runs 5–20 parallel Codex sessions ('swim lanes'), each assigned to a category: CI, features, bugs, P0 triage.; Swim lanes 1–2 handle low-touch refactoring (just commit, no babysitting); lanes 3–4 do conversational feature work; lane 5 triages new issues using GitHub/Discord agents.; Vincent reads agent outputs like Matrix code—'feeling' reasoning tokens to detect when an agent is waffling, lost, or bullshitting, then kills the session.; The bottleneck is no longer tokens or model quality; it's compute (Git worktrees) and human attention (factory manager brain space).
  • Evidence: Vincent's commit graph maps his sleep schedule; during the NVIDIA sprint, he and Peter ran ~15 foreground agents each, 60–70 total with sub-agents.; He adopted Git worktrees (70–80 active at once) but regrets it—crashed his machine; wishes he'd cloned the repo 10 times instead.; He doesn't use Codex plan/spec modes; he converses with agents, builds intuition over a year of token maxing.; During the Great Refactor, one maintainer moved the entire messaging channel folder, forcing all swim lanes to adapt mid-sprint.
  • Caveats: Requires infrastructure most teams lack: heavy test harness, compute to run 10+ sessions, tolerance for jank/crashes.; The 'feeling reasoning tokens' skill took Vincent a year of high-volume agent use to develop; new users won't have this intuition.; Factory-manager workflows assume eng leads have people-management experience (Vincent managed 30–40 people in airlines/AI teams); most ICs don't.
  • Implications: Ken should treat swim-lane orchestration as the new eng-manager skill set; expect training/cert programs, tools to visualize swim lanes, and demand for 'vibe maintainer' consultants.; Invest in infra: Git worktree managers, session recovery tools, agent drift detectors, and platforms that abstract swim-lane boilerplate.; Content play: Publish swim-lane playbooks, interview Vincent/Peter on failure modes, create benchmarks for agent bullshit detection.

The Great Refactor: 2,700 Commits, 1M LOC, 82% Codebase Touched Overnight

  • Claims: At NVIDIA, Vincent and Peter decided at 2 a.m. to refactor OpenClaw's entire architecture into a plugin system, touching 82% of the codebase.; The refactor moved messaging channels (MS Teams, Slack) to new folders, breaking all active work; agents adapted in real time.; At 1 a.m., tests were red; Vincent thought he 'vibed too hard' (flew too close to the sun). Overfitted AI-generated unit tests saved the launch by going green.; The refactor shipped 2,700 commits, close to 1M lines changed, enabling provider-owned plugin code (OpenAI, Mistral, Anthropic could own their integrations).
  • Evidence: Vincent's commit graph hit 3,000 commits/day during the sprint; Peter ran 15 Codex sessions, Vincent ran 10–15, totaling 60–70 agents with sub-agents.; NVIDIA provided external monitors; Vincent used his, Peter's Mac Studio was VPNed from home.; The overfitted unit tests (which AI loves to generate) acted as a harness: as long as tests went green, the team knew they were 'somewhat close.'; The plugin architecture was designed to avoid bloat: saying no to features by letting providers own/maintain their own code.
  • Caveats: Overfitted tests don't catch logic bugs—only structural correctness. If tests had been wrong or sparse, the refactor could have shipped broken code.; The sprint relied on 'blind faith in the harness'—Vincent contrasts this with his day-job evals (structured, telemetry-driven) but admits the walls are converging.; The refactor was a forcing function: one maintainer moving folders catalyzed the decision, not a planned roadmap.; Vincent and Peter had the risk appetite and compute budget to attempt this; most teams would never greenlight a 2 a.m. architecture overhaul.
  • Implications: Harness trust is the new infra primitive: without overfitted tests, the refactor fails. Invest in eval/test-gen companies (synthetic data, fuzzing, agent-authored suites).; The plugin model (provider-owned code) is a blueprint for open-source maintainability at agent scale: let AI companies own their integrations, avoid monolith bloat.; Ken: This story is a case study for agent-driven refactors. Package it as a talk, blog post, or playbook for eng leaders considering high-velocity agent adoption.; Blind faith in harness + people-management intuition = viable workflow. This isn't magic; it's process design + soft skills.

The Agent Development Environment: .skills, .files, and Bar Looping

  • Claims: Vincent treats .skills (prompt templates, workflows) like .dotfiles: versioned, improved via Codex log analysis, stored in GitHub (public + private).; He uses tools like .skills.sh (a loop mechanism) and GEPA (a skills gym he contributes to) to test and iterate skills.; His workflow: deploy skill → use in Codex → ask Codex to read last two weeks' logs → improve skill → redeploy.; 'Bar looping' (Vincent's term, coined with a maintainer) = opinionated, reward-driven agent loops vs. 'Ralph looping' (blind token burn).
  • Evidence: Vincent's .skills include co-created docs for technical writing (with dev-ex engineers); some are private, most are public on GitHub.; He uses Codex to analyze its own session logs and improve skills, treating agents as both workers and QA.; The bar-looping concept contrasts with giving an agent a task, burning tokens for 8–9 hours, and hoping—Vincent wants structured feedback (test pass = commit, PR graph = signal).
  • Caveats: Vincent doesn't explain how to measure skill quality, prevent drift, or resolve conflicts when multiple maintainers version skills differently.; Bar looping isn't formally defined—it's a vibe/direction, not a spec. Tools to operationalize it don't exist yet.; The .skills workflow is artisanal: works for Vincent but isn't plug-and-play for other teams.
  • Implications: Ken: .skills are the new dev supply chain. Expect marketplaces (agents store skills), version managers (skill lockfiles), and quality benchmarks (skills gym leaderboards).; Invest in companies building skills infra: portability across agents/IDEs, semantic search, testing harnesses.; Bar looping = next-gen agent orchestration. Whoever builds the 'reward loop platform' (structured feedback, test gates, PR de-dup) wins the agent-dev stack.; Content: Document Vincent's .skills workflow, publish templates, interview GEPA contributors, create a skills repo taxonomy.

PR/Issue Management at 6,000+ PR Scale: Semantic Graphs and De-duplication

  • Claims: OpenClaw has 6,000+ open PRs; every new maintainer tries to solve this with clustering, semantic graphs, or vector embeddings.; Vincent built a semantic graph with 706 edges for one PR, using vector embeddings across GitHub data.; The problem: everyone files duplicate issues/PRs, creating noise. The graph signals which issues have enough pressure to address.; There's a de-duplication process (not a roadmap) that helps maintainers decide what to work on.
  • Evidence: Vincent showed a slide of a semantic PR graph with 706 edges, illustrating the noise problem.; He mentions this is a 'running joke' among maintainers—everyone tries to solve the backlog, fails, then moves on.; The graph is used as a signal: if multiple 'clankers' (agents? contributors?) flag the same issue, it's worth addressing.
  • Caveats: Vincent admits the graph creates 'utter noise'—it doesn't solve the backlog, just visualizes the chaos.; No clear rubric for when to merge/close/ignore a PR; the process is still vibes-driven ('enough pressure = work on it').; Semantic graphs are compute-heavy and don't scale to real-time triage at 800 commits/day.
  • Implications: Ken: PR management at agent scale is unsolved infra. Invest in tools that auto-close dupes, cluster related issues, and surface high-signal PRs in real time.; This is a wedge for Linear, GitHub Copilot Workspace, or new YC startups: 'Semantic issue triage for agent-driven repos.'; The 'pressure signal' heuristic (multiple people file = important) could be formalized into a ranking algo—pitch this to GitHub PM team.; Content: Case study on OpenClaw's 6K PR backlog, extract patterns for other open-source projects hitting agent scale.

Evals and Harness Trust: Fake Slack, Synthetic Models, and Unit Test Overfitting

  • Claims: After the Great Refactor, OpenClaw built a 'fake Slack' with synthetic + real models to eval provider/channel integrations.; Vincent's day job is evals (structured, telemetry-driven); he contrasts this with his 'blind faith in the harness' on OpenClaw, but the two are converging.; Overfitted unit tests (which AI loves to generate) saved the refactor by acting as a green-light harness when 82% of code changed.; The bottleneck in 2026 will be token efficiency (not maxing), agent-in-the-loop, and process design.
  • Evidence: Vincent mentions the fake Slack runs evaluation loops to verify each provider (OpenAI, Anthropic, etc.) and channel (MS Teams, Discord) works.; At 1 a.m., tests were red; overfitted tests eventually passed, proving the refactor held together.; He explicitly says: '2025 was about token maxing. 2026 is about not wasting them. It's about token efficiency. It's about agent in the loop.'
  • Caveats: Overfitted tests don't catch logic bugs, edge cases, or production failures—only structural correctness.; The fake Slack eval is mentioned but not detailed—unclear if it runs on every PR, how long it takes, or what coverage it provides.; Vincent's 'blind faith' approach works because he has backup intuition (people-management skills, year of token maxing). Most teams won't.
  • Implications: Ken: Eval harnesses are the new CI/CD. Invest in companies building synthetic test envs (fake Slack/Discord/APIs), agent-authored suites, and coverage tools.; The shift from token maxing → token efficiency is the 2026 GTM narrative. Pitch this to LLM vendors, agent platforms, and enterprise AI buyers.; Overfitted tests as a harness = viable short-term strategy. Package this as a pattern: 'Generate overfitted tests during prototyping, refine in production.'; Content: Interview Vincent on harness trust, document the fake Slack architecture, create a taxonomy of agent eval strategies.

Notable Concepts & Terms

  • Dark Factory: A factory with no human workers—machines run autonomously. Vincent's metaphor for AI agents shipping code 24/7 without human intervention, analogous to the Industrial Revolution's shift from hand looms to mills.
  • Swim Lanes: Parallel Codex sessions (5–20) each assigned to a category (CI, features, bugs, P0s). The factory-manager mental model: agents are workers on different production lines, not a single supervised process.
  • Vibe Maintainer: An eng leader (e.g., Steve Yegge) who manages taste/direction/merges at high velocity using agent swarms, not by writing code. Vincent calls himself this; it's about soft skills, not technical chops.
  • Ralph Looping: Giving an agent a task, burning tokens for 8–9 hours, hoping it works. Named after autonomous agent loops; Vincent contrasts this with 'bar looping' (opinionated, reward-driven).
  • Bar Looping: Vincent's proposed alternative to Ralph looping: opinionated agent loops with structured feedback (e.g., test pass = commit), reward mechanisms, and de-duplication. Focus on token efficiency, not maxing.
  • .skills / .files: Versioned prompt templates, workflows, and agent instructions (like .dotfiles for devs). Vincent stores his publicly on GitHub, improves them via Codex log analysis, and treats them as infra.
  • Agent Development Environment (ADE): Vincent's framework for managing agent workflows: .skills, skills gym (GEPA), Codex log analysis, and deployment loops. The agent equivalent of IDEs/dev tooling.
  • The Great Refactor: OpenClaw's 24-hour sprint at NVIDIA: 2,700 commits, ~1M LOC changed, 82% codebase touched, plugin architecture launched. Saved by overfitted unit tests at 1 a.m.
  • Git Worktrees: Git feature that creates multiple working directories for one repo. Vincent used 70–80 at once (one per PR/agent) but regrets it—crashed his machine due to heavy test harness.
  • Overfitted Unit Tests: AI-generated tests that fit the code structure too closely (not robust). Vincent credits these with saving the Great Refactor—they went green when 82% of code changed, proving structural correctness.
  • Skills Gym (GEPA): A tool/platform Vincent contributes to for testing and iterating .skills (prompt templates). Analogous to a training ground for agent workflows.
  • Feeling Reasoning Tokens: Vincent's intuition for reading agent outputs (like Matrix code)—knowing when an agent is waffling, bullshitting, or lost based on how it explains itself, not just what it does.
  • Commit Maxing: Pushing the maximum number of commits possible (OpenClaw hit 800/day, Vincent hit 3,000/day). The 2025 strategy; Vincent says 2026 is about token efficiency instead.
  • Semantic PR Graph: Vector embeddings + graph of GitHub PRs/issues. Vincent built one with 706 edges for a single PR to de-duplicate and signal high-pressure issues. Still creates noise.

Operator Notes / Why Ken Should Care

  • Ken: This is the most detailed practitioner account of agent-driven dev at hyper-velocity. Vincent's swim-lane pattern, .skills workflow, and factory-manager mindset are the playbook for 2025–2026 eng orgs.
  • Invest in infra that makes Vincent's workflow less janky: Git worktree managers, swim-lane visualizers, agent drift detectors, .skills marketplaces, semantic PR triage tools.
  • The shift from 'review code' to 'manage agents like staff' is the new eng-manager skill gap. Expect training programs, certs, and consulting for 'vibe maintainers.'
  • Overfitted tests as a harness = short-term viable, long-term risky. Bet on companies building better eval frameworks (synthetic envs, agent-authored suites, coverage tools).
  • Vincent's 'feeling reasoning tokens' intuition took a year of token maxing to develop. This is a moat: early adopters who survive the jank will dominate. Late movers will struggle.
  • The Great Refactor (2,700 commits, 1M LOC, 82% codebase) is a case study for high-velocity agent adoption. Package this as content: blog, talk, playbook for eng leaders.
  • Bar looping (opinionated, reward-driven) vs. Ralph looping (blind burn) is the 2026 GTM narrative for agent platforms. Whoever formalizes bar looping wins the stack.
  • OpenClaw's PR backlog (6K+) is unsolved infra. Semantic triage, auto-close dupes, and pressure-signal ranking are wedges for Linear, GitHub, or YC startups.
  • Vincent's .skills are public on GitHub—download, study, and template for Ken's agent workflows. This is the supply chain for agent-driven dev.
  • The ChatGPT adoption curve (denial → ubiquity in 12 months) is repeating with agent-driven dev. Pitch: 'We're in the denial phase; infra built now wins the next stack.'
  • Vincent's commit graph maps his sleep schedule—this is the future of work. Expect burnout discourse, but also proof that humans + agents can 10–100x output if process is right.
  • Anthropic's C compiler, Spotify's no-code claim, Yegge's 50 PRs/day—these are public proof points. Interview them for Ken's content; document playbooks before they're common knowledge.
  • Vincent's day job (evals) + OpenClaw (blind faith) are converging. The future is structured evals + agent-in-the-loop, not pure autonomy or pure human control.
  • This talk is a forcing function for Ken's AI-ops thesis: agents need factory-manager workflows, not better models. Taste, process, and soft skills are the new bottleneck.

Watch Map

  • 00:00: Intro: Dark factories metaphor, shipping meme, velocity is not luck but engineering.
  • 01:30: Vincent's background: VR in 2013, East London roots, edge tech is janky.
  • 02:45: Industrial Revolution analogy: hand looms → mills → engineers as factory managers.
  • 03:30: Proof points: Anthropic C compiler, Spotify no-code, Yegge 50 PRs/day, OpenClaw 800 commits/day.
  • 04:45: Vincent's commit graph: 3,000 commits/day, stops when he sleeps, GitHub rate limits.
  • 05:00: Ralph looping vs. bar looping: opinionated, reward-driven agent loops.
  • 06:00: NVIDIA sprint story: building Nemo Claw, 15 Codex sessions each, 60–70 agents total.
  • 07:00: The Great Refactor: 2 a.m. decision, 2,700 commits, 1M LOC, 82% codebase, plugin architecture.
  • 07:45: 1 a.m. panic: tests red, 'did I vibe too hard?' Overfitted unit tests saved it.
  • 08:20: Swim lanes explained: 5–20 Codex sessions, CI/features/bugs/P0s, factory manager brain space.
  • 09:30: Git worktrees: 70–80 active, crashed machine, regrets adopting them.
  • 10:15: Feeling reasoning tokens: Matrix operator analogy, spotting agent bullshit.
  • 11:30: Agent Development Environment: .skills, .files, skills gym (GEPA), Codex log analysis.
  • 12:45: PR management: 6K+ PRs, semantic graphs (706 edges), de-duplication, pressure signals.
  • 13:30: Evals: fake Slack with synthetic + real models, overfitted tests as harness.
  • 14:00: Managing agents like staff: soft skills matter, how to spot bullshit, factory manager workflows.
  • 14:45: Closing: 2025 = token maxing, 2026 = token efficiency, agent-in-the-loop, process over model.

Source/Metadata

  • Title: Dark Factory: OpenClaw Ships Faster Than You Can Read the Diff — Vincent Koc, OpenClaw
  • Transcript words: 6362
  • Duration seconds: 1004
  • Timestamp note: No explicit timestamps in transcript; watch_map estimated based on 1004-second (16:44) duration and narrative flow.

Transcript

3061 words en Processed in 150.7s

[SPEAKER_00] They've got it. Cool. Amazing. So welcome everyone. I'm Vincent. What do I do? I'm one of the core maintainers at OpenClaw working with Peter. And as you've heard before, I have a day job as well. Same as Peter. He has a day job at OpenAI. But it's an open source project. Amazing things have been happening. I'm going to talk about what I call dark factories and how OpenClaw ships faster than you can read the diff. This meme is absolutely hilarious. So I think Peter posted this a week or two ago. I wake up. There's a new technological advancement. I wake up. It's this joke that we're shipping at insane speed and the velocity is just absolutely phenomenal. And some of you might think, oh, this is some luck or we're just looping to the max. I think there's actual engineering work here. And I'm going to talk about that. Now, as I mentioned, I'm Vincent. I'm your friend Niklanka. This is me using VR goggles back in 2013. So despite my accent that sounds somewhat Australian, I was born and raised in East London, not far from here. I actually went to college just down the road in Westminster. And, yeah, at some point I decided to live in Australia and my accent changed. But I used to love technology. I used to love being at the edge of technology. And this was one of the first few early VR goggles that came out. It came in this big box with a big warning sign on it saying, hey, use for five minutes at a time. Because it didn't have the anti-motion sickness built into it. And the funny thing with this one here was that I didn't use it for five minutes. I used it for three hours. And I played Team Fortress 2, had an absolute blast. And then I vomited for three hours after that. Because my vision turned into bivision. What I'm trying to say here is that anything on the edge is going to be janky. It's going to be horrific. It's going to be unchartered territory. And working on OpenClaw and being part of the team that ships probably an insane velocity of commits to a point where I get very limited by GitHub on an hourly basis is an interesting experience. And this experience, Britain's gone through before. We had the Industrial Revolution when mills and cotton were being produced at extreme amounts of volume. And there's a lot of history here around production and productionization at scale in the UK and in Europe. And I feel we're going through this moment again. We're going through this moment of how do we build at scale? And the ways we used to work before just don't work anymore. And it's strange because in my day job, I work in the space of evals, which everything is structured and there's telemetry and it has to be all perfect. And I work on a project where I have this blind faith in the harness. And it's this two walls, but they're starting to come together. We used to have hand looms in cottages, centralized mills everywhere. Craftsmen were the factory workers, but the bottleneck was the weaver's hands. We're now switching to a world where engineers writing code and editors, not so much. Swarms across repos, engineers are becoming factory managers, which I'm going to talk to, and the bottleneck becomes taste. That lovely word. Italian mother's hands, yes. So in context, what does this mean? Are you talking absolute nonsense? Are people building things at absolute scale? They are. What happened was very similar to the ChatGPT era, where everyone denied at scale that they were using ChatGPT. Everyone was in this absolute fear-mongering world. But what the reality was that everyone was using it. Everyone in secret was just like, oh, my God, what's going on? I need to talk to it. And the same thing is happening with autonomous agents at scale. Some organizations have openly come up with it. So, for example, Anthropic, with their recent work they did on building a new C compiler. We had Spotify saying they're no longer writing code by hand, supposedly. Steve Yeager, which I absolutely love, saying he pushes about 50 PRs a day total solo. He calls himself a vibe maintainer. I can relate to that. And OpenClaw, where we're pushing, at the peak we were doing 800 commits a day. And realistically, there's about 10 to 15 core maintainers all with day jobs. It's astronomical in terms of scale. And for me, this was March 15. What was that, two, three weeks ago, where I hit close to 3,000 commits per day. And if you actually look, my commits actually stop when I go to sleep. So if you want to see when I go to sleep and when I wake up and how many hours of sleep I have, you can just take a look at my commit history. Yes, it's astronomical. But the thing is, this is going to become the norm everywhere else. This is me telling you, you need to wake up. This scale of velocity is going to be normal. And trying to review PRs and go through all this nonsense may not work. But somewhere in the mix is engineering. There is a form of engineering that's going to happen. So we did commit maxing. Let's just go out there and smash as many commits as we can. And this reminds me of Ralph looping, right? This guy where you're like, hey, I'm just going to give you a task. I'm going to burn tokens for eight to nine hours. And you're waiting. You're hoping something happens. Maybe something happens. I don't know. But what if we had a bit more of an opinionated approach to this? What if we call it bar looping? I don't know. One of the other maintainers gave me this idea. Maybe we'll coin it. Do we need more than just tokens? What does that reward mechanism look like? How do we get a bit more opinionated? Yes, let's run loops, but let's be a bit more smart about how we do this. So right about the time you saw those 3,000 commits, this was the day before, I was at NVIDIA with Peter, and the gentleman you see on the left is one of the other NVIDIA gentlemen. And they were, hey, we're building Nemo Claw. I'm what? What's going on? And let's help you build it. And I was in the room. I was, I can't work on a laptop for hours on end. Can you bring me a screen? They bought me a screen. Peter didn't have a screen. So that's his laptop on the left. He asked for a screen. So they gave him an even bigger screen than mine, because, you know, why not? And we just got to work. So he's running about 15 codex sessions, and he's got his Mac Studio at home, he's VPNed into. I'm running another 10 or 15. And collectively, between Hey, we're building Nemo Claw. I'm like, what? What's going on? And let's help you build it. And I was in the room. I couldn't work on a laptop for hours on end. Can you bring me a screen? They bought me a screen. Peter didn't have a screen. So that's his laptop on the left. He asked for a screen. So they gave him an even bigger screen than mine. And we just got to work. So he's running about maybe 15 codex sessions, and he's got his Mac Studio at home, he's VPNed into. I'm running another 10 or 15. And collectively, between us, we're probably running with sub-agents included, maybe up to 60, 70 agents. But on the foreground, maybe 15 swim lanes. And we're just going for it. Funny thing is, we're working on Nemo Claw on one side, but one maintainer decided, I'm going to move some stuff around. I'm going to move a couple of folders around, and that was moving the entire channel. So all our conversations with MS Teams and Slack ended up moving to another location in the codebase. And we were like, oh my goodness, we're going to have to change stuff. And I found a really nice place to put my drink as well. The Nvidia people don't like this. So what ended up happening is what we call the great refactor. Essentially, we have lots of people raising PRs, and what they actually want is to build features. The thing is, we don't want to give everyone every single feature that they want, in which case, it becomes bloat. You heard Peter say earlier on, the challenge becomes, who do I say no to? It's not about saying yes. In a world where tokens are cheap, I can say yes to absolutely everyone and merge everything in. But that's going to turn this codebase into an absolute fire dump. So the vision was actually, we need to cut this codebase down. We need to rip it into pieces. And a plugin architecture somewhat made sense. Imagine if you're OpenAI or Mistral or Anthropic, what if you own that piece of the provider code and it was handed to you and it was separate from everything else. So this code change that occurred was a catalyst for us. It was 2 in the morning, we're tired, we thought, why not refactor the entire codebase? Sounds like a splendid idea. So 2,700 commits later, close to a million lines of code change, touching 82% of the core codebase, plugins were launched. The night before, I think it was 1 in the morning, I'm trying to go to sleep and the tests are not passing and I was like, was I Icarus and did I fly too close to the sun? As we call it, did I vibe too hard? I actually generally thought I vibed too hard. But as a team, we managed. We managed to bring this codebase back together again. But the saving grace was these awful unit tests that AI code loves to generate that actually ended up overfitting on our code. So when we completely ripped everything out, we still had these tests that were extremely overfitting, and as long as they would go green, we knew we were somewhat close. So how do we do this? In my case, I call it my factory. It's many codex sessions. Everyone asks me, what's this magic sauce? How do you do this? What's this crazy, insane thing? How are you guys building this? Very simple. I have swim lanes. It could be 5, it could be 10, it could be 20. But traditionally, they kind of cut themselves up into different pieces. So if I... Does this work, the laser? You can't really see it. But imagine you're a factory manager and you have a production line below. Essentially, you might have a case where you have, let's say, CI on one side, you might have features on one side, you might have bugs on another. So when I'm refactoring and doing stuff, right now the codebase is quite stable. I want to refactor some tests. So if I'm refactoring and doing stuff, that might be swim lanes one and two. I don't need to really babysit them too much. I just tell them, take your time, make sure the tests pass, just commit. Just push them through. Whereas with three and four, I might be looking at specific features and issues around, say, Docker or one of our messaging channels. In which case, I'm having a conversation with those agents. They're going off investigating, doing the work, coming back. And then maybe five is actually looking at new P0s and P1s. That might be using other data. They might be using GitHub. We have agents that run inside of a Discord channel. So when we do a release, we might be like, hey, what's happened in the last two hours that I need to be paying attention to? And this will scale up and down. But what ends up becoming quite interesting is tokens are no longer the problem. Depends who you ask. What really ends up becoming the problem is just raw compute and my brain space in order to keep an eye on all of these sessions. So in Harness, we trust. What ends up happening is I don't have this really insanely complicated process. The one thing I have complicated in my life is adopting Git work trees. And I kind of wish I hadn't. The only reason why I say this is when you're running an extremely heavy test harness, it ended up completely nuking my machine because I ended up running every PR I touch ends up becoming a new Git work tree. I end up with close to 70 or 80 active Git work trees in any given day on my machine. And that's kind of hell. So I had to actually build some magic sauce around my codex session. So my codex is aware of Git work trees. If I hit the escape key or it crashes, it will self heal, self recover, get sparse stuff. But realistically, I should have adopted what Peter and other people do and just clone the repo 10 times and point 10 different codex sessions to each one. But the trick here is that I haven't done any magical source. I don't use plan mode or spec mode. I have a conversation with the agent and we work through it and we find a way to make it work. So realistically, it looks a little bit like this from the matrix. And people go, oh, Vincent, how do you know it's kind of working? And this is going to sound somewhat a little bit lunatic. If anyone's watched the matrix and seen the scene where Neo goes over, it's like, how do you know? How do you read the text? And the guy's like, oh, you know, I've been doing this for a while. So I can see woman in red dress or guy walking dog. And you start to have this relationship where you can feel the reasoning tokens. I know it sounds the agent and we work through it and we find a way to make it work. So realistically, it looks a little bit like this from the matrix. And people go, oh, Vincent, how do you know it's working? And this is going to sound somewhat lunatic. If anyone's watched the matrix and seen the scene where Neo goes over, it's like, how do you know? How do you read the text? And the guy's like, oh, I've been doing this for a while. So I can see a woman in red dress or guy walking dog. And you start to have this relationship where you can feel the reasoning tokens. I know it sounds ludicrous. But there's times where I'm looking at the swim lane. I'm like, this sounds off. It doesn't sound off because of what it's doing. It sounds off because of how it's explaining itself to me. It's waffling. It's not making sense. It doesn't seem to know what it's doing. And this feels a lot like how I would manage people. If I had someone working for me and they started bullshitting, I'd be like, wait a minute, what's going on? So in these cases, I might just nuke the session and go, I'm not going to deal with this section of code. I'm going to leave that to another maintainer. Or I might come back to it four or five days later. But that experience feels very intuitive. And building that intuition, I've been able to get to because of the sheer volume of token maxing I've had to go through in the previous year. So there is engineering work. I call this the agent development environment. Essentially, the process goes, I have skills. I call it .skills, similar to .files. Both of my .skills and .files are available on GitHub. It's all open source. Go for it. Some of my skills are private. But there's skills in there for writing technical documentation, for example, that I've co-created with other developer experience and other engineers in the market. You can use a skills gym, something like a GEPA, which I'm also a contributor to. Or you could just say, codex, I've been using this skill in my last two weeks. Go through the codex sessions, read the logs, make improvements to the skill. I would then take that skill and deploy that into my open claw or take that into my personal environment. And I'll use something like .skills.sh as a mechanism to loop this. I've added some other testing and other elements on top of this. But there's a process to how I manage and maintain my skills as an engineer. The way we manage PRs has some level of engineering work to it. There's this running joke that every maintainer that joins the project decides to try and tackle, oh my God, we have 6,000 PRs. How are we going to solve it? I'm going to cluster everything and figure this out. How many? There you go. Oh no, thank you very much. So this was my flavor of trying to solve this. It's a semantic graphing, vector embedding on the entire GitHub stuff. This is one PR, has 706 edges. What ends up happening is that everyone else has the same problem, so they decide to send their flavor of the PR issue. It becomes utter noise. So there is process around how we consume what we're going to work on. We might not call it a roadmap, but we have a way of de-duplicating and seeing what's out there. This might be a signal for me to say, okay, if there's enough pressure coming on one issue, it must be big enough that all these other clankers decided it's a big problem. Maybe I should go and address it. There is evals, surprisingly. After all this refactoring work, we decided to make a fake slack of sorts with both synthetic models and real models so we can run evaluation loops to check that each of the providers and the channels work. And this question was asked to me recently. How do you manage 10 plus agents? And this is something that you're thinking. I asked them back, how do you manage 10 plus staff? And they had no answer for me. I'd worked in large organizations like airlines and other places managing large AI teams. I had experience managing up to 30, 40 people plus. So for me, it was not a new paradigm. But I think for engineers and people working with these coding agents at scale, it's the soft skills that matter. It's how do you ask your agent what's going on? How do you know when they're not bullshitting you? And how do you run that factory? So it's no longer about the model or the agent. It's about the process. 2025 was about token maxing. 2026 is about not wasting them. It's about token efficiency. It's about agent in the loop. Thank you. things have been happening. I'm going to talk about what I call dark factories and how OpenClaw ships faster than you can read the diff. This meme is absolutely hilarious. So I think Peter posted this a week or two ago. I wake up. There's a new technological advancement. I wake up. It's this joke that we're shipping at insane speed and the velocity is just absolutely phenomenal. And some of you might think, oh, this is some luck or we're just like Ralph looping to the max. I think there's actual engineering work here. And I'm going to talk about that. Now, as I mentioned, I'm Vincent. I'm your friend Niklanka. This is me using VR goggles back in 2013. So despite my accent that sounds somewhat Australian, I was born and raised in East London, not far from here. I actually went to college just down the road in Westminster. And, yeah, at some point I decided to live in Australia and my accent changed. But I used to love technology. I used to love being at the edge of technology. And this was like one of the first few sort of early VR goggles that came out. It came in this big box with a big warning sign on it saying, hey, use for five minutes at a time. Because it didn't have like the anti-motion sickness built into it. And the funny thing with this one here was that I didn't use it for five minutes. I used it for three hours. And I played Team Fortress 2, had an absolute blast. And then I vomited for three hours after that. Because my vision turned into BVision. What I'm trying to say here is that like anything on the edge is going to be janky. It's going to be horrific. It's going to be unchartered territory. And working on OpenClaw and being part of the team that ships probably, you know, an insane velocity of commits to a point where I get very limited by GitHub on an hourly basis is an interesting experience. And this experience, Britain's gone through before. We had the Industrial Revolution when mills and cotton were being produced at extreme amounts of volume. And there's a lot of history here around production and productionization at scale in the UK and in Europe. And I feel like we're going through this moment again. We're going through this moment of how do we build at scale? And the ways we used to work before just don't work anymore. And it's kind of strange because in my day job, I kind of work in the space of evals, which everything is sort of structured and there's telemetry and it has to be all perfect. And I work on a project where I'm, I have this blind faith in the harness. And it's this kind of two walls, but they're starting to come together. We used to have hand looms in cottages, centralized mills everywhere. Craftsmen were the factory workers, but the bottleneck was the weaver's hands. We're now switching to a world where engineers writing code and editors, not so much. Swarms across repos, engineers are becoming factory managers, which I'm going to talk to, and the bottleneck becomes taste. You know, that lovely word. Italian mother's hands, yes. So in context, like what does this mean? Like, are you talking absolute nonsense? Are people building things at absolute scale? They are. What happened was very similar to the ChatGPT era, where everyone denied it at scale that they were using ChatGPT. Everyone was in this absolute fear-mongering sort of world. But what the reality was that everyone was using it. Everyone in secret was just like, oh, my God, what's going on? I need to talk to it. And the same thing is happening with this autonomous agents at scale. Some organizations have openly come up with it. So, for example, Anthropic, with their recent work they did on building a new C compiler. We had Spotify saying they're no longer writing code by hand, supposedly. Steve Yeager, which I absolutely love, saying he pushes about 50 PRs a day total solo. He calls himself a vibe maintainer. I can kind of relate to that. And OpenClaw, where we're pushing, at the peak we were doing 800 commits a day. And realistically, like there's about 10 to 15 core maintainers all with day jobs. It's kind of astronomical in terms of scale. And for me, this was March 15. What was that, like two, three weeks ago, where I hit close to 3,000 commits per day. And if you actually look, my commits actually stop when I go to sleep. So if you want to see when I go to sleep and when I wake up and how many hours of sleep I have, you can just take a look at my commit history. Yeah, it's astronomical. But the thing is, this is going to become the norm everywhere else. Like this is like me telling you, you need to wake up that, you know, this scale of velocity is going to be normal. And trying to review PRs and go through all this nonsense may not work. But somewhere in the mix is engineering. There is a form of engineering that's going to happen. So we did commit maxing. You know, let's just go out there and smash as many commits as we can. And this reminds me of Ralph looping, right? This guy where you're like, hey, I'm just going to like give you a task. I'm going to burn tokens for like eight to nine hours. And you're waiting, you know. You're waiting. You're hoping something happens. Maybe something happens. I don't know. But what if we had a bit more of an opinionated approach to this? What if we call it bar looping? I don't know. One of the other maintainers gave me this idea. Maybe we'll coin it. Do we need more than just tokens? What does that reward mechanism look like? How do we get a bit more opinionated? Yes, let's run loops, but let's be a bit more smart about how we do this. So right about the time you saw those 3,000 commits, this was the day before, I was at NVIDIA with Peter, and the gentleman you see on the left is one of the other NVIDIA gentlemen. And they were like, hey, we're building Nemo Claw. I'm like, what? What's going on? And let's help you build it. And I was in the room. I was like, I can't work on a laptop for like hours on end. Can you bring me a screen? They bought me a screen. Peter didn't have a screen. So that's his laptop on the left. He asked for a screen. So they gave him an even bigger screen than mine, because, you know, why not? And we just got to work. So he's running about maybe 15 codex sessions, and he's got his Mac Studio at home, he's VPNed into. I'm running another like 10 or 15. And collectively, between us, we're probably running with sub-agents included, maybe up to 60, 70 agents. But on the foreground, maybe 15 swim lanes, if you want to call it that. And we're just going for it. Funny thing is, we're working on Nemo Claw on one side, but one maintainer decided, I'm going to move some stuff around. I'm going to move a couple of folders around, and that was moving the entire channel. So like, all our conversations with like MS Teams and Slack, ended up moving to another location in the codebase. And we were like, oh my goodness, we're going to have to change stuff. And I found a really nice place to put my drink as well. The Nvidia people don't like this. So what ended up happening is what we call the great refactor. Essentially, where we were like, hey, we have lots of people raising PRs, and what they actually want is to build features. The thing is, we don't want to give everyone every single feature that they want, in which case, it becomes bloat. You heard Peter say earlier on, the challenge becomes, who do I say no to? It's not about saying yes. In a world where tokens are cheap, I can just say yes to absolutely everyone and merge everything in. But that's going to turn this codebase into an absolute fire dump. So the vision was actually, we need to cut this codebase down. We need to rip it into pieces. And a plugin architecture somewhat made sense. Imagine if you're open AI or Mistral or Anthropic, what if you own that piece of the provider code and it was handed to you and it was separate from everything else. So this code change that occurred was like a catalyst for us. It was 2 in the morning, we're tired, we thought, why not refactor the entire codebase? Sounds like a splendid idea. So 2,700 commits later, close to a million lines of code change, touching 82% of the core codebase, plugins were launched. The night before, I think it was like 1 in the morning, I'm trying to go to sleep and the tests are not passing and I was like, was I Icarus and did I fly too close to the sun? As we like to call it, did I vibe too hard? I actually generally thought I vibed too hard. But as a team, we managed. We managed to bring this codebase back together again. But the saving grace was these awful sort of unit tests that AI code loves to generate that actually ended up over fitting on our code. So when we completely ripped everything out, we still had these tests that were like extremely over fitting, and as long as they would go green, we knew we were kind of somewhat close. So how do we do this? You know? In my case, I call it my factory. It's many codex sessions. Everyone asks me, like, what's this magic sauce? Like, how do you do this? What's this crazy, insane thing? Like, how are you guys building this? Very simple. I have swim lanes. It could be 5, it could be 10, it could be 20. But traditionally, they kind of cut themselves up into different pieces. So like, if I... Does this work, the laser? You can't really see it. But imagine you're a factory manager, and you have a production line below. Essentially, you might have a case where you have, let's just say, CI to one side, you might have features on one side, you might have bugs on another. So when I'm refactoring and doing stuff, right now the codebase is quite stable. I want to refactor some tests. So if I'm refactoring and doing stuff, well, that might be swim lanes one and two. I don't need to really babysit them too much. I just tell them, take your time, make sure the tests pass, just commit. Just push them through. Whereas with three and four, I might be looking at specific features and issues around, say, Docker or one of our messaging channels. In which case, I'm having a conversation with those agents. They're going off investigating, doing the work, coming back. And then maybe five is actually looking at new P0s and P1s. That might be using other data. They might be using GitHub. We have agents that run inside of a Discord channel. So when we do a release, we might be like, hey, what's happened in the last two hours that I need to be paying attention to? And this will scale up and down. But what ends up becoming quite interesting is tokens are no longer the problem. Depends who you ask. What really ends up becoming the problem is just raw compute and my brain space in order to sort of keep an eye on all of these sessions. So in Harness, we trust. What ends up happening is I don't have this really insanely complicated process. The one thing I have complicated in my life is adopting Git work trees. And I kind of wish I hadn't. The only reason why I say this is when you're running an extremely heavy test harness, it ended up completely nuking my machine because I ended up running like every PR I touch ends up becoming a new Git work tree. I end up with like someone close to like 70 or 80 active Git work trees in any given day on my machine. And that's kind of hell. So I had to actually build some like magic sauce around my codex session. So my codex is aware of Git work trees. If I hit the escape key or crashes, it will self heal, self recover, get, you know, sparse stuff. But realistically, I should have adopted what Peter and other people do and just like clone the repo 10 times and point 10 different, you know, codex sessions to each one. But the trick here is that, like I haven't done any magical source. I don't use plan mode or spec mode. I have a conversation with the agent and we work through it and we find a way to make it work. So realistically, it looks a little bit like this from the matrix. And people go, oh, Vincent, like, how do you know it's kind of working? And this is going to sound somewhat a little bit lunatic. If anyone's watched the matrix and seen the scene where Neo goes over, it's like, how do you know? How do you read the text? And the guy's like, oh, you know, I've been doing this for a while. So I can see like, woman in red dress or guy walking dog. And you start to have this like relationship where you can feel the reasoning tokens. I know it sounds somewhat ludicrous. But there's times where I'm looking at the swim lane. I'm like, this sounds off. It doesn't sound off because of what it's doing. It sounds off because of how it's explaining itself to me. It's waffling. It's not making sense. It doesn't seem to know what it's doing. And this feels a lot like how I would manage people. If I had someone working for me and they started downright bullshitting, I'd be like, wait a minute, what's going on? So in these cases, I might just nuke the session and go, you know, I'm not going to deal with this section of code. I'm going to leave that to another maintainer. Or I might come back to it four or five days later. But that experience feels very much like intuitive. And building that intuition, I've been able to get to because of the sheer volume of token maxing I've had to go through in the previous year. So there is engineering work. I call this the agent development environment. Essentially, the process goes, I have skills. I call it .skills, similar to .files. Both of my .skills and .files is available on GitHub. It's all open source. Go for it. Some of my skills are private. But there's skills in there for writing technical documentation, for example, that I've co-created with other developer experience and other engineers in the market. You can use a skills gym, something like a GEPA, which I'm also a contributor to. Or you could just say, codex, I've been using this skill in my last two weeks. Go through the codex sessions, read the logs, make improvements to the skill. I would then take that skill and deploy that into my open claw or take that into my personal environment. And I'll use something like .skills.sh as a mechanism to loop this. I've added some other testing and other elements on top of this. But there's a process to how I manage and maintain my skills as an engineer. The way we manage PRs has some level of engineering work to it. There's this kind of running joke that every maintainer that joins the project decides to try and tackle, like, oh, my God, we have 6,000 PRs. How are we going to solve it? I'm going to cluster everything and, like, figure this out. How many? There you go. There you go. Oh, no, thank you very much. So this was my flavor of, like, trying to solve this. It's like a semantic graphing, vector embedding on the entire GitHub stuff. This is one PR, has 706 edges. What ends up happening is that everyone else has the same problem, so they decide to send their flavor of the PR issue. It becomes utter noise. So there is even process around, like, how we even consume what we're going to work on. We might not call it a roadmap, but we have a way of kind of de-duplicating and seeing what's out there. This might be a signal for me to say, okay, if there's enough pressure coming on one issue, it must be big enough that all these other clankers decided it's a big problem. Maybe I should go and address it. There is evals, surprisingly. After all this refactoring work, we decided to make a fake slack of sorts with both synthetic models and real models so we can run evaluation loops to check that each of the providers and the channels work. And this question was asked to me recently. How do you manage 10 plus agents? And this is something that you're thinking. I asked them back, how do you manage 10 plus staff? And they had no answer for me. I'd worked in large organizations like airlines and other places like that managing large AI teams. I had experience managing up to 30, 40 people plus. So for me, it was not like a new paradigm. But I think for engineers and people working with these coding agents at scale, it's the soft skills that matter. It's how do you ask your agent what's going on? How do you know when they're not bullshitting you? And how do you run that factory? So it's no longer about the model or the agent. It's about the process. 2025 was about token maxing. 2026 is about not wasting them. It's about token efficiency. It's about agent in the loop. Thank you. So So ! So ! So