Open Reader

The Prime Intellect Stack — Will Brown, Prime Intellect

completed 46:52 Jul 13, 2026 Watch on YouTube

Current Status

completed

Video ID

V-EDrhIhHzQ

RAG / Chat

Enabled
The Prime Intellect Stack — Will Brown, Prime Intellect
Description

Deep dive into Prime Intellect's open-source ecosystem of post-training tools, including the verifiers and prime-rl libraries, as well as the Lab platform for self-serve training and inference. Speaker: Will Brown — Research Lead, Prime Intellect Will Brown leads Applied Research at Prime Intellect and builds open research infrastructure to enable every company to train, deploy, and self-improve their own frontier agentic models. He holds a PhD in Computer Science from Columbia University. X: https://x.com/willccbb LinkedIn: https://www.linkedin.com/in/willcb/ GitHub: https://github.com/willccbb Website: https://willcb.com TImestamps 0:00 Introduction and Overview of Prime Intellect 4:20 Defining the Environment in Post-Training 9:33 Decomposing Environments: Tasks, Harnesses, and Runtimes 12:46 Verifiers V1: The New Modular Pattern 17:46 Rewards, Metrics, and Group-Level Rewards 20:25 Tooling, User Simulators, and MCP Integration 22:00 The Interception Server Pattern 24:13 Trace Graphs and Handling Tokenization 25:35 The Renderers Library for Chat Templates 29:20 Primaril: Asynchronous Reinforcement Learning 38:02 Customizing Training Algorithms and Losses 42:35 The Lab Platform and Hosted Training

Summary

Generated by claude-sonnet-4-5

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: Prime Intellect provides an open-source, full-stack post-training infrastructure (Verifiers V1 + Prime RL) that makes frontier model RL training accessible, fast, and cost-effective by decoupling environments, harnesses, and runtimes while supporting modern algorithms like OPD, self-distillation, and GRPO at scale.
  • Why it matters: Makes $50k, 3-day GLM-5 RL runs (28 nodes, 131k context) feasible for enterprises; demonstrates how to architect production-grade agent post-training systems that don't exist elsewhere in open source.
  • Best use: Watch if building agent post-training systems, evaluating RL infrastructure vendors, or need to understand modern post-training flywheel architecture; skip detail on trace graphs/renderers unless implementing.

Executive Summary

Prime Intellect has rebuilt its open-source post-training stack (Verifiers V1 + Prime RL) around a fundamental insight: environments are not just for RL—they're the unifying abstraction for evals, SFT data generation, on-policy distillation, and reinforcement learning. The new architecture decouples task sets (agent-agnostic data/rules), harnesses (execution patterns: tool loops, CLI agents, RLMs, MCP servers), and runtimes (local/Docker/sandboxes), letting users mix-and-match without rebuilding core logic. This matters because it solves the 'can't iterate' problem: you can eval locally on a MacBook, then scale the same environment to cloud sandboxes for RL without rewriting code.

The core technical achievement is making large-scale async RL tractable and cost-effective for frontier models. Prime RL now does GLM-5 (131k context) RL steps on 28 nodes in under 5 minutes, enabling 1000-step runs in 3 days for ~$50k—cheaper than OpenAI API spend for many enterprises. This is possible because async RL (average 16 steps off-policy) overlaps long-tail agent rollouts (30 seconds to 3 hours) without GPU blocking, while the interception server pattern lets any harness (Codex, Claude Code, Langchain, DSPy) plug into RL training without knowing it's doing RL. The system handles FP8, wide expert parallelism, desegregated prefill, and router replay metadata at scale on Torch Titan (chosen for hackability over Megatron monolith).

Prime Intellect has also solved subtle but critical engineering problems that break other RL systems: retokenization drift (where messages→tokens→messages changes due to many-to-one tokenization, causing trainer/inference mismatch), group rewards (e.g., 'bonus for shortest correct answer' requires comparing rollouts, impossible in decoupled systems), and algorithm composability (SFT/OPD/GRPO/Echo now swap via config, not buried if-statements). The 'renderers' library (standalone, works with any inference engine) treats chat templates as programmable artifacts to manage token/message duality, inspired by OpenAI Harmony and used in Thinking Machines. This explains why OpenAI Responses became stateful—it's unavoidable at scale with agents.

The post-training flywheel architecture is: evals unlock RL → train expert models per task → distill experts into one model → deploy with LoRA hot-swapping → capture real-world feedback → refine environments → repeat. Multi-tenant LoRA training (live now) lets multiple users train on shared base weights with token-based pricing; full fine-tuning (coming in 2 weeks) supports custom algorithms/losses. The team operates 10k+ GPUs across global data centers, runs their own Intellect model series, and trains customer models. Team <40 people, research <10, hiring aggressively in SF. Open research + real business model via hosted training platform (Lab) and compute marketplace.

Key Takeaways

  • Claim: Environments are the unified abstraction for evals, SFT, OPD, and RL—not just RL—because they encapsulate task logic, execution harness, and verification/scoring in one reusable unit. | Evidence: Verifiers V1 decouples task sets (Hugging Face datasets, Harbor, NemoGym, OpenEnv) from harnesses (default tool loop, RLM, CLI agents like Codex/Cloud Code, MiniSuite, Langchain/DSPy) from runtimes (local/Docker/Prime Sandboxes). Same environment runs evals locally on MacBook, then scales to cloud RL without code changes. Supports MCP servers as tools and user simulators. | Caveat: Requires discipline to keep task sets harness-agnostic; some task/harness compatibility checks needed. Integration with Harbor/NemoGym means accepting their design constraints. | Implication: Ken can build one environment for a business workflow, use it for prompt optimization evals (API models), SFT data collection, and RL fine-tuning without rewriting. Lowers activation energy for post-training experimentation. | Timestamp: 00:00–15:00
  • Claim: Prime RL achieves 1000-step GLM-5 RL run (131k context, coding tasks) in 3 days on 28 nodes for ~$50k by doing async RL with <5 min step time, making frontier post-training affordable vs. API spend. | Evidence: GLM-5 (pre-5.2) step time <5 min on 28 nodes with 131k context on long-horizon coding tasks = 1000 steps in ~3 days at rental GPU pricing ~$50k. Comparable to monthly API token spend for many enterprises. Supports FP8, wide expert parallelism (multi-node MoE), desegregated prefill, router replay metadata offloading. | Caveat: $50k is not cheap for all use cases; assumes rental GPU pricing, not reserved/owned. Only makes sense if you have reproducible environments and can amortize training cost over inference budget. Async RL means 16-step average off-policy—not suitable if you need strict on-policy guarantees (but Will argues you don't for agents). | Implication: If Ken's inference spend is $50k+/month or he needs model behavior not available via API (reasoning trace, custom reward shaping, conciseness penalties), a 3-day RL run becomes justifiable. Reframes post-training as operational expense, not research moonshot. | Timestamp: 35:00–38:00
  • Claim: Retokenization drift (messages→tokens→messages changes output due to many-to-one tokenization) causes silent off-policy trainer/inference mismatches; solving it requires dual message/token streams and stateful APIs, which is why OpenAI Responses is stateful. | Evidence: Chat templates strip extra newlines or whitespace; re-tokenizing model output can produce different token IDs, causing logical branches in token space that don't exist in message space. Prime's trace graph stores both message-level (text) and token-level (IDs) representations. 'Renderers' library (standalone, works with any engine, inspired by OpenAI Harmony/Thinking Machines Tinker) treats chat templates as programmable artifacts with prefix hit detection to avoid retokenization. | Caveat: Only matters for RL training with log probs and multi-turn agents; not an issue for stateless single-turn inference. Adds complexity to system design. Some models (OpenAI reasoning models) don't return log probs anyway, so eval-only. | Implication: If Ken is building multi-turn RL or trying to understand why his trainer/inference results diverge, this is the culprit. Renderers library is standalone—can use it outside Prime stack to debug tokenization issues. Explains why stateful LLM APIs aren't just 'bad design.' | Timestamp: 23:00–27:00
  • Claim: Group rewards (e.g., 'bonus for shortest correct answer,' pairwise judging, ranking) are critical for agent RL but ignored by most frameworks because decoupled rollout architectures can't compare rollouts; Prime supports them first-class. | Evidence: Example: conciseness bonus—you don't know optimal token length per problem upfront, and it changes as model improves. Solution: take all correct rollouts in a group, give bonus to shortest. Requires variance across multiple samples, which RL provides but SFT doesn't. Most RL frameworks assume rollouts are independent; Prime's orchestrator/batch system allows group-level logic. | Caveat: Requires careful reward design to balance multiple objectives (correctness + efficiency). Group size affects variance/stability. Not all tasks benefit from group rewards—mainly useful when 'optimal' is relative (length, cost, speed) not absolute (correctness). | Implication: Ken can penalize models for overuse of reasoning tokens or reward efficiency without hardcoding token limits. Essential for preventing 'thinking forever' in open models during post-training. Differentiates Prime from simpler RL stacks. | Timestamp: 17:00–19:00
  • Claim: The interception server pattern (fake base URL → intercept requests → add log probs/temperature → forward to RL trainer) lets any harness (Codex, Claude Code, Langchain, DSPy) do RL without knowing it, enabling seamless eval→train→deploy workflows. | Evidence: Harness thinks it's calling OpenAI/Anthropic API (OpenAI chat completions, responses, or Anthropic dialect). Prime gives it fake base URL, intercepts each request, injects RL metadata (log probs, temperature), sends to inference server via renderers, returns response. Harness is just code—no RL-specific dependencies. Same harness runs eval (local) and RL (cloud sandboxes). | Caveat: Requires harnesses to be API-compatible (not all are). Some models don't return log probs (OpenAI reasoning models), so eval-only. Adds latency vs. direct inference, though async RL hides this. Fake base URL pattern can be confusing to debug. | Implication: Ken can take any existing agent codebase (Langchain, DSPy, Codex harness) and plug it into RL training without refactoring. Lowers switching cost between eval and RL. Also means you can test harness changes (e.g., system prompt, tools) in eval mode cheaply before committing to RL run. | Timestamp: 20:00–23:00
  • Claim: Modern post-training algorithms (SFT, OPD, self-distillation, GRPO, Echo) now swap via config in Prime RL by decomposing into 'algorithm' (data prep/scoring) and 'loss' (gradient target), not buried if-statements. | Evidence: Factorization: rollouts come from policy (on-policy RL/async RL) or teacher (SFT/context distillation/OPD). Advantage is reward-baseline (RL) or log prob ratio (OPD) or advantage=1 (cross-entropy/NLL for SFT). Echo: two algorithm components, cross-entropy on environment tokens + RL on action tokens. All configurable per-environment in TOML; add custom algorithms by writing a class that assigns advantages to rollouts. | Caveat: Assumes you understand RL/distillation algorithms well enough to configure them correctly. Factorization is cleaner but still requires domain knowledge (e.g., when to use OPD vs. GRPO). Max RL paper patterns not yet fully integrated but 'coming soon.' | Implication: Ken can experiment with new algorithms (e.g., policy distillation from o1 for reasoning tasks, self-distillation with hints) without waiting for framework updates. Research velocity increases if you can test ideas in days, not months. Also: enterprises can hire RL-literate engineers without needing framework engineers. | Timestamp: 40:00–44:00
  • Claim: Async RL (average 16 steps off-policy) is mandatory for agent RL because coding agent rollouts have long tails (30 sec to 3 hours), and waiting on slowest rollout wastes GPU; overlapping rollouts prevents blocking. | Evidence: Coding tasks (Codex, Cloud Code, OpenCode) can take 30 seconds to 3 hours per rollout. Synchronous RL forces all GPUs to wait for slowest rollout → GPU idle time. Async RL allows rollouts to finish long after they start, enter first available batch, use latest model weights in inference server. Prime typically runs 16 steps off-policy on average; DPPO paper shows this is stable up to thousands of steps (pushing toward 10k in experiments). | Caveat: Off-policy RL introduces staleness—rollouts use older model version than current trainer. Requires off-policy correction (DPPO, etc.) to stay stable. Not suitable if you need strict on-policy guarantees, though Will argues this is unnecessary for agents. Average 16 steps off-policy may not generalize to all tasks. | Implication: Ken should not build synchronous RL for agents—it's a dead end. If evaluating RL vendors, ask whether they support async and how far off-policy they've validated. Also: slow components (grading with judge, sandbox boot time) are fine in async RL, so Ken doesn't need to over-optimize them early. | Timestamp: 36:00–40:00
  • Claim: Training expert models per task, then distilling teachers into one model works better than training one model on all tasks simultaneously, especially for multi-objective agent workflows. | Evidence: Pattern: train individual RL experts on same base model for different environments/tasks, then do distillation from multiple teachers into one checkpoint. 'People have found this is more reliable' than single multi-task RL. Mentioned as common pattern in frontier model training. | Caveat: No quantitative comparison provided (e.g., expert-then-distill vs. multi-task RL win rate). Requires compute for N expert runs + distillation run. Distillation can lose capabilities if done poorly. Not clear when this pattern is overkill vs. necessary. | Implication: If Ken has multiple distinct agent tasks (e.g., code generation, debugging, planning), training separate experts then distilling may be more reliable than joint RL. Worth testing if multi-task RL plateau early. Also: suggests model merging/distillation tooling is critical for post-training stack. | Timestamp: 08:00–10:00

Detailed Brief

Verifiers V1 Architecture: Task Sets, Harnesses, Runtimes

  • Claims: Verifiers V1 overhauls old multi-turn environment pattern by decoupling task sets (data/rules, agent-agnostic), harnesses (execution pattern: tool loop, CLI, RLM), and runtimes (local/Docker/sandboxes).; Task sets integrate natively with Hugging Face datasets, Harbor, NemoGym, OpenEnv—any RL environment framework becomes a Prime task set.; Harnesses support default (system prompt + tools in loop), RLM (recursive language models), CLI agents (Codex, Cloud Code, OpenCode), MiniSuite Agent, and custom (Langchain, DSPy).; Runtimes use async IO + subprocesses; support local, local Docker, Prime Sandboxes, or custom sandbox layers.; UV script pattern (UV tooling) manages dependencies per harness/task without conflicts; grading can also be UV scripts.; Environments = task set + harness + runtime; plugging these together yields a 'trace' (rollout record).
  • Evidence: Dev release on main branch (prime-intellect-ai/verifiers GitHub), stable PyPI release 'any minute now.'; Decorator pattern (Pydantic-heavy, TOML configs, CLI overrides, type-check at validation time).; Examples: SWEGrep (agentic code search), Wordle, document search with judges, Harbor benchmarks with terminal agents.; Old 'rubric' pattern killed—didn't make sense for new architecture.; MCP (Model Context Protocol) used as backend for tools and user simulators; model sees user simulator as user, not tool.
  • Caveats: Task set/harness compatibility requires sanity checks (properties each supports/requires).; Local-to-cloud swap is convenient but adds abstraction—debugging fake base URL can be tricky.; UV script dependency isolation is powerful but assumes familiarity with UV tooling.; User simulators (LLM-as-user in multi-turn) are useful for realism but add cost/latency; not always needed.
  • Implications: Ken can prototype environments locally (fast, free) then scale to cloud (sandboxes, multi-node) without rewrites—critical for iteration speed.; Harbor/NemoGym integration means Ken can use existing benchmarks (e.g., SWE-Bench derivatives) as Prime environments immediately.; MCP support means Ken's MCP servers (tools, user simulators) plug directly into RL training—no custom integration needed.; Eval CLI hot-swap (local vs. sandboxes, different harnesses) lets Ken A/B test harnesses (e.g., does RLM beat tool loop for this task?) cheaply before RL.

Rewards, Metrics, Group Rewards, and Conciseness Patterns

  • Claims: Rewards are functions that take rollout records → return numbers (drive RL progress); metrics are logging (tool use count, errors).; Group rewards enable pairwise judging, ranking, shortest-correct-answer bonuses—ignored by most RL frameworks but critical for agents.; Conciseness bonus pattern: give reward to shortest correct answer in group, since optimal length is unknown/task-dependent/changes over time.; Models will 'think forever' without length penalties; open models often have very long CoT because it's useful but grows out of control.; Juggling multiple objectives (correctness + efficiency) is hard in reward design; group-level comparisons help balance them.
  • Evidence: Example: math problem—can't hardcode 'solve in <n tokens' because n is unknown, varies per problem, changes as model improves.; RL provides multiple samples (variance), allowing group rewards; SFT doesn't (single teacher trajectory).; Prime's orchestrator/batch system allows group-level logic; most frameworks assume independent rollouts.
  • Caveats: Group rewards require careful design to avoid degenerate solutions (e.g., all rollouts become maximally short but incorrect).; Group size affects variance/stability—too small = noisy, too large = slow.; Not all tasks benefit from group rewards; mainly useful when 'optimal' is relative (length, cost, speed) vs. absolute (correctness).; Length penalties can hurt tasks where longer reasoning is genuinely needed; need task-specific tuning.
  • Implications: Ken should include conciseness bonus in coding/reasoning RL to prevent token bloat, especially for open models (Llama, Qwen, DeepSeek).; If Ken's models are over-explaining or using too many tool calls, group rewards can shape this without hardcoded limits.; Metrics (tool use count, errors) should be first-class in dashboards—Ken needs visibility into behavior changes, not just loss curves.; Multi-objective RL (correctness + speed + cost) is tractable with group rewards; worth experimenting vs. single-reward RL.

Trace Graphs, Renderers, and Tokenization Management

  • Claims: Trace graph refactored to support sub-agents, parallel branching trees, and linear token dependencies for RL.; Branches are at message level (logical text); tokens are separate stream; dual representation prevents retokenization drift.; Retokenization drift: message→tokens→message can change due to many-to-one tokenization (extra newlines, whitespace stripped by chat templates).; Chat templates (Jinja) are painful; renderers library (standalone, Python, no Prime dependencies) treats them as programmable artifacts.; Renderers inspired by OpenAI Harmony (GPT OSS) and Thinking Machines cookbooks (Tinker); works with any inference engine.; Stateful LLM APIs (OpenAI Responses) exist because tokenizer subtleties create unavoidable issues at scale with agents—not just 'bad design.'
  • Evidence: Renderers library: prefix hit detection in message space even if not in token space after retokenization.; Trace stores both message-level (harness sees text) and token-level (trainer sees token IDs) representations.; Example: model says something, you convert to text, re-tokenize → token IDs change → trainer/inference mismatch → either off-policy or logical branch in token space.; OpenAI Responses is stateful to avoid retokenization drift in multi-turn with reasoning traces.
  • Caveats: Renderers add complexity—users must manage dual streams even if they don't care about tokens.; Only critical for RL with log probs and multi-turn agents; single-turn or eval-only workflows can ignore this.; Some models (OpenAI reasoning) don't return log probs, so renderer doesn't help—eval-only.; Stateful APIs (prefix caching, KV cache reuse) have benefits beyond tokenization, but this is one unavoidable reason.
  • Implications: Ken should use renderers library if building multi-turn RL to avoid silent trainer/inference divergence—saves debugging time.; If Ken's RL runs are unstable or losing reward unexpectedly, retokenization drift may be culprit; check message vs. token consistency.; Standalone renderers library means Ken can adopt it outside Prime stack (e.g., with VLLM, TGI, custom inference) for same benefits.; Understanding tokenizer issues explains why OpenAI/Anthropic made certain API design choices—not arbitrary.

Prime RL: Orchestrator, Async RL, Inference/Trainer Separation

  • Claims: Prime RL is async-first (not one-foot-in like other frameworks); orchestrator separates inference (server) and trainer (server)—they don't share GPUs or know each other.; Orchestrator manages run: environments → inference server → batches → trainer; trainer determines loss function from batch.; Inference server always uses latest model weights; rollouts finish long after start, enter first available batch (async RL = regionally off-policy).; Average 16 steps off-policy; validated stable up to thousands of steps (DPPO paper); pushing toward 10k in current experiments.; Async RL mandatory for agents: coding rollouts have long tails (30 sec–3 hours), synchronous RL wastes GPU waiting on slowest rollout.; Separation of concerns: scale environments, inference replicas, sandboxes independently; trainer doesn't care about environment/inference scaling.
  • Evidence: GLM-5 step time <5 min on 28 nodes, 131k context, long-horizon coding = 1000 steps in 3 days at ~$50k rental pricing.; Desegregated prefill, wide expert parallelism (multi-node MoE), FP8, router replay metadata offloading supported.; Torch Titan base (not Megatron) for hackability; research team <10 people, company <40 people.
  • Caveats: Off-policy RL = staleness; rollouts use older model → requires DPPO or similar for stability.; Not suitable if strict on-policy guarantees needed (though Will argues unnecessary for agents).; $50k for 1000-step run is not cheap for all use cases; only makes sense if training cost < inference cost amortized.; Async RL means you can't easily reproduce exact training run (rollouts interleave non-deterministically).; Torch Titan vs. Megatron tradeoff: hackability vs. maturity/features; assumes team can implement needed features.
  • Implications: Ken should not build synchronous RL for agents—it's a dead end; ask vendors if they support async and how far off-policy validated.; If Ken's agent rollouts are slow (3+ min), async RL is mandatory to avoid GPU idle time; synchronous RL will waste 10x+ compute.; Separation of inference/trainer means Ken can scale them independently (e.g., 10x inference replicas for long rollouts, 1x trainer).; $50k/3-day run is comparable to API spend for many enterprises—reframes post-training as operational expense, not research project.; Torch Titan choice signals 'move fast' culture; if Ken needs bleeding-edge model support (e.g., new MoE architecture), Prime can add it quickly vs. waiting for Megatron upstream.

Algorithm Composability: SFT, OPD, Self-Distillation, GRPO, Echo

  • Claims: Modern post-training algorithms (SFT, on-policy distillation, self-distillation, GRPO, Echo, Max RL) now swap via config by decomposing into 'algorithm' (data prep/scoring) and 'loss' (gradient target).; Factorization: rollouts from policy (on-policy/async RL) or teacher (SFT/OPD); advantage = reward-baseline (RL) or log prob ratio (OPD) or 1 (cross-entropy for SFT).; OPD: get reference log probs from teacher by sending sequences as prefill (ask for 1-token response = prefill only), use as advantage.; Self-distillation: send different prompt (hint) to teacher using renderers, get same sequence back, slice to original form for reference log probs.; Echo: two algorithm components, cross-entropy on environment tokens + RL on action tokens (world modeling with RL).; All configurable per-environment in TOML; add custom algorithms by writing class that assigns advantages to rollouts.
  • Evidence: Table of algorithms: actor = policy/teacher, advantage = reward/log prob ratio/1, loss = policy gradient/cross-entropy.; GRPO supported alongside OPD; can mix-and-match per environment.; Max RL paper patterns coming soon but factorization already supports it.
  • Caveats: Requires understanding of RL/distillation to configure correctly—not a magic 'choose algorithm' button.; Some algorithms need specific teacher models (e.g., OPD from o1 requires OpenAI API key, log probs if available).; Self-distillation with hints assumes you have good hints; poor hints = poor distillation.; Echo (world modeling) is research-stage; not clear when it helps vs. pure RL.; Factorization is cleaner but doesn't eliminate need for algorithm-specific debugging.
  • Implications: Ken can test OPD from o1 (if reasoning trace visible via hints) vs. pure RL vs. self-distillation vs. GRPO in same codebase, just config changes.; If Ken has teacher model (e.g., Claude Opus) and wants to distill into smaller model (Llama 3.3 70B), OPD path is now straightforward.; Self-distillation useful if Ken has prompt that improves model performance ('think step-by-step') and wants to bake it into weights.; Echo (world modeling) worth watching if Ken's agents need better environment understanding (e.g., debugging, planning tasks).; Research velocity increases: test new algorithm ideas in days by writing advantage-assignment class, not rewriting trainer.

Hosted Training Platform: Multi-Tenant LoRA and Full Fine-Tuning

  • Claims: Hosted training platform = Prime RL without managing GPUs; multi-tenant LoRA (live now, RL-focused) and full fine-tuning (coming in 2 weeks).; Multi-tenant LoRA: multiple users train on same base model weights (shared inference pool, hot-swap LoRAs), token-based pricing, no GPU reservation.; Full fine-tuning: supports full weight training, custom algorithms/losses, all Prime RL features; GPU-based pricing, auto-scaling, magic restarts.; Unified billing for sandboxes, judges, training, inference; dashboard with logs/metrics.; Develop environments on CPU (laptop), push to platform as environment packages, specify in configs.; Three levels of customization: (1) change reward function (environment space), (2) configure trainer (loss, LR, TOML), (3) hack trainer (algorithm, loss function, deeper).
  • Evidence: Multi-tenant LoRA uses same pattern as multi-tenant inference (one big model, shared KV pool, LoRA adapters hot-swapped).; Full fine-tuning supports 'changing as much as you want in Prime RL'—implies full code access, not just config.; Environment packages = portable, versioned environments (task set + harness + runtime); push to platform like Docker images.
  • Caveats: Multi-tenant LoRA limited to LoRA (no full weight changes)—fine for many tasks but not all.; Full fine-tuning requires GPU reservation (not token-based), so more expensive if utilization low.; Platform lock-in risk: if Ken builds on hosted training, migrating off requires recreating orchestrator/infra.; Auto-scaling, magic restarts, unified billing are convenient but opaque—Ken doesn't control low-level infra.; Coming in 2 weeks = typical startup schedule slip; assume 4 weeks to be safe.
  • Implications: Ken can start with multi-tenant LoRA (low activation energy, token-based pricing) for RL experiments, graduate to full fine-tuning if needed.; If Ken's team is small (no infra engineers), hosted training eliminates GPU management, auto-scaling, restart logic—massive time savings.; Environment packages are portable—Ken can develop locally, test, then push to platform; same workflow as Docker for apps.; Three-level customization model is good design: most users change rewards, some configure trainer, few hack algorithm/loss—matches skill distribution.; Lock-in risk mitigated by open-source stack—Ken can self-host Prime RL if needed, though loses platform convenience.

Notable Concepts & Terms

  • Verifiers V1: Complete overhaul of Prime's environment library, decoupling task sets (data/rules), harnesses (execution patterns), and runtimes (local/Docker/sandboxes) for flexible eval and RL workflows.
  • Prime RL: Open-source async RL training framework (Torch Titan base) optimized for frontier models at scale; separates orchestrator, inference server, and trainer; supports FP8, wide expert parallelism, desegregated prefill.
  • Task Set: Agent-agnostic data and rules for an environment; integrates with Hugging Face datasets, Harbor, NemoGym, OpenEnv; represents backend/server/state without owning model execution.
  • Harness: Execution pattern for model in environment (default tool loop, RLM, CLI agents, Langchain/DSPy); decoupled from task set; plugs into runtime via interception server.
  • Runtime: Where harness code executes (local, Docker, Prime Sandboxes, custom); uses async IO + subprocesses; UV script pattern for dependency isolation.
  • Interception Server: Fake base URL (OpenAI/Anthropic compatible) given to harness; intercepts requests, adds log probs/temperature, forwards to RL trainer; harness doesn't know it's doing RL.
  • Renderers: Standalone library treating chat templates as programmable artifacts; manages message/token duality, prevents retokenization drift; works with any inference engine; inspired by OpenAI Harmony.
  • Trace Graph: Data structure storing rollout records at both message level (logical text) and token level (IDs) with branching support for sub-agents; prevents trainer/inference mismatch from retokenization.
  • Group Rewards: Rewards comparing multiple rollouts in a batch (e.g., bonus for shortest correct answer, pairwise judging); critical for agent RL but ignored by most frameworks due to decoupled architectures.
  • Async RL: Off-policy RL where rollouts finish long after starting, enter first available batch using latest model weights; prevents GPU blocking on slow rollouts; Prime runs ~16 steps off-policy average.
  • DPPO: Distributed Proximal Policy Optimization variant for stable async RL; Prime validated stable up to thousands of steps, pushing toward 10k; handles staleness from off-policy rollouts.
  • On-Policy Distillation (OPD): Algorithm where rollouts come from student model but advantage is log prob ratio vs. teacher model; student learns to match teacher on its own sampled trajectories.
  • Self-Distillation: OPD variant where teacher is same model with different prompt (e.g., hint added); student learns to internalize prompt improvement into weights.
  • Echo: RL algorithm with two objectives: cross-entropy loss on environment tokens (world modeling) + RL on action tokens; research-stage for better environment understanding.
  • Multi-Tenant LoRA: Training pattern where multiple users share base model weights and inference pool; each has own LoRA adapter (hot-swappable); enables token-based pricing without GPU reservation.
  • Router Replay: Storing MoE routing decisions per rollout (metadata for every layer, multiple experts per layer) for training; nasty systems problem due to storage/memory overhead; Prime offloads to object storage.
  • Conciseness Bonus: Group reward giving higher reward to shortest correct answer in batch; solves problem of unknown optimal length per task; prevents 'thinking forever' in open models during RL.
  • UV Script: UV tooling pattern for running Python scripts with isolated dependencies; used by Prime for harnesses, tools, user simulators, grading to avoid dependency conflicts.
  • User Simulator: MCP server acting as simulated user in multi-turn agent training; model sees it as user (not tool); uses LLM + context to respond realistically; critical for multi-turn agent evals/RL.
  • Retokenization Drift: Phenomenon where message→tokens→message changes output due to many-to-one tokenization (chat templates strip whitespace); causes trainer/inference mismatch in RL; solved by dual message/token streams.

Operator Notes / Why Ken Should Care

  • Prime Intellect is credibly positioned as infrastructure vendor for frontier-model post-training: 10k+ GPUs, <40 people, research team <10, real customers, open-source stack, hosted platform. Rare combo of open research + viable business model.
  • If Ken is evaluating RL vendors or build-vs-buy for post-training, Prime is the only open-source option at this scale. Key differentiators: async RL (16 steps off-policy validated), group rewards, algorithm composability, trace graphs, renderers, interception server.
  • The 'environments = evals = RL = SFT data gen' insight is architecturally sound and solves the 'can't iterate' problem. Ken should adopt this mental model even if not using Prime stack—it clarifies what to build.
  • Hosted training platform (multi-tenant LoRA live, full fine-tuning in 2 weeks) is the GTM wedge: low activation energy (no GPU management), token-based pricing (LoRA), upgradeable to full fine-tuning. Smart land-and-expand motion.
  • Technical depth is real: retokenization drift, router replay metadata, trace graphs, renderers library. These are not 'nice-to-haves'—they're the difference between RL that works at scale vs. RL that breaks subtly after 500 steps.
  • Async RL is the right architecture for agents (long-tail rollouts 30 sec–3 hours), but 16 steps off-policy is aggressive. Ken should ask: what's the stability/sample-efficiency tradeoff? Are there tasks where this fails?
  • GLM-5 benchmark ($50k, 3 days, 1000 steps, 131k context) is the headline number but lacks comparison. What's the baseline? How much does OpenAI spend for equivalent? What's the win rate improvement? Need more context.
  • Team size (<40 people, research <10) is both a feature (agile, move fast) and a risk (key-person dependency, support quality). Ken should assess: who are the 3-5 must-retain people? What happens if they leave?
  • Open-source strategy is both moat (ecosystem lock-in, community contributions) and risk (competitors can fork, replicate). Prime's moat is hosted platform + customer relationships, not code. Ken should evaluate platform stickiness.
  • Hiring aggressively in SF signals growth/funding. 'Exciting news later this week' suggests funding round or acquisition. Ken should monitor for changes in strategy, pricing, roadmap priorities.
  • The 'open super intelligence stack' framing is ambitious but increasingly credible as models hit frontier performance. If Llama 4, Qwen 3, DeepSeek v4 continue improving, Prime's thesis (open post-training > closed API) becomes more valuable.
  • MCP integration (tools, user simulators) is forward-looking—positions Prime for agent world where harnesses are composable, not monolithic. Ken should watch whether MCP becomes standard or fragments.
  • Cookbook repo (alpha) is critical for adoption—tutorial/example quality determines whether engineers can self-serve. Ken should evaluate cookbook completeness: does it cover Ken's use cases? Are examples production-ready?
  • Three-level customization (rewards, config, algorithm/loss) matches skill distribution well: most users change rewards (product eng), some config (ML eng), few hack trainer (research eng). Good API design.
  • Interception server pattern (fake base URL) is clever but fragile—requires harnesses to be API-compatible, adds latency, can be confusing to debug. Ken should test whether this works for Ken's harnesses (Langchain, DSPy, custom).
  • Renderers library (standalone, works with any engine) is underappreciated—solves real problem (retokenization drift) and is reusable outside Prime. Ken should adopt even if not using Prime RL.
  • Group rewards are the 'secret sauce' for agent RL (conciseness bonus, pairwise judging) but require careful reward design. Ken should budget time for reward engineering—it's not plug-and-play.
  • Full fine-tuning (coming in 2 weeks) will determine whether Prime competes with Modal, RunPod, others on training infra. Key questions: pricing vs. alternatives? Auto-scaling reliability? Magic restart = checkpointing?
  • Multi-tenant LoRA token-based pricing is strong GTM (low friction, pay-as-you-go) but only works for LoRA use cases. Ken should clarify: when is LoRA sufficient vs. full fine-tuning needed? What's the transition path?
  • Post-training flywheel (evals → RL → deploy → feedback → refine envs → repeat) is the right mental model but assumes Ken can close the loop (capture feedback, update environments). Most companies struggle here—Ken should assess feasibility.
  • Will Brown's presentation style: technical depth, open about tradeoffs, grounded in production experience. Credible as applied research lead. Ken should engage if considering Prime—Will seems accessible.
  • Prime's compute marketplace (10k+ GPUs, global data centers) is both a strength (vertical integration, control) and a risk (capital-intensive, commoditized). Ken should clarify: does Prime own GPUs or broker? What's margin structure?
  • The 'we figured out a good business model to do real open research' claim is bold. Ken should validate: what's ARR/revenue? Who are customers (names, use cases)? Is this sustainable or VC-subsidized?
  • Torch Titan (not Megatron) choice trades maturity for hackability. Ken should assess: is Prime's team strong enough to maintain/extend Torch Titan? What features are missing vs. Megatron? Does this limit model support?
  • Router replay metadata offloading (object storage) is a real systems problem for MoE RL. Ken should ask: what's the storage cost? Latency impact? Does this work for 1T+ parameter models?
  • Desegregated prefill (inference optimization) is standard now but worth highlighting—signals Prime is keeping pace with inference SOTA (VLLM, TGI). Ken should verify: what's prefill latency? Token throughput?
  • FP8 support (training + inference) is critical for cost/efficiency at scale. Ken should ask: what's the accuracy/stability tradeoff? Does FP8 work for all models or just some?
  • Wide expert parallelism (multi-node MoE) is hard to get right. Ken should ask: what's the communication overhead? Does this scale linearly? What's the largest MoE Prime has trained?

Watch Map

  • timestamp unavailable: Timestamps not provided in transcript; chapter markers unavailable.
  • 00:00–05:00: Intro: Prime Intellect overview, open super intelligence stack, compute marketplace, Lab platform, Intellect model series.
  • 05:00–10:00: Post-training loop: environments as evals, SFT, RL, distillation; expert models then distill; deploy and iterate flywheel.
  • 10:00–20:00: Verifiers V1 deep dive: task sets (Hugging Face, Harbor, NemoGym), harnesses (tool loop, RLM, CLI agents), runtimes (local, Docker, sandboxes); rewards, metrics, group rewards (conciseness bonus).
  • 20:00–30:00: Tools, user simulators (MCP), interception server pattern (fake base URL, harness doesn't know it's RL), eval CLI hot-swap (local to cloud), trace graphs, renderers library (message/token duality, retokenization drift, OpenAI Responses stateful APIs).
  • 30:00–40:00: Prime RL architecture: orchestrator, async RL (16 steps off-policy), inference/trainer separation, GLM-5 benchmark ($50k, 3 days, 1000 steps, 131k context), async RL rationale (long-tail rollouts 30 sec–3 hours), DPPO stability (thousands of steps).
  • 40:00–50:00: Parallelisms: expert parallel, context parallel, wide expert (multi-node MoE), desegregated prefill, FP8, router replay metadata offloading; Torch Titan (not Megatron) for hackability; algorithm composability (SFT, OPD, self-distillation, GRPO, Echo) via loss + algorithm factorization.
  • 50:00–end: Hosted training platform: multi-tenant LoRA (live, token-based pricing), full fine-tuning (coming 2 weeks, GPU-based, custom algorithms/losses), environment packages, three-level customization (rewards, config, trainer hack); hiring (SF, <40 people, research <10), open research + real business model.

Source/Metadata

  • Title: The Prime Intellect Stack — Will Brown, Prime Intellect
  • Transcript words: 17178
  • Duration seconds: 2812
  • Timestamp note: No timestamps or chapter markers present in transcript; watch_map is inferred from content flow. Video duration ~47 minutes (2812 seconds).

Transcript

9528 words en Processed in 604.4s

[SPEAKER_00] Hey guys, how's it going? Thanks for showing up. This was a little bit of a last minute assembly. A few days ago, I was talking to Swix, and I was like, hey, can I still do a workshop? And he's like, we got one slot left, it's Monday at 4:30, and I was like, I'll take it. And then, yeah, I wanted to just do a bit of an update on some of the stuff we've been building at Prime Intellect. So, if you don't know me, hi, I'm Will Brown. I lead applied research at Prime Intellect. We do a lot of stuff around every part of the AI research infrastructure stack. Today is going to be about post training, which is where I spend a lot of my time thinking and building. And especially want to be talking about the post training tools that we build that are fully open source: the verifiers and Prime RL libraries, which go hand in hand, both on the environment side and the training info side, and show off some things we've been cooking over the past few months that I think is the way that things have evolved as the agent use cases have gotten more complex, but also clearer in terms of what people want out of agents, and the sorts of things that are needed to do the post training that is needed to power the real world applications people are building nowadays. And so, broadly at Prime Intellect, our goal is to make doing large scale open source AI research easier, and to enable companies to train their own models and deploy them and have them improve based on the scenarios that they actually see in production in terms of use cases for applications and products and internal tasks and workflows, and to give people an option to not just use the open source models that are getting quite good, but to take them and make them even better on their own use cases. And so, we use the phrase "the open super intelligence stack" to describe what we mean by this. And I think when we said this phrase like a year ago, it felt a little more marketing, and now it feels a little more like, oh yeah, that's what it is. The models are getting very, very good. They are superhuman in many ways at lots of things. And what we want to do is give people an open toolkit that they can use to do real training with them, and to have the control that they need to deploy it where they need to deploy it, and customize it as much as they need to, to get the job done. And so, this is the stack that we build. And it all sits on top of compute. So, we operate a global marketplace of data centers around the world. A lot of these are quite large data centers. We currently operate over 10,000 GPUs, many in hundreds or thousands within a cluster. We have our primary all training framework. We have environments built with the verifiers library and our environments hub platform. We have our platform for research workflows that we're now calling Lab, which is an assembly of many pieces, including the environments hub, hosted training evaluations, as well as inference and sandboxes. And all of this is in service of empowering and unlocking frontier model training. And so, we do this ourselves. We have our Intellect model series with some exciting things there coming soon. And we also train models with our customers, where we have lots of people we work with whose goal is to do large scale model training on their own workflows. And so, to do all of this, we need to give people the tools that they can assemble into the pipelines, the workflows, the research that allows them to actually get the results that they need at scale with everything they need to do. And so, this talk is going to be about going deep into verifiers and Prime RL and showing off some of these new things, but all under the umbrella of what does modern post training look like? What does it mean to take a model and train it to be better at your task? What are all the parts? What are all the gotchas? And how do you orchestrate this into a system that is actually easy for people to use without needing to go build a massive research team and to be able to have it be accessible for sorts of things that anyone who's an AI engineer at any startup or enterprise that wants to invest in post training can actually do. And so, there's a cookbook repo that is kind of an alpha release right now. It's still changing a bit, but it's a preview of all the stuff we've been building over the past several months. And so, today we'll be following along that framing a good bit. And so, I think the first thing we'll talk about is what is an environment. People talk about environment in the context of RL and think of RL environments, but environments are more than just for RL. They're for all sorts of things in post training and evaluation. We're going to talk about what we're going to call the V1 version of the verifiers library, which is a full overhaul. Everything else still from before still works, but we kind of wanted to redo it all. And so, we have a new way of doing everything that we think is going to make a lot more sense and be a lot more powerful for what people are looking to do going forward, as well as talk about how Prime RL has evolved as a library. And so, Prime RL is our full stack open source training framework to support asynchronous reinforcement learning. And we've got a lot of fun new bells and whistles to show off in terms of both scale and features. A lot of this is in service of custom algorithms, so making it much easier to do the sorts of things that people are interested in for modern post training. If you have been following the news on policy distillation or self-distillation or all these other fun new algorithms that people are coming out with, it is the age of research indeed. And we don't want to just train small models. We want to train big models. We want to train them really efficiently because as models get bigger, the compute starts adding up. And we've got a lot of fun new bells and whistles to show off in terms of both scale and features. A lot of this is in service of custom algorithms, so making it much easier to do the kinds of things that people are interested in for modern post training. If you have been following the news on policy distillation or self-distillation or all these other fun new algorithms that people are coming out with, it is the age of research indeed. And we don't want to just train small models. We want to train big models. We want to train them really efficiently because as models get bigger, the compute starts adding up. And if you want to make this accessible to people, especially if you want to be able to iterate on it, it has to be fast. It has to be cheap. It has to be affordable and reliable. And all of these funnel into our lab platform. And we'll talk about both some of the things that we've already released there, as well as some things that are coming soon. And so, the post training loop in my mind revolves around environments in the sense that environments are a language for specifying what you want your model to do. They are an encapsulation of the data you might have, the scenario you might want your agent to be in, the way it'll interact with that environment, as well as how to score what good looks like, to determine what was good and bad. And often this is the first thing you want to do with an environment is just evals. And so I think a lot of people are maybe nervous about getting into post training. They're like, oh, it seems like a lot of work. There's a whole new tool chain. What if I'm already using the frontier models and I want good results out of them, or I'm getting good results out of them, or I want to see what I can do at the harness level first, or prompt optimization. And that's all good. We're not necessarily asking people to just throw everything away. I think in many cases what people will find and what we see with our customers is that the systems that work best for them involve using both. And you want to be able to make these decisions about where is the right place to train, where's the right place to use a frontier model that is available via some API. And so evaluations are very key to this. And so evals are the thing that opens the door to post training. And so environments and evals are essentially the same thing. But once you have evals, now this is the same kind of unit of logic that you actually need to do post training anyways. And so building evals is just good for your product hygiene no matter what you're doing. If you want to decide whether to use GPT or Claude or decide do you need Opus or Sonnet or Mythos for a task, like if you want to min max on intelligence versus dollars, evals are a very good way to do this. But evals also then unlock this flywheel. And in terms of modern post training, I think historically people have done SFT then RL as the main frontier model recipe. Although unpolicy distillation has certainly found its way into a lot of workflows. And I think some people are also very eager about algorithms like self-distillation. We can talk a bit about that and when it makes sense and when it doesn't. But in particular, one area where it does make sense to do the unpolicy distillation thing is when you're training experts where you have multiple different things you want your model to be good at. And people have found that if you have a bunch of different environments that are all different things and you want to have one model be really good at them, a nice way to do this is train individual RL experts on top of the same base model and then do distillation from those teachers into the same checkpoint. That just generally ends up being more reliable. And then once you have this, you want to deploy the trained model which could be a full base model with full weight training or it could be a LoRa adapter and you want to serve this at scale. And ultimately what is useful about this whole process is it's not just a thing you do once. I think some people also say, oh, why should I do post training if the frontier models are going to get better? Well, your model should get better too. It's not like everything's going to get better. The point of this is to have flywheels that make everything get better. And so what you really want is to be able to not just post train like today, but to be able to have this iterative process of model refinement and the sort of thing where you can have the training compute end up being a pretty small fraction of your overall inference budget that you amortize out such that your model is always getting better and better as you are getting more signal from the real world. And getting the signal from the real world isn't trivial. That's kind of largely an open question as to how you go about getting information from real world feedback into your environments. It's an engineering problem. It's a research problem, but it's the sort of thing we're all here at this conference to think about and learn about. And so I'll touch on that in the talk as to how we think about this. But really the goal is going to be thinking about what do these tools look like? How do you actually do this? What are the parts? And how do we build it? And so environments as evals, what is an environment? I think it's useful to decompose environments into tasks and a harness. And this is foreshadowing some of the refactoring we've done in verifiers over the past months. If any of you have used the verifiers library before, you may be familiar with the multi-turn environment pattern, our tool environment pattern, where there's one loop that is owned by the environment that you can plug in various tools into. And that was really great for a very long time for getting started for people, especially back in the day when people were mostly just trying to graduate from single-turn into multi-turn tool use. But what we found, as we iterated on different patterns and extended it in various different ways, we found ourselves repeating a lot of work of adding patterns for a CLI agent or adding patterns for MCP. And we wanted to be able to step back and rethink how should an environment work. And what it really is, there's a notion of a harness. There's also a notion of a task. And I think one of the reasons that this is subtle and tricky and was a design problem that we went over many iterations over the past six months, really, is certain things, it's not clear where they live. The day when people were mostly just trying to graduate from single-turn into multi-turn tool use. But what we found, as we iterated on different patterns and extended it in various different ways, we found ourselves repeating a lot of work of adding patterns for a CLI agent or adding patterns for MCP. And we wanted to be able to step back and rethink how should an environment work. And what it really is, there's a notion of a harness. There's also a notion of a task. And I think one of the reasons that this is subtle and tricky and was a design problem that we went over many iterations over the past six months, really, is certain things, it's not clear where they live. There's certain things that might belong to the harness and might belong to the task. There are certain tools that, in some cases, I want this task I'm going to do to use a certain tool. In some cases, the harness has certain tools. Same with skills or system prompts or many other pieces of the puzzle in assembling the full world that your agent is going to be operating in or your model is going to be operating in. But ultimately, we're going to call all of these parts of the environment. And the goal of this is to have some notion of verification, where you plug in a model into the environment, which includes a harness, you give it a task, it does a rollout, and then you verify what it did. And this same process works both for evaluation offline, just understanding which model is better, as well as for doing reinforcement learning, RL, as well as for generating data for SFT. I think in many cases, people think of SFT as this thing where they want to upload a dataset, but really often what they're doing there is they're essentially cobbling together something that's essentially an environment, and they're doing rollouts in it, and saving it offline, and putting it in one format, and uploading it, and then changing it to another format, and then plugging it into a trainer. And the way we've approached this is saying, well, you can just cut all that out and just treat it like a problem where you're doing rollouts in an environment. Just in this case, there's a teacher. And the teacher can be another model, it could be replaying from another dataset, but ultimately it's about collecting rollouts, and training on those rollouts. And then on policy distillation, again, takes the same form, where you are doing rollouts in an environment just as you would for RL, but the scoring is from a teacher, and the log problems of the teacher, the likelihood of the teacher, rather than the reward signal itself. And so these all, in our framework, are Python packages. So you can have any dependencies you want. You can pull data from anywhere you want. There's a lot of flexibility that we've unlocked in terms of what these tasks can look like, what these harnesses can look like. And our goal is to just make this a really flexible toolkit for all the kinds of evaluation things people want to do, both for API models as well as for post-training. And so Verifiers V1 is what we're calling it, which is, it's not actually released as V1 yet, but we took inspiration from VLLM doing this, and decided that we were going to have this be the new pattern that we want to have everything be centered around. And the key pieces we broke things down into were a task set, a harness, and a runtime. And so these are all composable. You can mix and match them, and they're all individually loadable in different ways. But the way to think about it is that task sets are the data and the rules of what should be done that are agent agnostic. So they're the sort of thing you could plug an agent into. And we wanted to take a very general approach in supporting a lot of the great work being done throughout the ecosystem. So we integrate natively with Huggingface datasets, with Harbor, with NemoGym, and OpenEnv, and most other tools that you see out in the wild that are under the umbrella of an RL environment, we would call these a task set. We generally have found that it's useful to have these be harness agnostic, where they represent the backend or the server or some state that you're interacting with, but they don't own everything about what the model is doing. And so it doesn't, in some cases, it doesn't make sense to plug a model into a task set, especially because we're gravitating towards an agent world where everything is running in a terminal, or it has skills, or it is using CLI tools. And these things often look more complex than just basic loops. But we also want to support basic loops. So we want to allow both the old way of doing things and the new way of doing things. And so everything that was the old way is now the default harness, where it's system prompt and tools in a loop. But the harness pattern also supports much more flexible execution of things like recursive language models, or CLI agents, like Codex, Cloud Code, OpenCode, or classics from the resource literature, like Mini Suite Agent, or building your own with arbitrary Python libraries like Langchain or DSPy. And so we've been able to decouple these into a pattern where you get to write your harness independently of your task set. There are some basic sanity checks about properties that harnesses either do or don't support and task sets do or don't require. And these click together. And the runtime is where this executes. And so we've still been embracing a lot of the async IO patterns from before, but we've leaned a little more into having things be sub processes, where you can still run everything locally. You don't have to use sandboxes, but you can use local Docker, or you can use our own Prime sandboxes layer. You can use any other sandbox layer you'd like, or build from scratch. And so the harness, the runtime backend, just is a place where the harness can run its code. And so the harness just needs to be able to run code somewhere as a script, essentially. We've used a lot of the UV tooling where UV script is a very powerful pattern to mix and match and contain dependencies. But what happens is once you plug these together, you run a rollout on a task from a task set, and you get a trace. This is live on the verifiers main branch for prime intellect AI slash verifiers on GitHub, as well as it's released as a dev release. The stable main release will be coming to PyPI any minute now, but you can install the dev and play around with it if you want. And so what do these look like? So tasks are just like a row of a dataset. And the very basic version of it is you just are loading a dataset from Hugging Face or anywhere else. And so the new pattern here is from verifiers v1, just to keep the old stuff separate. The old stuff still works just fine, but this is how we have been able to decouple and iterate on the new version. AI slash verifiers on GitHub, as well as it's released as a dev release. The stable main release will be coming to PyPI any minute now, but you can install the dev and play around with it if you want. And so what do these look like? So tasks are just like a row of a data set. And the very basic version of it is you just are loading a data set from Hugging Face or anywhere else. And so the new pattern here is from verifiers v1, just to keep the old stuff separate. The old stuff still works just fine, but this is how we have been able to decouple and iterate on the new version. As well as we've really embraced this decorator pattern. We found it to be very useful. But we also, if you were a rubric fan, we killed rubric. Didn't make sense anymore if you were using old verifiers rubric patterns. But still, you have functions and loaders. We are very heavy on Pydantic, so everything is super typed. We have lots of powerful config features where you can have everything in a toml file, you can override it in the CLI, and everything is clean and guaranteed to type check at a validation time rather than waiting for something to fail later down the road. And so examples of this are things like suite grep where you can do agentic code search, you can do the classic games like Wordle, you can do search over documents with judges, you can do complex things like Harbor that support a lot of popular benchmarks now that need agents running in a terminal. And all of these are going to be combinations of the task set pattern with pick your own runtime and pick your own harness. And so rewards and metrics, I think, are also just functions that take in the records of what's happened in a rollout and return numbers. So rewards are the main thing that will drive progress in RL. Metrics are just logging what has happened, so counting tool use and counting errors. These sorts of things are very useful to be able to expose in your dashboards. And then group rewards. I think this is something that we have fought hard to make sure still is first class because we see it as very important to a lot of the research pattern people want to do, but I think is also ignored in a lot of tooling out there where in many RL frameworks it's actually quite hard to do group rewards because things are very decoupled and things assume that all rollouts are going to live independently and that they don't need to talk to each other. But there's a lot of things where you really want to do pairwise judging or you want to do ranking or you want to give a bonus to the shortest correct answer in terms of tokens used. And so these sorts of things are really flexible in terms of the way we've really designed for flexibility in supporting the things that we see as the most exciting papers we've read or all the algorithms that we think people may want to innovate on while still allowing people to have the core primitives that they expect out of an RL framework. And so in group rewards I think this conciseness pattern is one that I find very useful a lot. I think a big pattern that comes up a lot when people are doing post training is models will love to think and think and think if you let them. And if you don't give them some pressure to be more efficient I think a lot of people will notice that open models often have really long chains of thought because on one hand this is a useful strategy for a model but it's also the sort of thing that will grow out of control if you don't counteract it. And so in reward design one of the big things people will want to do is something like a length penalty or a conciseness bonus. And so one of the reasons this is tricky is because you don't know the optimal length for a problem. Like if I give you a math problem I could say solve it in less than n tokens but also who knows what the right n is. It's also going to change as the model gets smarter over time. It's going to be different for every problem. And so you can't know this up front and the only way you can do it is take advantage of variance. And so one of the nice things about RL is you have multiple samples typically and this allows you to use the fact that you have multiple samples to shape the reward. And so if you have multiple robots in a group what you could do is look at all the ones that were the correct answer or just all the ones in general and give a bonus to the ones that are the most concise. Where if you also have a correctness reward like these are going to ensure that you're both incentivizing correctness as well as incentivizing efficiency. And so juggling multiple objectives simultaneously is one of the hard challenges in RL in reward design. But doing things like group level comparisons and these sorts of bonuses are quite useful in many cases. I also want to talk about tools and user simulators which I think have been becoming more important in a lot of complex applications where you have models that are in many cases there's a core agent harness but there's also in many cases you are putting a model in a setting where it's going to be in some task where a user is giving it additional tools. Whether these are in some cases you might want to model these as skills in some cases you might want to model them as MCP servers. We use MCP as a back-end framework that can interact with the runtime both for tools and for user simulators. So user simulators especially if you want to do training where there's a user in the loop you don't want just your agent to go do some tasks but you want to be able to do some tasks that involves understanding how a user will interact. You can essentially have this user be an MCP as well where we make it so the model sees it as a user not as a tool but behind the scenes it is a server that has some script or something and it has some LLM that is going to get some context and it's going to be able to be a user in the context of a rollout. In many cases benchmarks have found that this is very useful for simulating the realism of having a multi-turn setting where there are users in the loop especially given that people are building products now where there are these users in the loop and so you want a first class way to incorporate this into your RL environments. And so we've found that it's useful to have all of these things be modular and pluggable and so the harness can connect to each of these which run as a UV script. We also have UV script support for grading in addition to the basic reward function patterns and we also are using this idea called an interception server. So the harness we want people to be able to use real harnesses and not have to break the harness and retrofit it into an RL harness and so ideally we don't know anything about the harness code and so the pattern that we use with the interception server is that we are responsible for giving each harness rollout a fake base URL which could be OpenAI compatible or Anthropic compatible. The harness just thinks it's talking to some endpoints so any harness that can just talk to some endpoint we're good to go and then of these which run as a UV script. We also have UV script support for grading in addition to the basic reward function patterns and we also are using this idea called an interception server. So the harness we want people to be able to use real harnesses and not have to break the harness and retrofit it into an RL harness and so ideally we don't know anything about the harness code and so the pattern that we use with the interception server is that we are responsible for giving each harness rollout a fake base URL which could be OpenAI compatible or Anthropic compatible. The harness just thinks it's talking to some endpoints so any harness that can talk to some endpoint we're good to go and then we intercept each request. We can do some backend maneuvering to make sure that we're getting the log probs and setting the right temperature and then we send this to our inference server with the RL trainer and then as it completes we send back the request and so the harness doesn't know that it's doing RL. The harness just is a harness running as if it would be running in a real world environment and so you can move between the RL setting and the deployment setting where your harness is just code. It doesn't need to be anything specialized to verifiers or RL. We also have found it really useful to go from local to global and have this hot swap pattern so we have this new eval CLI where you can choose the harness you want. You can have a task set where you're saying I'm going to run an eval on this set of tasks. I want to use recursive language models and RLM. I want to use Codex. I want to run it locally. I want to run it in sandboxes. I want to run it in Docker. These are all interchangeable and so we found that this is super useful in the iteration loop as you go from testing something out on a small scale towards scaling it up towards being able to understand questions like what's the best harness for this model? Does this harness generalize across tasks? As well as being able to have the convenience of local prototyping where you can run things fast on your MacBook without having to wait for a cloud job to finish but also you can go right to the cloud when you need to. And so there's been a lot of patterns we've had to innovate on as we've done this overhaul and we were quite happy with how it's turned out. It's made our lives a lot easier for both client projects and research and just being able to have a lot more flexibility and power and control over the kinds of agents we want to be training. And so one of the things behind the scenes is what we call the trace graph. And so we had been having this grow out of control in terms of the old way of doing things and we decided this was another opportunity to overhaul our system to have really good support for sub-agents and parallel branching trees while also still preserving the linear sequential dependencies that you need for RL with careful token control. And so here there's a notion of branches that are at the message level. So conceptually the things that matter logically in environment space and in harness space are messages which are just text. The harnesses don't think about tokens but if you've done any RL experimentation you may have encountered issues where re-tokenization or messages if a model will say something and you turn it into text and you put it back through tokenizer it can change a little bit because tokenization is many to one. And so this causes lots of very subtle numerical problems especially late in large-scale training runs. And so you want a really nice back and forth between messages and tokens. And so the trace data structure that we created here partly is to enable this where we can store things both at trace level and then map them back into token level in the right sequences as needed. And we also released a library called renderers recently which is a standalone toolkit that anyone can use that we have found to be the sort of thing that we're working with some of the inference tooling to support. So renderers are all about rethinking tokenizers and chat templates where behind the scenes it's making calls to the tokenizer but chat templates if people have spent time debugging with them it sucks. Jinja is awful. It's very painful and there's so many subtle things that we kept running into where a model would sometimes have an extra newline and the chat template would strip it out and this would cause a mismatch in your trainer and inference that would either force you to go off policy because you now have a trainer inference mismatch or it would cause a logical branch where a thing that is a branch in token space even though it shouldn't be in logic space because of tokenizer subtleties. And so renderers as an abstraction was pioneered by OpenAI's Harmony with the GPT OSS release and used prominently in Thinking Machines cookbooks as well for Tinker but we found it was useful to make it a standalone thing and so this is a Python library that doesn't depend on any other prime stuff. You could use it with any inference engine you want just as a standalone thing that is really designed for being able to manage this token in token out concatenation without thinking about it too much yourself because we turn each of these chat templates for these popular models into programmable artifacts where you can do things like look up a history of secret you can use the history of a trace to be able to understand what is the right tokenization do I essentially have a logical prefix hit in message space even if I don't in tokenization space after retokenizing and so this is the sort of thing where I think people have gone back and forth on whether they want LLM APIs to be stateful in general. I think a lot of people were hoping that we could just have every model API be stateless. I think maybe people are less concerned about this now because we're moving towards this agent world where agents themselves are going to be stateful APIs. But I think this has revealed to us going through all the things here why OpenAI Responses decided to be stateful. There are some unavoidable issues that come up when you're doing large-scale agentic rollouts where you need to manage this very carefully and it's unavoidable because of how tokenizers work. And so you want to maintain these dual streams of the logical text and the tokens. And you want these to be cleanly interoperable where users don't have to think about the tokens very much but the trainer gets to see everything nicely in token space as well as the inference engine. And so from the harness interception server we have clients that can be used both for training and inference and so you can swap between these modes without thinking about it because certain models don't need to in a training setting you need to be able to get log probs and set the temperature. Some model APIs won't let you do this. They won't return log probs because OpenAI models with reasoning won't show you the reasoning trace. So there's no way they give you the log probs for everything and so that's fine. But the trainer gets to see everything nicely in token space as well as the inference engine. And so from the harness interception server we have clients that can be used both for training and inference and so you can swap between these modes without thinking about it because certain models don't need to. In a training setting you need to be able to get log probs and set the temperature. Some model APIs won't let you do this. They won't return log probs because OpenAI models with reasoning won't show you the reasoning trace. So there's no way they give you the log probs for everything and so that's fine. It's just eval only and so we have this client layer. You can go between eval and train to be able to support all these models. But we still use the interception server pattern either way because it allows us to have this notion of a dialect where you can choose OpenAI chat completions or responses or Anthropic and all of these are easily supportable as translation layers between a raw request into something that'll get passed through a renderer potentially if you're on the training client side and formatted into a message via tokens. And so this brings us to PrimeRL. PrimeRL is our training framework that consumes the environment. So once you have an environment with your task set and your harness and your interception server and your runtime and your render and all those things, this plugs into what we call the orchestrator. And so PrimeRL has been async from the ground up. So I think async RL is one of those things that people were kind of one foot in one foot out. And a lot of training frameworks if you see them will still support synchronous training. Some people I think have their reasons for wanting to do synchronous training. I don't agree with them. I think you kind of want to bite the bullet of the off-policyness anyway for reasons that come up with agents. In terms of you want to be able to overlap long rollouts and not always be waiting on your slowest rollout. And this means you can't be fully on-policy unless you want to accept always waiting on your slowest rollout. And so this is really why we went all in on async. The orchestrator's job is to allow the inference and trainer to be separate processes, separate servers. They don't share GPUs, they don't really know about each other all that much, they just consume from each other. But the orchestrator's job is to manage the run. The orchestrator will make sure that the environment is running with the endpoint mapping to the inference server. It'll do rollouts. It'll package these up into a batch. It'll send this back to the trainer and it'll be up to the trainer to figure out what to do with the batch, which will be printing some sequence to feed into a loss function based on the specification. And so the server pattern we use for environments is an engine that can send requests to the inference and send batches back to the orchestrators. It's very client-server. And we found that this is a really useful way to allow scaling concerns to be decoupled. For example you can have a lot of environments running or you can have one environment running. You can have a bunch of inference replicas or you can have one. You can have sandboxes or no sandboxes. And the trainer doesn't care about this, the inference doesn't care about this. It's just separation of concerns at a system level. It allows you to not really worry about these things as combined units, versus in some cases people will want to have training and inference on the same stack where, especially the closer that you fold in your logic all in one, then sometimes you can't even run your environments if you're not doing RL, but that means you can't really experiment. You also can't use them as evals. There's a lot of reasons why pulling everything apart into these pieces and having nice APIs for them to talk to each other makes your life way easier. And we've also been scaling it. We've been doing a lot of work on GLM 5 series and Claude 3.5 2.5, 2.6 series just to make sure that we can do really good large-scale RL efficiently. And so some results we found recently, this was the run on GLM 5 before the 5.2 one came out, it supports 5.2 as well. But on the latest PrimeRL version we can do a GLM 5 step on 28 nodes in less than five minutes for long horizon coding tasks with 131k context. Which means you can do a thousand step run in three days and that costs for rental prices about $50k. And so $50k is not cheap. But if you're doing a full run on a frontier-sized model on a proper real world agent environment, it's a lot cheaper than what OpenAI is charging for it. It's a lot cheaper than some of the clusters people are selling and it's the sort of thing that starts making sense for a lot more enterprises if you can actually do this. It becomes pretty justifiable if you can get to the point where you have the tooling chain to be able to build the reward valves and do all this stuff. And so this is the sort of thing where let's say you wanted to have a bunch of tasks that are representative of your coding workflows and you want to do a big RL run, this is a pretty big RL run. But it's also the sort of thing that people spend this much on tokens in a month sometimes. And so you can start finding a lot of savings if you think about doing large-scale post-training. And that's what we're here to help people do. More on the async side, one of the reasons why you really want to do async is that there's a long tail of how long your coding agents take. If you fire up a coding agent task, think about your code or Codex tasks, how many minutes is it going to take? Sometimes it'll be 30 seconds, sometimes it'll be two minutes, sometimes it'll be a goal that goes for three hours. And these can all be rollouts. So one of the goals of async RL is to have your forward progress speed not be tied to the speed of your individual rollout. And so this means you can allow your rollouts to finish long after they start and go into the first batch that can accept them, even if that is from a much different copy of the model. And so the inference server is always taking the latest version of the model. What people have generally found, and what we've done with our experimentation and found as well, is that you can go regionally far off-policy. I think 16 is where we typically are often operating, like an average. But this means you can have a lot more room to not worry about certain things about speed. Not be tied to the speed of your individual rollout And so this means you can allow your rollouts to finish long after they start and just go into the first batch that they can accept them Even if that is from a much different copy of the model And so then the inference server is just always taking the latest version of the model and so What people have generally found and what we've done with our experimentation and found as well is that you can go regionally far off policy I think 16 is where we typically are often operating is an average But this means that you just also can have a lot more room to not worry about certain things about speed You don't need to worry about the boot up time your sandboxes as much or the wait sync time Or the time of any environment or your grading. It's fine if you have these things that take time because you can't If you're doing grading with a judge you can't force your judge to be super fast all the time And so you want to have a system where it's okay If there are pockets of your life cycle that don't use GPU time, but do use time And you can overlap these without wasting GPU cycles And so that's really the key benefit of a base in car L in our eyes And we've done a lot of work on the loss function side. The DPPO paper is one that I think has gotten popular that we've been using a lot as well of how do you make sure that this stuff stays stable? We found that it's very stable up to thousands of steps We're pushing towards 10,000 and kind of current experiments in terms of the scale that we're trying to get the stuff to reliably As well as doing this at big batch sizes where you need to pull out all the bells and whistles on parallelisms and so what we found is that I can go back to Doing expert parallel on the trainer as well as contact parallel And then on inference doing a wide extra parallel for the big MOEs meaning multi node experts across multiple nodes as well as desegregated pre-fill and you can throw all the inference bells and whistles that people would do for normal serving into your RL stack and get same wins there as well So here to fully recap all the advancements we've been pushing into the stack We've been leaning in towards FP8 Wide expert YDP Desegregate pre-fill Lots of stuff at the routing and KB offloading and management of just where things live Especially things like router replays. The router replay is actually a really nasty systems problem because it requires tracking a lot of metadata per rollout because you have this for every single layer So it's a big multiple over just tokens and log probs that you actually have to store Because especially if you're routing multiple experts per layer And so the storage concerns for these as well as for multimodal if you have images that you need to store, there's a lot of stuff where you want to offload the heavier artifacts onto some object storage system or other file system And not just have it floating around in memory And so we've rebuilt a lot of our systems to support this as well as just really pushing There's a lot of special cases that you need to worry about where certain models will want certain types of context parallelism which have different considerations about what you can do in terms of other aspects of the system And so we've just been trying to really refine the recipes for making sure that this works really well Especially for models like GLM And we've done this all on top of a torch titan base I think a lot of people ask why do you use torch titan and not megatron? And it's because torch titan is really easy to hack and megatron is this monolith That I think some people will hack it, but We just started with torch titan a long time ago and kind of find the pieces we want to bring in And especially when a new model comes out it's like well you want to train on it or you want to say you want to there's some new paper that you read that has some new idea You want to be able to make everything really hackable and modular and so our team is not that big. Our whole research team that maintains our RL is less than 10 people And the company is less than 40 people And so there's a lot of work that we want to make sure people can parallelize but also move quickly and so This is all stuff that we've figured out over the past few months as we've really been pushing it for scale Another thing that's fun is algorithms. I think a lot of what researchers care about. I think researchers, some people will love thinking about the system problems of scale If you do talk to us, but if you don't, I think what a lot of people want to spend their time thinking about is algorithms in terms of the on policy distillation stuff or OPSD or if you saw the echo paper I think that got a lot of people excited of thinking about how do you do world modeling with RL And folding these in as well as just basic stuff like SFT. There's also this max RL paper from a while back that was super cool And so we were seeing all these papers and we were we just want to do all of these. We want it to be much easier to mix and match these and not have to add another if statement buried super deep in the code And pipe a bunch of stuff all the way through and so we decomposed things into the loss which is the thing that is taking the gradient as well as the algorithm which we say is the thing that's preparing the data And so we have different losses we can pipe things to in terms of the signal and the masking And you can use these to assemble different algorithms And so now you go to just a class where you can have a class that does different things in terms of scoring or the groups And you pick which loss you want to target with it And so we support all of these, all the popular ones, and you can add your own that look like adding in a function to assign advantages to a rollout Yeah, and so then from your algorithm then from your configs you can just say hey I want this algorithm and it's just going to pick one from the registry. You can add your own to the registry if you want And then this will be the algorithm used for your training run And you can do this on a per environment basis if you want But all of these algorithms that people are looking at fall into this table where there's questions about where your rollouts are coming from. Are they coming from your current policy model? Or are they coming from some other source like a teacher? And so any algorithm people call on policy or slightly off policy in terms of the async RL stuff Yeah, and so then from your algorithm, from your configs you can just say hey, I want this algorithm and it's just going to pick one from the registry. You can add your own to the registry if you want. And then this will be the algorithm used for your training run. And you can do this on a per environment basis if you want. But all of these algorithms that people are looking at fall into this table where there's questions about where your rollouts are coming from. Are they coming from your current policy model? Or are they coming from some other source like a teacher? Any algorithm people call on policy or slightly off policy in terms of the async RL stuff, this is one where your actor in the RL sense is going to be the model you're training. Your policy is your actor. In other cases you're doing stuff where your actor is some other model. So if you're doing context distillation or you're doing SFT, these are ones where you are going to have some other model or prompt potentially be the teacher that is generating the data that you're going to be training on. And so all these fit within this family. Within this, the other thing is the advantage. The advantage is just, if you generalize it to a score, you can now call cross entropy loss or negative log likelihood as everything being advantage one. All of the OPD algorithms, you can think of the log prob ratio as being the advantage. And with RL, your reward minus some baseline, usually group mean or something like it, is your advantage. And so we just took a step back and looked at all these things and we're like, oh, we can just factor this all out pretty nicely and have almost everything else be shared. But these things get swapped. And so your infra doesn't need to change just because your loss function needs to change. And so for on policy distillation, it's just plugging in with a different loss target and it's talking to the teacher and being like, okay, the thing I need to get is reference log probs from some teacher and I already have my sequences. So I just need to send these to a teacher as pre-fill. You can get pre-fill by just asking for a one token response. Now I get my sequences and I can just stick these in as my reference log probs. For self distillation you can have a hint that you're putting in before your teacher. So you're not saying the same prompt, you're sending a different prompt using renderers to pack these together. And then you get back the same sequence that you can slice out into the original form. So that you can add these as your reference log probs there. You could do echo where you have two different algorithm components. You could have one that is targeting cross entropy on your environment tokens while doing an RL objective on your action tokens. And you can mix and match these. So you can have different teachers. You can decide whether you're going to be sampling from your student to your teacher. You can decide which algorithm you're going to be using on a per environment basis. And you can have this one where we have both a normal OPD as well as GRPO. And yes, so this whole family we now support within Prime RL natively. And one thing you may have also poked around or seen or tried is we have our hosted training platform. Which means you don't have to worry about GPUs at all. And so this is just hosted Prime RL. The version we have that is the broad self-serve version today is multi-tenant LoRA, which is focused mostly on RL. What we have coming quite soon that we'll be rolling out is full fine-tuning, which supports changing as much as you want in Prime RL in terms of the model and everything else, where we still give you all the same abstractions for not needing to think about the GPUs and auto scaling and magic restarts and having a dashboard where you can log everything and have unified billing for sandboxes and judges and all these things. But also you get to develop your environments on CPU on your laptop, push them to the platform as environment packages and specify them in your configs. And then this allows you to have a lot of flexibility in deciding when you want to go to different levels of the stack. In many cases you don't actually want to change anything in the trainer. You just want to change your reward function and you can do this in environment space. Maybe you want to configure the trainer but not change it, in which case you want to poke into slightly more complex knobs for things like different loss functions or learning rates. Additionally, you may want to actually go deeper into the trainer and do new algorithms at the environment level, the algorithm class level, or the loss function level, or maybe something else entirely that needs going even deeper. And so we support all of this with the full fine-tuning as well. And so multi-tenant LoRA, if you're unfamiliar, is a very useful pattern for allowing multiple people to do training runs on the same architecture, on the same model weight copy. You have a base model. This is how all inference for token-based pricing usually works. You're doing multi-tenant inference where you have one big copy of Claude that is just getting everyone's requests and hitting a shared KV pool that is managed. This is really nice with LoRA because you can just have multiple LoRAs available that you can hot swap and each person can have their own LoRA without needing to replace the base model. And you get one inference pool that serves everybody at once even if they're using different models. This allows you to do things like token-based pricing and not need to reserve GPUs. For full fine-tuning, it does have to be GPU-based. But we've been getting a lot more GPUs, and so we have the ability to let people run stuff on their GPUs and have more than just a LoRA adapter as well as the ability to really customize at every layer you might want, like how your algorithm is going to work. But so the multi-tenant LoRA one is already live and you can use it today. Full fine-tuning is coming out in the next couple weeks. The v1 stuff from before is already out as an alpha feature, with a stable release coming soon. And then the cookbook I mentioned earlier is also out. And that's mostly what I want to talk about. It's a bit informal and if people have questions, I can dig into any individual parts people think are curious. We're also hiring. We are a small team based mostly in San Francisco, becoming a much larger team quickly. And we'll have some other exciting news about that coming later this week. Full fine-tuning is coming out in the next couple of weeks. The v1 stuff from before is already out as an alpha feature, with a stable release coming soon hopefully. And the cookbook I mentioned earlier is also out. That's mostly what I want to talk about. It's informal and if people have questions, I can dig into any individual parts people think are curious. We're also hiring. We are a small team based mostly in San Francisco. We're becoming a much larger team quickly. We'll have some other exciting news about that coming later this week to demonstrate our commitment to scaling the team. We are growing and I think we're in a unique position where we do very open research work. All this code is on GitHub if you want to go play with it. But we're also a real company that trains big models and makes money. So I think we've figured out a good business model to do real open research. It's been an incredible journey thus far and I would love to have excited, passionate, talented people on the team. Thank you. throughout the ecosystem. So we integrate natively with Huggenvase datasets, with Harbor, with NemoGym, and OpenEnv, and most other tools that you see out in the wild that are kind of under the umbrella of an RL environment, we would call these a task set. We generally have found that it's useful to have these be harness agnostic, where they represent the backend or the server or some state that you're interacting with, but they don't own everything about what the model is doing. And so it doesn't, in some cases, it doesn't make sense to plug a model into a task set, especially because we're kind of gravitating towards an agent world where everything is running in a terminal, or it has skills, or it is using CLI tools. And these things, like, often look more complex than just basic loops. But we also want to support basic loops. So we want to kind of allow both the old way of doing things and the new way of doing things. And so everything that was the old way is now the default harness, where it's system prompt and tools in a loop. But the harness pattern also supports much more flexible execution of things like recursive language models, or CLI agents, like Codex, Cloud Code, OpenCode, or classics from the resource literature, like Mini Suite Agent, or building your own with arbitrary Python libraries like Langchain or DSPy. And so we've been able to decouple these into a pattern where you get to write your harness independently of your task set. There are kind of some basic sanity checks about properties that, like, harnesses either do or don't support and task sets do or don't require. And these kind of click together. And the runtime is where this executes. And so we've still been embracing a lot of the async IO patterns from before, but we've leaned a little more into having things be sub processes, where you can still run everything locally. You don't have to use sandboxes, but you can use local Docker, or you can use our own Prime sandboxes layer. You can use any other sandbox layer you'd like, or kind of build from scratch. And so the harness, the runtime backend, just is a place where the harness can run its code. And so the harness just needs to be able to run code somewhere as a script, essentially. We've used a lot of the UV tooling where UV script is a very powerful pattern to be able to kind of mix and match and kind of contain dependencies. But what happens is once you plug these together, you run a rollout on a task from a task set, and you get a trace. This is live on the verifiers main branch for prime intellect AI slash verifiers on GitHub, as well as it's released as a dev release. The stable main release will be kind of coming to PyPy any minute now, but you can install the dev and play around with it if you want. And so what do these look like? So tasks are just like a row of a data set. And the very basic version of it is you just are loading a data set from Hugging Face or anywhere else. And so the new pattern here is from verifiers v1, just to keep the old stuff separate. The old stuff still works just fine, but this is how we have been able to kind of decouple and iterate on the new version. As well as we've really embraced this decorator pattern. We found it to be very useful. But we also, if you were a rubric fan, we killed rubric. Didn't make sense anymore if you were using old verifiers rubric patterns. But still, you have functions and loaders. We are very heavy on Pydantic, so everything is super typed. We have lots of powerful config features where you can have everything in a toml file, you can override it in the CLI, and everything is kind of clean and guaranteed to kind of type check at a validation time rather than waiting for something to fail later down the road. And so examples of this are things like suite grep where you can do agentic code search, you can do the classic games like Wordle, you can do search over documents with judges, you can do complex things like Harbor that support a lot of popular benchmarks now that need agents running in a terminal. And all of these are going to be combinations of the the task set pattern with pick your own runtime and pick your own harness. And so rewards and metrics, I think, are also kind of just functions that take in the kind of records of what's happened in a rollout and return numbers. So rewards are the main thing that will drive progress in RL. Metrics are just kind of like logging what has happened, so counting tool use and counting errors. These sorts of things are very useful to be able to expose in your dashboards. And then group rewards. I think this is something that we have fought hard to kind of make sure still is first class because we see it as very important to a lot of the research pattern people want to do, but I think is also ignored in a lot of like tooling out there where in many RL frameworks it's actually quite hard to do group rewards because things are very decoupled and things kind of assume that all rollouts are going to live independently and that they don't need to talk to each other. But there's a lot of things where you really want to do pairwise judging or you want to do ranking or you want to give a bonus to the shortest correct answer in terms of tokens used. And so these sorts of things are really flexible in terms of the we've really designed for flexibility in supporting the things that we see as like the most exciting papers we've read or all the algorithms that we think people may want to innovate on while still allowing people to have like the core primitives that they kind of expect out of an RL framework. And so like in group rewards I think this concise this pattern is one that I find very useful a lot. I think like a big pattern that comes up a lot when people are doing post training is models will love to like think and think and think if you let them. And if you don't give them some kind of pressure to like be more efficient I think a lot of people will notice that like open models often have really really long chains of thought because on one hand it's like this is a useful strategy for a model but it's also the sort of thing that will grow like out of control if you don't counteract it. And so in reward design like one of the big things people will want to do is something like a length penalty or a conciseness bonus. And so one of the reasons this is tricky is because you don't know the optimal length for a problem. Like if I give you a math problem I could say oh solve it in less than n tokens but also like who knows what the right n is. It's also going to change as the model gets smarter over time. It's going to be different for every problem. And so you kind of can't know this up front and the only way you can do it is take advantage of variance. And so one of the nice things about RL is you have multiple samples typically and this allows you to use the fact that you have multiple samples to shape the reward. And so if you have multiple robots in a group what you could do is look at all the ones that were the correct answer or just all the ones in general and give a bonus to the ones that are the most concise. Where if you also have a correctness reward like these are going to ensure that you're both incentivizing correctness as well as incentivizing efficiency. And so juggling multiple objectives simultaneously is kind of one of the hard challenges in RL in reward design. But doing things like group level comparisons and kind of these sorts of bonuses are quite useful in many cases. I also want to talk about like tools and user simulators which I think have been becoming more important in a lot of complex applications where you have models that are in many cases there's like a core agent harness but there's also in many cases you are putting a model in a setting where it's going to be in some task where a user is giving it additional tools. Whether these are in some cases you might want to model these as skills in some cases you might want to model them as MCP servers. We use MCP as a kind of a back-end framework that can interact with the runtime both for tools and for user simulators. So user simulators especially if you want to do training where there's a user in the loop you don't want just your agent to go do some tasks but you want to be able to do some tasks that involves understanding how a user will interact. You can essentially have this user be an MCP as well where we make it so the model sees it as a user not as a tool but behind the scenes it is a server that has some script or something and it has some LLM that is going to get some context and it's going to be able to be like a user in the context of a rollout. In many cases benchmarks have found that this is very useful for simulating the realism of like having a multi-turn setting where there are users in the loop especially given that people are building products now where there are these users in the loop and so you want kind of a first class way to incorporate this into your URL environments. And so we've found that it's useful to have all of these things be kind of modular and pluggable and so the harness can connect to each of these which run as a as a UV script. We also have UV script support for grading in addition to the basic reward function patterns and we also are using this idea called an interception server. So the harness we want people to be able to use real harnesses and not have to like break the harness and like retrofit it into like an RL harness and so ideally we don't know anything about the harness code and so the pattern that we use with the interception server is that we are responsible for giving each harness rollout a fake base URL which could be open AI compatible or anthropic compatible. The harness just thinks it's talking to some endpoints so any harness that can just talk to some endpoint we're good to go and then we intercept each request. We can do some some back-end maneuvering to make sure that we're getting the log probs and setting the right temperature and then we send this to our inference server with the RL trainer and then as it completes we send back the request and so the the harness doesn't know that it's doing RL. The harness just is a harness running as if it would be running in a real world environment and so you can kind of very easily move between the RL setting and the deployment setting where your harness is just code. It doesn't need to be anything specialized to verifiers or RL. We also have found it really useful to be able to kind of go from this like local to global and like have this hot swap pattern so we have this new like eval CLI where you can just like choose the harness you want. You can have a task set where you're saying I'm going to run an eval on this set of tasks. Okay I want to use recursive language models and RLM. I want to use codex. I want to run it locally. I want to run it in send boxes. I want to run it in Docker. These are all just like interchangeable and so we found that this is super useful in the iteration loop as you go from testing something out on a small scale towards scaling it up towards being able to understand questions like what's the best harness for this model? Does this harness generalize across tasks? As well as being able to both have the convenience of like local prototyping where you can kind of run things fast on your MacBook without having to like wait for a cloud job to finish but also you can like go right to the cloud when you need to. And so there's been a lot of fun patterns we've had to kind of innovate on as we've done this overhaul and we were quite happy with how it's turned out. It's made our lives a lot easier for both client projects and research and just being able to have a lot more flexibility and power and control over the the kinds of agents we want to be training. And so one of the fun things behind the scenes is what we call the trace graph. And so we had kind of been having this grow out of control in terms of the old way of doing things and we decided this was another opportunity to like really overhaul our system to like have really good support for sub-agents and parallel branching trees while also still preserving the kind of linear sequential dependencies that you need for RL with careful token control. And so here there's a notion of branches that are kind of like at the message level. So conceptually the things that matter logically in environment space and in harness space are messages which are just text. The harnesses don't think about tokens but if you've done any RL experimentation you may have encountered issues where re-tokenization or like some messages if a model will say something and you turn it into text and you put it back through tokenizer it can change a little bit because tokenization is many to one. And so this causes lots of very subtle numerical problems especially late in large-scale training runs. And so you want a really nice back and forth between messages and tokens. And so the trace data structure that we created here partly is to enable this where we can store things both at trace level and then map them back into token level in the right sequences as needed. And we also released a library called renderers recently which is a standalone toolkit that anyone can use that we have found the sort of thing that we're working with some of the inference tooling to support. So renderers are really all about like essentially rethinking tokenizers and chat templates where behind the scenes it's just making calls to the tokenizer but chat templates if people have spent time debugging with them it sucks. Jinja is awful. It's very very painful and there's so many subtle things that we kept running into where like a model would sometimes have an extra new line and the chat template would strip it out and this would like cause a mismatch in your trainer and inference that would either force you to go off policy because you now have a trainer inference mismatch or it would cause a logical branch where a thing that is a branch in like it becomes a branch in token space even though it shouldn't be in logic space because of tokenizer subtleties. And so renderers as an abstraction it was kind of pioneered by OpenAI's Harmony with the GPT OSS release and used prominently in Thinking Machines cookbooks as well for Tinker but we found it was useful to just kind of make it a standalone thing and so this is just a python library that doesn't depend on any other prime stuff. You could use it with any inference engine you want just as a standalone thing that is really designed for being able to manage this token in token out concatenation without thinking about it too much yourself because we kind of turn each of these chat templates for these the popular models into programmable artifacts where you can do things like look up a history of secret you can use the kind of history of a trace to be able to understand like what is the right tokenization like do I essentially have like a logical prefix hit in message space even if I don't in tokenization space after retokenizing and so this is the sort of thing where I think people have gone back and forth on like whether they want LLM APIs to be stateful in general I think a lot of people were hoping that like we could just have every model API be stateless I think maybe people are less concerned about this now because we're moving towards this agent world where agents themselves are going to be stateful APIs But I think this has revealed to us like going through all the things here like why opening eye responses decided to be stateful There are some kind of like unavoidable issues that kind of come up when you're doing a large-scale agentic rollouts where you you do need to kind of manage this very carefully and it's kind of unavoidable just because of how tokenizers work And so you want to be able to maintain these dual streams of the logical text and the the tokens And you kind of want these to be cleanly interoperable where users don't have to think about the tokens very much But the trainer gets to see everything nicely in token space As well as the inference engine And so from the harness interception server we have clients that can be used both for training and inference and so you can kind of like Swap between these modes without thinking about it because Certain models like don't need to in in a training setting you need to be able to get log probs and set the temperature Some model APIs won't let you do this They won't return log probs because like open AI models with reasoning like won't show you the reasoning trace So there's no way they give you the log probs for everything and so like that's fine It's just eval only and so we have this client layer You can go between eval and train to be able to support all these models But we still use the interception server pattern either way because it allows us to have like This notion of a dialect where like you can choose Open AI chat completions or responses or anthropic and all of these are kind of easily supportable as just like Translation layers between a raw request into something that'll get passed through a renderer potentially if you're on the training client side And formatted into a message via tokens And so this brings us to prime rl so prime rl is our training framework that is consumes the environment So once you have an environment with your task set and your harness and your interception server and your runtime and your render and all those things This plugs into what we call the orchestrator And so prime rl has been async from the ground up So I think async rl is one of those things that I think people were kind of one foot in one foot out And a lot of training frameworks if you see them will still kind of support synchronous training Some people I think have their reasons for wanting to do synchronous training. I don't agree with them I think you kind of want to bite the bullet of the off policyness anyways for reasons that come up with agents In terms of you want to be able to overlap long rollouts and not always be waiting on your slowest rollout And this kind of means you can't be fully on policy unless you want to kind of accept always waiting on your slowest rollout And so this is really why we went all in on async and so the orchestrators job is to Allow the inference and trainer to just be separate processes separate servers They don't share gpus they don't really know about each other all that much they just consume from each other But the orchestrators job is to really like manage the run and so the orchestrator will Make sure that the environment is running with the endpoint mapping to the inference server It'll do rollouts It'll package these up into a batch It'll send this back to the trainer and it'll be up to the trainer to figure out what to do with the batch Which will be kind of printing some sequence to feed into a loss function Based on the the specification And so the the server pattern we use for environments is just an engine that can like send request the inference and like send batches back to the orchestrators It's very client server And we found that this is just a really useful way to Allow scaling concerns to be decoupled as well And so like for example you can have a lot of environments running or you can have one environment running you can have You can have a bunch of inference replicas you can have one in sepulchra you can have sandboxes or no sandboxes And the trainer doesn't care about this the inference doesn't care about this It's just separation of concerns at a system level Allows you to kind of not really worry about these things as Like combined units versus like in some cases People will want to like have training an inference on the same stack where it's like Especially the closer that you kind of fold in your logic all in one then it's like sometimes you can't even run your environments If you're not doing rl, but that means then you can't really experiment you also can't use them as evals There's a lot of reasons why just like pulling everything apart into these pieces and just like having nice api's for them to talk to each other Makes your life way easier And we've also just been like really scaling it And so we've been doing a lot of work on Like glm 5 series and kimmy k2.5 2.6 series just to make sure that we can do like really good large-scale rl efficiently And so some results we found we have recently this was The run was on glm 5 before the 5.2 one came out it supports 5.2 as well But on the latest primaral version we can do A glm 5 step on 28 nodes in less than five minutes for long horizon coding tasks with 131k context Which means you can do a thousand step run in three days and that costs For rental prices about 50k and so 50k is not cheap But it's like if you're doing a full run on a frontier sized model On like a proper real world agent environment Like it's a lot cheaper than what open.ai is raising for it It's a lot cheaper than like some of the clusters people are Selling and it's the sort of thing that like just starts making sense for a lot more enterprises if you can like actually do this It becomes pretty justifiable if you can kind of get to the point where You have the tooling chain to be able to like build the ready valves and like do all this stuff And so this is like the sort of thing where it's like let's say you do wanted to like you have a bunch of tasks that are like Representative of like your coding workflows and you want to like do a big rl run like this is a pretty big rl run But it's also the sort of thing that like you could people spend this much on tokens in a month sometimes And so you can start finding a lot of savings if you kind of think about doing large-scale post training And that's kind of what we're here to help people do I guess more on the async side One of the reasons why you really want to do async is that There's a long tail of how long your coding agents take like if you fire up a coding agent task Just like think your quad code or codex tasks like how many minutes is it going to take you? Sometimes it'll be 30 seconds sometimes it'll be like two minutes sometimes it'll be a goal that goes for like Three hours And these can all be rollouts and so one of the goals of async rl is to have your like forward progress speed Not be tied to the speed of your individual rollout And so this means you can kind of allow your rollouts to finish Long after they start and just go into the first batch that they can accept them Even if that is from a much different copy of the model And so then the the inference server is just always taking the latest version of the model and so What people have generally found and what we've done with our experimentation and found as well is that you can go regionally far off policy Like I think 16 is where we typically are often operating is like an average But this means that you just also can have a lot more room To not worry about certain things about speed Like you don't need to worry about the boot up time your sandboxes as much or the wait sync time Or like the time of any environment or your grading like it's it's fine if you have these things that take time because you can't If you're doing like grading with a judge you can't like force your judge to be super fast all the time And so you want to have a system where it's okay If there are pockets of your life cycle that don't use gpu time, but do use time And you can kind of overlap these without kind of wasting gpu cycles And so that's really the the key benefit of a base in car l in our eyes And we've done a lot of work on the loss function side the dppo paper is one that I think has gotten Popular that we've been using a lot as well of like how do you make sure that this like stuff stays stable? We found that it's very stable up to thousands of steps We're pushing towards 10,000 and kind of current experiments in terms of the the scale that we're trying to kind of get the stuff to reliably As well as doing this at big batch sizes where you kind of need to pull out all the bells and whistles on parallelisms and so what we found is that I guess I can go back to Like doing like this expert parallel on the trainer as well as contact parallel And then on inference doing a wide extra parallel for the big moes meaning multi node experts across multiple nodes as well as Desegregated pre-fill and you can kind of throw all the inference bells and whistles that people would do for normal serving Into your rl stack and get same wins there as well So here to kind of fully kind of recap all the advancements we've been pushing into the stack We've been leaning in towards FP8 wide expert YDP Desegregate pre-fill Lots of stuff at the routing and kb offloading and management of just kind of where things live Especially things like router replays the router replay is actually a really nasty systems problem because it requires Tracking a lot of metadata per rollout because you have this for every single layer So it's like it's a big multiple over just to be like tokens and log probs That you actually have to store Because especially if you're routing multiple experts per layer And so the storage concerns for these as well as like for multimodal if you have like Images that you need to store like there's a lot of stuff where you want to kind of offload the heavier artifacts onto some like object storage system Or other file system And not just have it floating around in memory And so we've rebuilt a lot of our systems to support this as well as just really pushing There's a lot of kind of like Special cases that you need to worry about where it's like certain models will want certain types of context parallelism Which have different considerations about what you can do in terms of like other aspects of the system And so we've just been trying to really refine the recipes for making sure that this works really well Especially for models like GLM And we've done this all on top of a torch titan base I think a lot of people are like why do you use torch titan and not megatron? And it's because torch titan just like really easy to like hack and megatron is kind of this monolith That I think some people will hack it, but it's I think We just started with torch titan a long time ago and kind of find the pieces we want to bring in and Especially when a new model comes out It's like well you want to train on it or you want to let's say you want to there's some new paper that you read it like has some new idea You want to be able to like make everything really hackable and modular and so like our our team is not that big like our whole research team that Maintains from our L is like less than 10 people And the company is less than 40 people And so there's a lot of work that we want to make sure people can paralyze but also move quickly and so This is all stuff that we've kind of been able to figure out over the past few months as we've really been pushing it for scale Another thing that's I think fun is algorithms, so I think a lot of what researchers care about I think researchers some people will love thinking about the system problems of scale If you do talk to us, but if you don't I think what a lot of people want to spend their time thinking about is algorithms in terms of The on policy distillation stuff or like opsd or if you saw the echo paper I think that got a lot of people excited of like thinking about how do you do world modeling with rl And folding these in as well as just basic stuff like sft. There's also this max rl paper from A while back that was super cool And so we were just we were seeing all these papers and we were like we just want to do all of these we want it to be Much easier to kind of like mix and match these and did not have to like add another if statement buried super deep in the code And like pipe a bunch of stuff all the way through and so we kind of decompose things into the loss Which is like the thing that is taking the gradient as well as the algorithm which we say is the thing that's kind of like preparing the data And so we have different losses we can kind of like pipe things to in terms of like the signal and the masking And you can use these to kind of assemble different algorithms And so now go to just like a class where you have you can have a class that does different things in terms of scoring or the groups And you kind of pick which loss you want to target with it And so we support all of these all the popular ones and you can kind of add your own that look like adding in a function to like assign advantages to a rollout So Yeah, and so then then from your algorithm then from your configs you can just kind of say hey I want like this algorithm and it's just going to pick one from the registry you can add your own to the registry if you want And then this will be the algorithm used for your training run And you can do this on a per environment basis if you want But all of these algorithms that people are looking at kind of fall into this like table where there's questions about like What are your where your rollouts coming from are they coming from your current policy model? Or are they coming from some other source like a teacher? And so Any algorithm people call like on policy or like slightly off policy in terms of the async RL stuff This is one where like your your actor in RL sense is going to be the model you're training your policy your policy is your actor In other cases you're doing stuff where your your actor is some other model So if you're doing context distillation or you're doing sft like these are ones where you are You're going to have some other like model or Prompt potentially be the teacher that is generating the data that you're going to be training on And so all these kind of fit within this family within the other thing is the advantage and so the advantage is just like if you generalize it to like a score Like you can now call like just cross entropy loss or negative log likelihood is like everything is like advantage one There's a lot of all of the like opd algorithms can kind of you can kind of think of it as the log prob ratio as being the advantage And with RL like your Your reward minus some baseline usually group mean or something like it is your advantage And so we just kind of like took a step back and looked at all these things and we're like Oh, we can just kind of like factor this all out pretty nicely And have almost everything else be shared But these things get kind of swapped and so your infra doesn't need to change just because your loss function needs to change um And so for on policy distillation like it's just plugging in with a different loss target and it's like talking to the teacher and being like Okay, the thing I need to get is reference log probs from some teacher and I already have my like sequences So I just need to send these to a teacher as pre-fill in terms of like a like you can get pre-fill by just like asking for like a one token response Now I get my sequences and I can just like stick these in as my reference log probs Um for self distillation you can have like a hint that you're kind of putting in before your um your uh Like this teacher so you're not saying the same prompt you're sending a different prompt Using renderers to kind of pack these together And then you get back the same like sequence that you can slice out into the original form So that you can kind of add these as your reference log probs there You could do echo where you have like two different algorithm components you could have one that is targeting the uh the doing cross entropy on your environment tokens while doing like an rl objective on your your action tokens um And you can mix and match these so you can have um like different uh like you can decide which teacher is going to be done You're whether you're going to be sampling from your student to your teacher You can like decide which algorithm you're going to be using on a per environment basis And you can have like this one we have like both a normal like opd as well as grpo Um and then Ah yes, so like this whole family we now support within prime rl natively um And one thing you may also have poked around that or seen or tried is uh we have our hosted training platform Which means you don't have to worry about gpus at all And so this is just hosted prime rl the version we have that is kind of the the broad uh self-serve version today Is multi-tenant laura which is focused mostly on rl Um what we have coming quite soon that we'll be rolling out is full fine tuning Which supports changing as much as you want in prime rl in terms of the model and everything else where we still give you all the same abstractions for uh Not needing to think about the gpus and kind of auto scaling and uh Magic restarts and having a dashboard where you can log everything and have unified billing for sandboxes and judges and all these things But also you get to develop your environments on cpu on your laptop push them to the platform as environment packages and specify them in your configs and then this allows you to um have a lot of flexibility in like Deciding when you want to go to different levels of the stack so in many cases you don't actually want to change anything in the trainer You just want to change your reward function and you can do this in environment space Maybe you want to configure the trainer but not change it in which case you want to kind of poke into like slightly more complex knobs for Things like different loss functions or learning rates Additionally, you may want to like actually go deeper into the trainer and like Do new algorithms at the environment level the algorithm class level or the loss function level or maybe something else entirely that needs going even deeper And so we we support all of this with the the the full fine tuning as well And so multi-tenant laura if you're unfamiliar is a A very useful pattern for allowing multiple people to do training runs on the same architecture on the same model weight copy and so you have a base model That's kind of this is how all inference for like token based pricing usually works as you're doing multi tenant inference where you have one big copy of Claude That is just getting everyone's requests and hitting a shared like kv pool that is managed With this is really nice with laura because you can just have multiple lauras available that you can hot swap and each person can have their own laura Without needing to replace the base model and you get one kind of inference pool that serves everybody at once even if they're using different models This allows you to do things like token based pricing and not need to reserve gpus For full fine tuning it kind of does have to be gpu based But we've been getting a lot more gpus and so we have the ability to let people run stuff on their gpus and Kind of have more than just a laura adapter as well as the ability to kind of really customize at at every layer you might want Like how your algorithm is going to work But so we the multi tenant laura one is already live and you can use it today Full fine-tuning is coming out in the next coming next couple weeks The v1 stuff from before is already out as a kind of a alpha feature that is a stable release coming the next I don't know quite soon hopefully and then the cookbook I mentioned earlier is also out And that's mostly what I want to talk about It's a bit informal and I can people have questions. I can dig into any individual parts people think are curious We're also we're hiring we are small team based in mostly San Francisco Becoming a much larger team quickly and We'll have some other exciting news about that coming later this week to I don't know Demonstrate our commitment to scaling the team But yeah, we are we're growing and we're I think in a unique position of like we do very open research work like all this code is on github if you want to go play with it But we also like our a real company that trains big models and makes money And so yeah, I think we figured out a good business model to do real open research And it's been an incredible journey thus far and would love to have Excited passionate talented people On the team Thank you