The Frontier AI Inference Cloud for Agents — Byung-Gon (Gon) Chun, FriendliAI
Description
Byung-Gon Chun's team invented continuous batching, now standard across the industry, and the work that followed inspired one of the most widely used open source serving frameworks. So when he says agents have changed the economics of inference, he knows the tooling from the inside. His demonstration: the same coding agent, the same task of building a tower defense game, run once on a closed frontier model and once on an open weight model. Both finished at a usable level. The open weight run came in roughly five and a half times cheaper. That is the promise of open weights. But the model is only part of the bill, because agentic inference is a different problem from chat. In chat the unit was the request. In an agent the unit is the task: a loop of plan, act, observe, repeat, often running for minutes or hours, with sub agents fanning out in parallel and every observation appended to a context that only grows. His internal traces show consecutive steps sharing enormous prefixes, and recomputing that prefix on every call is compute spent on work already done. Nobody cares about the latency of one call; they care when the task is done. FriendliAI rebuilt its stack around that metric. Prefix caching so a shared prefix is computed once. A hierarchical KV cache across GPU, host memory and disk. Cache aware routing that sends a request to the replica already holding its prefix, instead of spreading load evenly and destroying locality. And agent aware scheduling that knows a call belongs to a longer program. One customer's split test found it seven times faster with a lower error rate. Speaker info: - https://www.linkedin.com/in/byung-gon-chun - https://bgchun.github.io Timestamps: 0:00 - The team that invented continuous batching 2:06 - One task, two models, one bill 3:18 - The unit is the task, not the request 5:25 - Consecutive steps share a huge prefix 7:02 - Four pillars of an agentic inference cloud 9:19 - Cache aware routing versus a naive load balancer 10:04 - A
Summary
Generated by gpt-5.6-terraAt-a-Glance
- Verdict: Skim
- Core thesis: Agent workloads require an inference stack optimized for end-to-end task completion—not isolated request latency—and FriendliAI argues that open-weight frontier models plus cache-centric serving can deliver that performance at materially lower cost.
- Why it matters: The talk identifies reusable infrastructure concerns for production agents: growing shared context, KV-cache locality, tool-call gaps, distributed routing, and scheduling decisions made at the agent-session level.
- Best use: Use it as a concise vendor-led framing of agentic inference architecture and a checklist for evaluating inference providers, rather than as independently validated performance research.
Executive Summary
FriendliAI CEO Byung-Gon Chun argues that production agents have changed the optimization target for inference. Chat systems are primarily measured per request, whereas an agent executes a task through repeated LLM calls, tool calls, accumulating context, and sometimes parallel sub-agents. The meaningful user metric is therefore end-to-end task completion time, which is variable and cannot be managed as a predictable stream of independent requests.
The central technical observation is that adjacent steps in an agent session commonly reuse a large prompt prefix. Recomputing that prefix repeatedly wastes prefill compute, so FriendliAI’s proposed stack centers on prefix/KV caching, memory-efficient and hierarchical KV-cache management, and routing requests toward replicas that already hold the relevant cache. It also proposes agent-aware scheduling that can make decisions using the likely trajectory of the larger program rather than treating every model call independently.
The commercial argument is that frontier-capable open-weight models have become good enough for many agent workflows while being much cheaper than closed alternatives. The speaker illustrates this with a tower-defense coding task said to cost about $0.27 on GLM 5.2 via FriendliAI versus $1.50 on Claude Opus 4.8, while claiming both produced usable outcomes. FriendliAI positions its infrastructure as the layer that makes those models sufficiently fast and reliable for production.
The talk is strongest as an architecture thesis, especially around cache locality and task-level optimization. Its benchmark and customer-performance claims—including a claimed sevenfold Cursor speed advantage—are vendor assertions without workload definitions, latency distributions, configuration details, or independent validation, so they should inform diligence rather than settle a provider decision.
Key Takeaways
- Claim: Agentic inference should be optimized around end-to-end task latency rather than latency for an individual model request or token. | Evidence: The speaker describes an agent task as a recurring plan → tool/action → observation → LLM loop, potentially involving tens or hundreds of inference steps over minutes or hours; sub-agents may run concurrently and the number of calls varies by task. | Implication: Ken should evaluate inference infrastructure with task-level benchmarks that include tool waits, retries, context growth, and completion quality—not only TTFT, throughput, or per-call p99 latency. | Caveat: Task-completion time also depends on tool execution, environment reliability, orchestration design, and model behavior, not solely on the inference provider.
- Claim: Prefix reuse is a major efficiency opportunity in long-running agents because successive calls generally share most of their accumulated context. | Evidence: The speaker says internal GLM 5.2 coding-agent prompts grow as observations are appended, while consecutive steps retain a large shared prefix; caching the prefix KV state lets later calls process only the new suffix. | Implication: Agent runtimes should preserve stable session/context identities and avoid orchestration choices that unnecessarily invalidate or fragment reusable prefixes. | Caveat: The benefit depends on workflow continuity and cache retention: highly branching, short-lived, or non-sticky sessions may have less reusable prefix state.
- Claim: KV-cache management is a first-class capacity and performance problem for agent workloads, not an implementation detail. | Evidence: FriendliAI lists frugal GPU-memory packing, KV quantization, hierarchical caching across GPU, host memory, and disk, and distributed caching across replicas as required to keep large, active contexts usable. | Implication: When comparing serving stacks, Ken should request cache-hit rates, cache-tier behavior, eviction policies, context-length performance, and effective concurrency under long-running agent sessions. | Caveat: Moving cache state below GPU memory can expand capacity but may introduce latency tradeoffs; the transcript supplies no measurements for the tradeoff by cache tier.
- Claim: Global load balancing that ignores cache locality can make agent inference slower and more expensive even if it appears evenly balanced. | Evidence: The speaker contrasts naive load balancing across GPU clusters with cache-aware routing that sends a later request from task A to a pod already holding task A’s prefix, while still preventing a single pod from becoming a hotspot. | Implication: The control plane for an agent platform should treat routing as a joint cache-locality, load, and resilience optimization problem rather than a generic round-robin or least-loaded dispatch decision. | Caveat: Cache-affinity routing must be balanced against load concentration and failure recovery; strict stickiness can hurt tail latency when the preferred replica is overloaded or unavailable.
- Claim: Scheduling can improve task completion if the serving layer understands an LLM call as part of an agent program rather than as an independent request. | Evidence: FriendliAI proposes agent-aware optimization such as preempting the right work, speculatively prefilling likely next-step context, and making cache-eviction choices based on agent-level context. | Implication: There is a meaningful architecture boundary between an agent orchestrator and inference service: exposing session lineage, likely next actions, priority, and cancellation signals can enable better serving decisions. | Caveat: These are presented as opportunities and product direction; the transcript does not provide an implemented policy, prediction accuracy, or quantified gains from agent-aware scheduling.
- Claim: Open-weight models can make capable agent workflows materially more economical, provided the serving stack preserves production-level speed and reliability. | Evidence: For a tower-defense coding task, the talk claims GLM 5.2 on FriendliAI cost approximately $0.27 versus $1.50 for Claude Opus 4.8, or 5.6 times less, while both results were described as clearly usable; it also names GLM 5.2, Minimax, and Kimi as frontier open-weight options. | Implication: Ken should consider a model-routing portfolio in which open-weight models handle validated task classes, but use outcome-based evaluations and fallback policies rather than assuming price parity implies capability parity. | Caveat: A single coding-task comparison does not establish equivalent quality, reliability, safety, tool-use behavior, or total cost across a representative production workload.
- Claim: FriendliAI claims a production performance advantage from its agent-centric cloud design and cites Cursor as validation. | Evidence: The speaker says Cursor tested several providers for GLM 5 and found FriendliAI "consistently seven times faster" with a significantly lower error rate than other third-party providers and direct model-lab usage; FriendliAI also cites LG as a customer. | Implication: If FriendliAI is under consideration, ask for a controlled pilot on Ken’s own representative tasks and require task-level p50/p95 completion time, success rate, error taxonomy, cost, and failover results. | Caveat: The transcript provides no benchmark protocol, model version/configuration, traffic mix, definition of speed or error rate, statistical range, or independent evidence; this should be treated as marketing evidence.
Detailed Brief
Deployment and procurement posture
- Claims: FriendliAI presents three consumption models: a serverless model API, isolated dedicated endpoints with guaranteed SLAs, and BYOG deployment on customer-owned GPUs.; The company positions the same inference stack as usable by AI-native startups and large enterprises with different isolation and infrastructure requirements.
- Evidence: The serverless option is described as the fastest route to core frontier open-weight models.; Dedicated endpoints are positioned for production workloads requiring isolated deployment and guaranteed SLAs.; BYOG is described as running FriendliAI inference on an organization’s own infrastructure.
- Caveats: No details are provided on security controls, tenant isolation mechanics, regional availability, compliance posture, data retention, pricing, SLA terms, or operational burden of BYOG.; The presentation does not explain whether feature parity—including distributed caching and agent-aware scheduling—is identical across the three deployment modes.
- Implications: A procurement decision should distinguish experimentation needs from production requirements for data residency, isolation, capacity guarantees, and operational ownership.; Ken should verify control-plane integration requirements before assuming that a serverless-to-dedicated-to-BYOG migration will be operationally seamless.
What the talk omits from the task-latency model
- Claims: The presentation focuses on inference-side improvements but implicitly recognizes that agents spend meaningful time outside the model during tool execution.; Long-horizon workflows can run for minutes or hours, making reliability and recovery as important as raw serving speed.
- Evidence: The described agent loop alternates between LLM inference and one or more non-LLM tool executions.; The deep-research example is said to include multiple stages, sub-agents, repeated inference and tool calls, and persistent shared context.
- Caveats: The transcript does not cover durable session state, replay after failure, tool idempotency, distributed tracing, cancellation, human approval gates, or security boundaries around tool access.; It also does not quantify how much task latency comes from model execution versus external tools, so the economic and speed benefits may vary substantially by workload.
- Implications: Inference optimization will have the largest payoff in model-heavy loops with substantial context reuse; tool-bound workflows require equal attention to orchestration and external-system latency.; For long-horizon agents, an infrastructure evaluation should include recovery behavior and observability across the full task, not only live serving performance.
Notable Concepts & Terms
- Agentic inference: Inference serving designed for multi-step autonomous tasks with iterative model calls, tool use, growing context, and possible sub-agent parallelism.
- End-to-end task latency: The user-relevant completion metric for agents: elapsed time until a whole task finishes, rather than latency for one API request.
- Prefix caching: Reusing the previously computed KV state for the shared portion of consecutive prompts so only newly added context must be prefetched.
- KV cache: The model attention state retained from prior tokens; it is central to efficient continuation of large, repeated agent contexts but consumes substantial memory.
- Hierarchical KV caching: Managing cached state across GPU memory, host memory, and disk to exceed GPU-memory limits, with potential tier-dependent latency tradeoffs.
- Cache-aware routing: Sending a request to infrastructure that already contains its relevant prefix cache, rather than balancing traffic without regard to locality.
- Agent-aware optimization: Scheduling, prefetch, preemption, and eviction decisions informed by an agent session’s broader program state and likely next steps.
- BYOG: Bring Your Own GPU: deploying the provider’s inference software on customer-controlled hardware rather than consuming only a managed endpoint.
Operator Notes / Why Ken Should Care
- Add task-level benchmark instrumentation to agent evaluations: task success, end-to-end p50/p95 completion time, model/tool time split, token cost, retries, cache-hit rate, and recovery outcomes.
- Require session-affinity and cache-locality behavior from any inference control plane handling iterative agents; test its behavior under overload, replica loss, and branching workflows.
- Run a representative open-weight versus closed-model routing experiment before committing to a cost thesis, with quality thresholds and automatic fallback paths defined per task class.
- If evaluating FriendliAI, ask for a reproducible pilot that defines the workload, model parameters, context lengths, concurrency, error categories, and exact basis for its claimed speed advantage.
- For long-running agent systems, ensure durable orchestration state and idempotent tool execution are designed independently of the inference provider’s caching and scheduling capabilities.
Source/Metadata
- Title: The Frontier AI Inference Cloud for Agents — Byung-Gon (Gon) Chun, FriendliAI
- Transcript words: 1824
- Duration seconds: 898
- Timestamp note: No timestamps or chapters were present in the supplied transcript.
Transcript
# Transcript Let's get started. Hi, everyone. Thank you for coming. This is the late afternoon in the last day. So I really appreciate it. I'm Gaon, founder and CEO of Friendly AI. Today I want to talk about agentic inference. So I'll first walk through what changed, why it matters, and how we rebuilt the inference cloud for agents. Before we go deeper, let me briefly introduce Friendly AI. Friendly AI is the frontier AI inference cloud for agents. So we run inference for agents at scale, faster, cheaper, and more reliably. We are born from a research team at Seoul National University, and those research roots still define us. We are the team that invented continuous batching, the inference optimization that is now standard across the industry, and our ORCA work inspired 3LLM, a widely used open source framework. Today we operate globally, headquartered in San Francisco with a team in Seoul to scale frontier inference. As you know, 2026 is the year agents go into massive production, and it's driven by two trends coming together. First, agents are going exponential. AI agents are driving explosive adoption across software operations and knowledge work. Second, open-weight models have reached the frontier and make agents economical. They now rival closed frontier models in capability, which means you can run frontier quality agents on open models with much lower token costs. Let me make the open-weight model part concrete. Open-weight models are now strong enough for these types of real agentic workflows. Here we gave the exact same task, building a tower defense game with a coding agent to two models. On the left is GLM 5.2, an open-weight model running on Friendly AI. On the right is Anthropic Claude Opus 4.8. The important point is not that the outputs are identical. The point is that both complete the task at a level that is clearly usable. For many agentic workflows, open-weight models have crossed the quality threshold. But the economics are very different. For the same task, Opus 4.8 costs about $1.50. For the same task, GLM 5.2 on Friendly AI costs $0.27. That's 5.6 times cheaper. So this is the promise I mentioned earlier. Open-weight models give you frontier quality agents at a fraction of cost. But model cost is only one part of the story. To make agents actually fast and reliable, the inference stack itself has to change. So let's look at what actually happens inside an agentic workload. So first, let's look at changes in the workload. In the past, the dominant usage was chat. The basic unit was a request. A person asks a question, the model answers, and the person reads it. Latency meant how fast did I get one response. Agents are different. The basic unit is a task. A task may involve many model calls, many tool calls, and it may run autonomously for a while. So the user does not really care about the latency of one individual request. The user cares about when the whole task is completed. That means we have to optimize for tasks, not just individual requests. Let's look at agentic workload more closely. An agent really runs a session made up of tasks. Each task typically runs in a loop. First, it plans, which usually means an LLM call. Then it acts maybe by calling a tool. Then it absorbs the result and adds that back into the context. And it repeats this until the task is done. So we are constantly alternating between LLM inference and one or more non-LLM tool executions. So there is a gap between LLM calls. An agent can also create sub-agents and run them in parallel. Agent inputs also look very different from chat. The graph here shows the prompt and completion length distributions of our internal coding agent runs with GLM 5.2, which we use day to day. They are much longer. They grow as the task progresses, since every observation gets appended back into the context. There is an important pattern here. Consecutive agent steps usually share a huge prefix. If we recompute that same prefix every time, we are burning a lot of compute on work we already did. So this is one of the biggest opportunities in agentic inference. So how token hungry are agents? Now let's look at a long horizon test example like deep research. We ran the spec decoding framework in LLM using code with GLM 5.2 on Friendly AI. There are multiple stages and each stage is composed of sub-agents which run multiple inferences and tool calls. So it might run tens or even hundreds of inference steps, sometimes over minutes or hours. And the shared context keeps going the whole time. For the user, what matters is not the latency of a single token or one call. What matters is when is my task completed. So agentic inference is not just chat with more requests. It's a different problem. The context grows over time. Tool work is interleaved between model calls. The number of model calls depends on the input. So you can't really plan around a fixed request rate plan. And the real metric is end-to-end task latency, not a single request latency. This is where Friendly AI comes in. We rebuilt the frontier inference cloud specifically for agentic workloads around the challenges I just walked through. And we set one goal: optimize end-to-end task latency, the task, not just the request. So how do we do that? Let me show you the key engineering behind it. Here's the engineering map for how we think about it. We built the stack layer by layer around agentic workloads. There are four big pillars I'm going to cover today. Prefix caching, key value, also called KV cache, management, cache-aware routing, and agent-aware optimization. And of course, underneath, we need model layer optimization like sparse attention for long contexts, techniques to reduce errors, fast decoding, resilient serving, and more. In this talk, I'm going to focus on the four pillars. Let's start with prefix caching. Since agent steps share a large prefix, we compute key value for the prefix once and cache it. Then on later steps, we reuse the cached key value and only process the new suffix. Reading from cache is much cheaper than recomputing prefill, so this improves time-to-first token and reduces compute on every step. And the longer the task runs in agents, the more valuable this becomes. But caching only works if the KV cache actually fits and can move around efficiently. So we need strong KV cache management. We use frugal memory management to pack more active contexts onto each GPU memory. We use KV quantization to reduce the memory footprint. We use hierarchical caching across GPU memory, host memory, and disk, so we can go beyond GPU limits. And we also use distributed caching, so one prefix can be served across replicas, not just inside one instance. At global cluster scale, routing becomes really important. A naive load balancer may spread requests evenly across GPU clusters, but it can destroy cache locality. A cache-aware router at a global scale does something smarter. It sends a request to a pod that already has the right prefix cached, turning a core prefill into one cache hit. At the same time, it still has to balance load, so one pod doesn't become a hotspot. In this example, the two requests of task A go to the same pod for cache locality. The next piece is agent-aware optimization, and this is the next frontier of agent inference. Today, most systems schedule each LLM call as if it were independent. They don't really understand that this call is part of a longer agent program. But if the optimizer knows the agent-level context, it can make better decisions. For example, preempting the right work, speculatively prefilling context for a likely next step, or making a better cache eviction decision based on agent-level context. So the goal is to reduce end-to-end task latency, not just make one call look fast. When we put all of this together, this is the payoff. We are using the same model, GLM 5.2, with code to create a simple mobile game. We ran the same task with model APIs of Friendly AI and another well-known inference provider. As you can see, Friendly AI completes the same task end-to-end much faster, thanks to our agent-centric cloud design. So what does this unlock in practice? A stronger production agent stack. Take an agent you already like. Now plug in open-weight frontier models like GLM 5.2, Minimax, and Kimi served on Friendly AI. The model gives you frontier quality capability and better economics. Friendly AI gives you the speed, reliability, and end-to-end task performance needed in production. That combination—quality, speed, reliability, and cost— is what makes agents actually useful and economical in production. Friendly AI is currently powering teams in production from AI-native startups to global enterprises. I'd like to highlight a couple here. Cursor is a hugely popular agentic AI coding tool serving millions of users. LG is a global enterprise whose businesses range from electronics to healthcare to energy. Very different companies, but they all need the same thing: fast, reliable, cost-effective agentic inference. This testimonial from our client, Cursor, says it all. Over the past year, Cursor has tested several inference providers hosting both open and closed models. In a split test of GLM 5 usage compared against other third-party providers and direct usage from the model lab, Friendly AI was consistently seven times faster with a significantly lower error rate. Today, Friendly AI is a core component of the Cursor stack. And you can consume this however fits your stack. Model API is the fastest way to start. Core frontier open-weight models through our serverless API. Dedicated endpoints give you your own isolated deployment with guaranteed SLAs for production workloads. And BYOG, bring your own GPU, lets you run Friendly inference on your own infrastructure. Same stack, three ways to deploy. To wrap up, there are three things to remember. First, frontier open-weight models make production agents economically scalable. Second, agents are not just chat with more calls. Agentic inference requires optimizing end-to-end task latency with the challenges I mentioned. Third, Friendly AI is built as an inference cloud for this world. Fast, reliable, cost-effective agentic inference. Thank you for attending my session. If you're building agents, give frontier open-weight models a try on Friendly AI today. You can get started at Friendly AI in minutes. Thank you. I'll be around after the session. Thank you. Thank you. Thank you.