# Transcript Let's get started. Hi, everyone. Thank you for coming. This is the late afternoon in the last day. So I really appreciate it. I'm Gaon, founder and CEO of Friendly AI. Today I want to talk about agentic inference. So I'll first walk through what changed, why it matters, and how we rebuilt the inference cloud for agents. Before we go deeper, let me briefly introduce Friendly AI. Friendly AI is the frontier AI inference cloud for agents. So we run inference for agents at scale, faster, cheaper, and more reliably. We are born from a research team at Seoul National University, and those research roots still define us.
We are the team that invented continuous batching, the inference optimization that is now standard across the industry, and our ORCA work inspired 3LLM, a widely used open source framework. Today we operate globally, headquartered in San Francisco with a team in Seoul to scale frontier inference. As you know, 2026 is the year agents go into massive production, and it's driven by two trends coming together. First, agents are going exponential. AI agents are driving explosive adoption across software operations and knowledge work. Second, open-weight models have reached the frontier and make agents economical.
They now rival closed frontier models in capability, which means you can run frontier quality agents on open models with much lower token costs. Let me make the open-weight model part concrete. Open-weight models are now strong enough for these types of real agentic workflows. Here we gave the exact same task, building a tower defense game with a coding agent to two models. On the left is GLM 5.2, an open-weight model running on Friendly AI. On the right is Anthropic Claude Opus 4.8. The important point is not that the outputs are identical. The point is that both complete the task at a level that is clearly usable.
For many agentic workflows, open-weight models have crossed the quality threshold. But the economics are very different. For the same task, Opus 4.8 costs about $1.50. For the same task, GLM 5.2 on Friendly AI costs $0.27. That's 5.6 times cheaper. So this is the promise I mentioned earlier. Open-weight models give you frontier quality agents at a fraction of cost. But model cost is only one part of the story. To make agents actually fast and reliable, the inference stack itself has to change. So let's look at what actually happens inside an agentic workload. So first, let's look at changes in the workload. In the past, the dominant usage was chat.
The basic unit was a request. A person asks a question, the model answers, and the person reads it. Latency meant how fast did I get one response. Agents are different. The basic unit is a task. A task may involve many model calls, many tool calls, and it may run autonomously for a while. So the user does not really care about the latency of one individual request. The user cares about when the whole task is completed. That means we have to optimize for tasks, not just individual requests. Let's look at agentic workload more closely. An agent really runs a session made up of tasks. Each task typically runs in a loop. First, it plans, which usually means an LLM call.
Then it acts maybe by calling a tool. Then it absorbs the result and adds that back into the context. And it repeats this until the task is done. So we are constantly alternating between LLM inference and one or more non-LLM tool executions. So there is a gap between LLM calls. An agent can also create sub-agents and run them in parallel. Agent inputs also look very different from chat. The graph here shows the prompt and completion length distributions of our internal coding agent runs with GLM 5.2, which we use day to day. They are much longer. They grow as the task progresses, since every observation gets appended back into the context.
There is an important pattern here. Consecutive agent steps usually share a huge prefix. If we recompute that same prefix every time, we are burning a lot of compute on work we already did. So this is one of the biggest opportunities in agentic inference. So how token hungry are agents? Now let's look at a long horizon test example like deep research. We ran the spec decoding framework in LLM using code with GLM 5.2 on Friendly AI. There are multiple stages and each stage is composed of sub-agents which run multiple inferences and tool calls. So it might run tens or even hundreds of inference steps, sometimes over minutes or hours.
And the shared context keeps going the whole time. For the user, what matters is not the latency of a single token or one call. What matters is when is my task completed.
So agentic inference is not just chat with more requests. It's a different problem. The context grows over time. Tool work is interleaved between model calls. The number of model calls depends on the input. So you can't really plan around a fixed request rate plan. And the real metric is end-to-end task latency, not a single request latency. This is where Friendly AI comes in. We rebuilt the frontier inference cloud specifically for agentic workloads around the challenges I just walked through. And we set one goal: optimize end-to-end task latency, the task, not just the request. So how do we do that? Let me show you the key engineering behind it.
Here's the engineering map for how we think about it. We built the stack layer by layer around agentic workloads. There are four big pillars I'm going to cover today. Prefix caching, key value, also called KV cache, management, cache-aware routing, and agent-aware optimization. And of course, underneath, we need model layer optimization like sparse attention for long contexts, techniques to reduce errors, fast decoding, resilient serving, and more. In this talk, I'm going to focus on the four pillars. Let's start with prefix caching. Since agent steps share a large prefix, we compute key value for the prefix once and cache it.
Then on later steps, we reuse the cached key value and only process the new suffix. Reading from cache is much cheaper than recomputing prefill, so this improves time-to-first token and reduces compute on every step. And the longer the task runs in agents, the more valuable this becomes. But caching only works if the KV cache actually fits and can move around efficiently. So we need strong KV cache management. We use frugal memory management to pack more active contexts onto each GPU memory. We use KV quantization to reduce the memory footprint. We use hierarchical caching across GPU memory, host memory, and disk, so we can go beyond GPU limits.
And we also use distributed caching, so one prefix can be served across replicas, not just inside one instance. At global cluster scale, routing becomes really important.
A naive load balancer may spread requests evenly across GPU clusters, but it can destroy cache locality. A cache-aware router at a global scale does something smarter. It sends a request to a pod that already has the right prefix cached, turning a core prefill into one cache hit. At the same time, it still has to balance load, so one pod doesn't become a hotspot. In this example, the two requests of task A go to the same pod for cache locality. The next piece is agent-aware optimization, and this is the next frontier of agent inference. Today, most systems schedule each LLM call as if it were independent. They don't really understand that this call
is part of a longer agent program. But if the optimizer knows the agent-level context, it can make better decisions. For example, preempting the right work, speculatively prefilling context for a likely next step, or making a better cache eviction decision based on agent-level context. So the goal is to reduce end-to-end task latency, not just make one call look fast. When we put all of this together, this is the payoff.
We are using the same model, GLM 5.2, with code to create a simple mobile game. We ran the same task with model APIs of Friendly AI and another well-known inference provider. As you can see, Friendly AI completes the same task end-to-end much faster, thanks to our agent-centric cloud design. So what does this unlock in practice? A stronger production agent stack. Take an agent you already like. Now plug in open-weight frontier models like GLM 5.2, Minimax, and Kimi served on Friendly AI. The model gives you frontier quality capability and better economics. Friendly AI gives you the speed, reliability, and end-to-end task performance needed in production.
That combination—quality, speed, reliability, and cost— is what makes agents actually useful and economical in production. Friendly AI is currently powering teams in production from AI-native startups to global enterprises. I'd like to highlight a couple here. Cursor is a hugely popular agentic AI coding tool serving millions of users. LG is a global enterprise whose businesses range from electronics to healthcare to energy. Very different companies, but they all need the same thing: fast, reliable, cost-effective agentic inference. This testimonial from our client, Cursor, says it all.
Over the past year, Cursor has tested several inference providers hosting both open and closed models. In a split test of GLM 5 usage compared against other third-party providers and direct usage from the model lab, Friendly AI was consistently seven times faster with a significantly lower error rate. Today, Friendly AI is a core component of the Cursor stack. And you can consume this however fits your stack. Model API is the fastest way to start. Core frontier open-weight models through our serverless API. Dedicated endpoints give you your own isolated deployment with guaranteed SLAs for production workloads. And BYOG, bring your own GPU,
lets you run Friendly inference on your own infrastructure. Same stack, three ways to deploy. To wrap up, there are three things to remember. First, frontier open-weight models make production agents economically scalable. Second, agents are not just chat with more calls. Agentic inference requires optimizing end-to-end task latency with the challenges I mentioned. Third, Friendly AI is built as an inference cloud for this world. Fast, reliable, cost-effective agentic inference. Thank you for attending my session.
If you're building agents, give frontier open-weight models a try on Friendly AI today. You can get started at Friendly AI in minutes. Thank you. I'll be around after the session. Thank you. Thank you. Thank you.