Open Reader

Productionizing LLM Gateways: Architecture, Tradeoffs and Hard Lessons — Kanish Manuja, Twilio

completed 16:23 Aug 28, 2026 Watch on YouTube

Current Status

completed

Video ID

zrZ1amZBSPw

RAG / Chat

Enabled
Productionizing LLM Gateways: Architecture, Tradeoffs and Hard Lessons — Kanish Manuja, Twilio
Description

Something went wrong, please try again. Kanish Manuja opens on that message and then explains why it exists, which is more interesting than laziness. Once a response starts streaming you have committed to that provider. Tokens already sent cannot be recalled, so the fallback you carefully built is unavailable exactly when you need it. Streaming buys perceived speed by trading away your levers, and that error string is what the trade costs. His frame for an LLM gateway is a permanent fight between availability, latency, guardrails and cost, where degradation forces you to give one of them up. The advice is refreshingly specific about where normal engineering instincts mislead. Retrying a slow expensive call eats the latency budget and multiplies spend, and tripping a circuit breaker is silly when a healthy second provider is sitting right there, so prefer per request fallback. Do not measure gateway wide latency, because a reasoning model's normal is a chat model's outage; track P99 per model per route and set timeouts the same way, since a missing timeout is his top cause of silent outages. Treat guardrails as services that fail too, and decide in advance whether you fail open or closed. He also argues most teams asking for a central gateway actually want centralized governance, which does not require centralizing the traffic. Speaker info: - https://www.linkedin.com/in/kanish-manuja-a99bb923/ Timestamps: 0:00 - The message behind the message 1:21 - Four things you cannot all maximize 2:33 - Why retries and breakers mislead here 3:39 - Per request fallback, and where failure counts live 4:49 - Fallbacks are not transparent 5:55 - Give the backup provider more headroom, not less 7:08 - Mixed workloads and the aggregate latency lie 8:17 - Reasoning and router models, 2 seconds to 60 9:28 - Hedging the tail 10:40 - Guardrails that fail open or closed 11:53 - Time budgets, fallbacks and placement 13:02 - The gateway as a new dependency 14:11 - Load shedding under a r

Summary

Generated by gpt-5.6-terra

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: A production LLM gateway must expose explicit tradeoffs among availability, latency, security guardrails, and cost rather than pretending it can maximize all four during degradation.
  • Why it matters: This is a concrete operations-oriented framework for designing resilient multi-model routing, observability, guardrails, and failure containment in agent and LLM systems.
  • Best use: Use it as an architecture and incident-readiness checklist for any gateway, model-router, or centralized AI control-plane design.

Executive Summary

Twilio principal engineer Kanish Manuja frames the LLM gateway as middleware between applications and model providers that handles routing, authentication, fallback, rate limits, and governance. His central point is that outages and degradation force choices between availability, latency, guardrails, and cost; a gateway should make those choices configurable by route and use case.

For provider reliability, he argues conventional API retry patterns are insufficient because LLM calls are slow and expensive. The preferred pattern is per-request cross-provider fallback, supported by circuit breaking and fleet-aware health state. But fallback is not inherently transparent: providers differ in tool-call schemas, token limits, and stop reasons, so gateways need normalization and rigorous compatibility testing. Streaming further constrains recovery because a provider cannot be swapped after tokens have reached the client.

Latency management must be route- and model-specific. A gateway-wide aggregate latency metric conceals radically different workloads, from sub-second embeddings to long-running reasoning calls. Manuja recommends P99 tracking and timeouts per model and route, fixed reasoning settings where possible, and optional tail hedging when a request consumes most of its latency budget.

The talk also treats guardrails and the gateway itself as dependencies that can fail. Teams need explicit fail-open versus fail-closed policies, guardrail time budgets and fallback mechanisms, bounded queues and load shedding, isolated API limits to prevent noisy-neighbor failures, and decentralized traffic paths with centralized governance rather than one company-wide gateway deployment.

Key Takeaways

  • Claim: Design LLM gateways around explicit degradation tradeoffs among availability, latency, guardrails, and cost. | Evidence: Manuja identifies these four concerns as being in direct tension during an incident; for example, parallel multi-provider requests can reduce latency risk but approximately double request cost. | Implication: Ken should require route-level policy controls rather than one global resilience policy for all model calls. | Caveat: There is no universal setting: the correct behavior depends on the route's business criticality, acceptable security exposure, and latency budget.
  • Claim: Use per-request multi-provider fallback and circuit breaking instead of relying on blind LLM retries. | Evidence: Retries with exponential backoff consume latency budget and multiply cost for slow LLM calls. The proposed sequence is provider A, then provider B on failure; when a primary is persistently unhealthy, remove it from the request path, place it in cooldown, and reintroduce it later. | Implication: Build provider routing that can select a healthy alternate per request, while making the fallback provider at least as well provisioned as the primary because it must absorb failover traffic. | Caveat: Parallel requests to multiple providers are appropriate only when latency is important enough to justify the extra spend.
  • Claim: Cross-provider failover requires a compatibility and normalization layer, not merely an OpenAI-compatible endpoint. | Evidence: The speaker cites differences in tool-calling schemas, token limits, and stop reasons even as providers converge on OpenAI-style APIs. | Implication: Treat fallback paths as first-class integration test targets, including tool use and streaming failure behavior, rather than assuming provider interchangeability. | Caveat: Streaming requests have an irreducible recovery limitation: after output begins reaching the client, the gateway cannot transparently switch providers midstream.
  • Claim: Measure and control latency per model and route, not through a gateway-wide aggregate. | Evidence: A single gateway may serve sub-second embedding and classification requests, roughly three-second chat requests, and much longer reasoning requests. Manuja calls aggregate gateway latency misleading and recommends P99 per model and route plus per-route timeouts. | Implication: Ken should segment SLOs, dashboards, timeout budgets, and alert thresholds by workload class; otherwise quiet latency failures will be hidden by blended metrics. | Caveat: Reasoning-model latency can be highly variable: the same prompt reportedly ranged from two to 60 seconds in production, and what is normal for reasoning may look like an outage for chat.
  • Claim: Make reasoning and router-model behavior as deterministic as possible, then hedge only the tail where justified. | Evidence: Reasoning models may not allow temperature zero and router models can obscure underlying model selection. The recommendation is to fix reasoning level per route and optionally launch a second request when the first has consumed about the P90 portion of its latency budget. | Implication: For high-value, latency-sensitive workflows, define deterministic route configurations first and reserve hedging for routes where the business value exceeds duplicated inference cost. | Caveat: Tail hedging improves P99 but deliberately increases request volume and cost.
  • Claim: Guardrails need their own reliability design and an explicit fail-open or fail-closed policy. | Evidence: Guardrail services can fail just like model providers. The talk recommends bounded guardrail time budgets, secondary checks or providers, and cached decisions. It distinguishes input prehooks, concurrent checks, and output posthooks. | Implication: Set the default to the worst outcome each route can safely tolerate, and tailor guardrail placement to the workload rather than applying a uniform security pipeline. | Caveat: Fail-open improves availability but can expose users to unfiltered content or security risk; fail-closed protects policy but turns a guardrail outage into application unavailability. Concurrent guardrails also fit poorly with streaming.
  • Claim: Avoid making a centrally deployed gateway a company-wide traffic choke point; centralize governance instead. | Evidence: Manuja warns that a central gateway is itself a single point of failure. He recommends decentralized gateway deployments or code/plugins while centralizing governance such as cost tracking and rate-limit management. | Implication: Separate the control plane from the data plane: maintain centrally managed policy and visibility while keeping application traffic paths isolated and failure-contained. | Caveat: A single team can still operate the platform; the objection is to one shared traffic deployment, not to shared ownership or policy.

Detailed Brief

Gateway failure containment and operational safeguards

  • Claims: The gateway adds a new dependency to the request path, so its own overload behavior must be designed rather than assumed.; Noisy tenants and retry storms can turn a localized provider issue into a gateway-wide outage.; Capacity planning for fallback is a resilience requirement, not an optional optimization.
  • Evidence: The speaker recommends segregating API keys and shared limits as granularly as possible by route and use case.; He calls load shedding a feature that should be validated in runbooks and game days because scaling out alone may not solve a retry storm.; Web-server queues should be bounded rather than allowed to accept unlimited requests; traffic prioritization can preserve critical workloads under load.; Failure counters may be local to serving instances or shared across the fleet; shared state enables faster fleet-wide failover, while local counters change behavior as deployment size changes.
  • Caveats: Fleet-wide health state can improve response speed but introduces shared infrastructure and coordination considerations.; Priority policies require a clear, pre-agreed definition of which use cases are important enough to retain during overload.
  • Implications: Incident exercises should include provider degradation, retry amplification, fallback saturation, queue exhaustion, and guardrail-provider failure rather than only full provider outages.; Rate-limit and credential partitioning should align with business isolation boundaries so one route or tenant cannot consume another's recovery capacity.

Guardrail placement choices

  • Claims: Input, concurrent, and output guardrails serve different purposes and impose different latency and streaming constraints.; Guardrails should not become the rate-determining step for ordinary LLM requests.
  • Evidence: A prehook evaluates input before the model call and is described as safest, but it adds serial latency.; Parallel guardrails can reduce end-to-end latency for structured outputs, but the speaker says streaming does not work well with this pattern.; Posthooks are positioned for output monitoring and auditing.
  • Caveats: The transcript does not prescribe a universal implementation for preventing unsafe partial output in streamed responses; that remains a design constraint when combining streaming with enforcement.
  • Implications: Use non-streaming structured-output flows where concurrent validation materially improves latency, and retain pre-execution enforcement where input risk cannot be tolerated.

Notable Concepts & Terms

  • LLM gateway: Middleware between applications and model providers that can enforce routing, authentication, fallbacks, rate limits, and governance.
  • Per-request fallback: Attempting an alternate provider for an individual failed request, which is better suited to expensive, slow LLM calls than repeated retries against one provider.
  • Circuit breaker with cooldown: Removing a persistently failing provider from routing temporarily, then testing it again after a cooldown rather than continuing to send production traffic.
  • Tail hedging: Launching a duplicate request after the original consumes much of its latency budget, intended to reduce P99 latency at additional cost.
  • Fail open / fail closed: The policy decision for guardrail failure: continue serving without the check or deny service to preserve safety/compliance.
  • Prehook / parallel check / posthook: Three placements for guardrails: before inference for safety, concurrently for lower latency in suitable non-streaming flows, or after inference for monitoring and auditing.
  • Centralized governance, decentralized traffic: A control-plane/data-plane split in which policy, cost tracking, and rate-limit management are shared centrally without routing all company traffic through one gateway deployment.

Operator Notes / Why Ken Should Care

  • Create a route-by-route degradation matrix covering provider fallback order, timeout, streaming behavior, cost ceiling, guardrail failure mode, and traffic priority.
  • Run compatibility tests for every intended provider fallback path, specifically including tool calls, token-limit handling, stop reasons, structured output, and partial-stream failure.
  • Replace blended gateway latency reporting with per-model, per-route P50/P95/P99 dashboards and route-specific alerts.
  • Verify that fallback provider quotas, throughput, credentials, and regional capacity can absorb primary-provider traffic during a real incident.
  • Add game-day scenarios for retry storms and guardrail outages; validate bounded queues, load shedding, and critical-traffic prioritization.
  • Review any proposed shared LLM gateway architecture for a control-plane/data-plane split before approving a single enterprise traffic deployment.

Source/Metadata

  • Title: Productionizing LLM Gateways: Architecture, Tradeoffs and Hard Lessons — Kanish Manuja, Twilio
  • Transcript words: 2313
  • Duration seconds: 983
  • Timestamp note: No timestamps or chapters were present in the supplied transcript.

Transcript

2193 words en Processed in 93.5s

I'm Kanish Manojo. I'm a principal engineer at Twilio. Let's start with a quick show of hands. Who here has seen the message, "Something went wrong, please try again?" Well, we have a few lucky ones and a few that have had a good lunch. So, behind that simple message is actually a system that is very complex, that serves you that message despite the model providers being down. And that's what we're going to productionize today, or discuss productionizing today. So, what is an LLM gateway? An LLM gateway is an entry point or a middleware between your apps and the model providers behind them. It does a bunch of things: routing, authentication, fallback, rate limits, all kinds of governance that you can think of. And right at the heart of the gateway is a fight between four things. It's availability, latency, your guardrails, and costs. In case of a degradation, you cannot maximize all four. You need to pick what you want. So, with this talk, if you use an LLM gateway, I want to help you make that tradeoff for your use case. And if you design a gateway, I want you to design or provide those levers to your callers and customers so that your customers are happy. Let's start with availability. If you have a single model provider, their ceiling is your ceiling. Their outage is your outage. So, in typical software engineering, the way you tackle an unreliable dependency is by retrying. Retrying with exponential backoffs, with jitters. And when all of that fails, you have a circuit breaker that trips after you've seen sufficient failures, and you stop calling the damn thing. This is not enough for LLMs. LLMs are very different compared to your fast, cheap APIs that you retry on. Retrying an LLM API eats into your latency budget really fast. Also, tripping over a circuit breaker when you have another perfectly fine model provider to route to doesn't make sense. You should use the second model provider. And third, as I said, the calls are slow and expensive. So, blind retries just multiply your cost and your tail latencies. So, what is a better idea here? It is actually a per-request fallback. What that means is you can actually try model provider A and then, in sequence, try model provider B if your request to model provider A fails. Another option to consider here is you can fire requests to both providers in parallel. But that's only if you're highly, highly obsessed with latencies, because that's just going to double your cost. Some of the similar circuit-breaking patterns apply here to LLMs as well. If you know that your primary has been failing for some time, it doesn't make sense to try it again. You take it out of the load balancer or your request path and put it in a cooldown, and then after a few minutes have passed, try putting that back again. One interesting choice that you have to make here is where your failure counts live. You can decide to have the failure counts live in memory on the instances that are serving your traffic, or you can have shared infra where your failure counts are shared across the fleet. There are trade-offs. If you want quick failovers, then fleet-wide helps. And with local state counters on instances, the issue that you run into is whenever you change your deployment size, your configuration and your expectations change. So, something to consider. What that clean diagram did not really show you are some of the other gotchas that I'm going to discuss. So, fallbacks are not transparent. While the industry is converging on an OpenAI API-compatible format, I would say there are still nuances. So, you need to really test your fallbacks well. They can have differences in your tool-calling schemas, token limits, stop reasons, and what have you. So, with LLM gateways, you can have a normalization layer that can ensure that you can do cross-provider fallbacks as well. Another thing is streaming. Essentially, nobody wants to wait for 30 seconds to have a wall of text appear in front of them. So, there are use cases where streaming is absolutely required. But it comes at a cost. You trade away your levers. Once you've decided to go with provider A, you have to continue going with provider A. You cannot midstream change the providers. Whatever has been sent to the client, it's done. And that's where the "something went wrong" message comes in. That's the one that you see. It's not because of laziness. It's by design that you see that. And it's one of the trade-offs. I would like to call out one other thing where I've seen teams trip over and over again. They really provision and test their primary providers really well. But the second provider, the fallback provider, doesn't necessarily get the same level of love. And I would argue that your throughput, your capacity, or your headroom should be even higher for the second provider, or the fallback provider. Because that's your last line of defense. If that goes down, your application goes down. Let's discuss latencies. Availability failures are right in your face. They fail. You get alarmed. You get paged. But high latencies can be the quiet ones. And they need to receive more love than, I would say, tuning your services for just availability. One thing to call out: a gateway may run mixed workloads. And you can have embedding requests that take less than a second. You can have classification requests that take less than a second. You have chat requests taking three seconds. And reasoning requests taking a long time. Quick show of hands if you measure your aggregate latency for your entire service. Well, that was a trick question. Sorry. You shouldn't. It doesn't make sense. It's a lie. You should be tracking your P99 per model, per route, not a gateway-wide number. A gateway-wide number doesn't make sense, especially if you're running mixed workloads. And I hope you're not, for those who've raised your hand. Another thing that can really, I cannot emphasize this enough, is for you to set timeouts per model class, per route. That's the number one root cause of your silent outage. If you don't have a timeout, your gateway thinks your request is being happily served, while it is not. And I'll leave you with this message for latencies, specifically. A reasoning model's normal is actually a chat model's outage. So, you definitely need to track latency per route. Okay. This is the most painful slide, or the slide that has given me the most scars, which is reasoning and router models. So, this is where, truly, the latency is unpredictable. And reasoning models, they do not give you, they're highly undeterministic, more undeterministic than your normal models. You cannot set the temperature to zero in many cases. And the same prompt can take somewhere from two seconds to 60 seconds. And we have seen that in production, where P99 suddenly popped to 60 seconds for no good reason. So, while there's no magical solution to it, I would recommend that you at least start with fixing the reasoning level per route. So, with router models, they hide that abstraction from you, like they pick which models to run. And I would highly recommend that you at least make requests as deterministic as possible with an undeterministic system. Another idea is hedging the tail. You can fire another request if your primary request actually consumed, let's say, P90 of your latency budget. This can hedge the tail. This can really hedge the P99 tail for your services. All right, this is one of my favorite ones. To keep your model secure, you need to have guardrails. And guardrails are necessary for preventing your services from prompt injection attacks, keeping PII filters in place, having toxicity filters, keeping the LLMs from swearing at your customers. All those good things. But just like a model provider, there are trade-offs too. Guardrails are just another service. That can go down, that can be unreliable. And that's where you need to choose: do you fail open or do you fail closed? When I say fail open, you can still serve the request even if your guardrails are down. Fail closed, you block the request and say, "Hey, I'm not available." That's the trade-off between availability and security to a certain extent. While there's no universal answer, it really depends on your use case. You can decide, for example, that if a toxicity filter is not up and running, you can still serve that request. So, the default choice should be the worst case that you can live with. There are a few things that you can actually do to improve the behavior of your systems in the face of guardrails being down and managing the unreliability of the guardrails themselves. So, the first is time budget. Your request should never be bound by your guardrail timing. It should always be the LLM that is the rate-determining step. So, make sure that you have timeouts in place and those guardrails run with a specific time budget. Another important thing is fallback. You've heard, you probably know, and I've talked about it, we always discuss fallbacks with regard to model providers. But guardrails are critical services too, where you can consider fallbacks, have secondary providers, secondary checks, cached decisions to keep your service available when a guardrail provider is down. Another interesting choice that pops up with regard to guardrails is the placement of the guardrails. Typically, you can place the guardrail in three ways. You can have a prehook that runs where the guardrail actually runs on the input. And that's probably the safest, but it does add serial latency to your requests. Another one is in parallel. This is one of my favorites, but just to call out, streaming wouldn't work well here in parallel. So, if you're especially producing structured output, please don't stream them. Try to save your latencies and run these guardrails concurrently for your structured outputs. Another one is posthooks. These are best for output monitoring, auditing your outputs, and so forth. So far, I've discussed all the things that can go wrong with regard to our dependencies. We haven't discussed that we are actually adding another dependency in the request path itself, which is the central, or which is the LLM gateway itself. There are a few things where we have been bitten, and we've learned some lessons that I want to share with you if you're working on an LLM gateway or using one. One is shared limits. Make sure that your API keys are segregated per route, per use case, to the most granular possible, to the most granular thing that you can imagine. Having a noisy tenant can be one of the biggest problems here. Another thing is load shedding. This is a feature that you should, as part of your runbooks and game days, make sure that the gateway that you're using supports. Because when you have a retry storm, it becomes really hard to just scale out. You cannot simply scale out services that are under a retry storm. And all these web servers, they have an internal queue, and they're configurable. Make sure that they're bounded, and they cannot accept requests that are unbounded. And if you want to have some custom logic, you can even have traffic prioritization here as well, to make sure under load, your most important use cases get served well. Last thing that I wanted to discuss is the whole idea of a central gateway itself. It is a single point of failure. So, if you're thinking of having a central gateway for your entire company, for all LLMs, I would recommend rethinking that and seeing what the reasons are that you want it. What I've noticed is that in most scenarios, it's not the central gateway that they want. They want centralized governance. And there is a path forward where you can actually decentralize the gateway and still centralize governance. So, do not try to centralize your traffic, but you can have plugins, you can have custom code. That can centralize your governance. Governance can be in the form of cost tracking, rate limit management, and there are other solutions possible. So, explore those before you chart on having one central gateway for your entire company. It can be managed by a single team, but I wouldn't recommend deploying it as a single deployment for the entire company, even though it's distributed. With that said, I want to end this talk on a personal note. It is my son's birthday today, and I'm here talking to strangers about circuit breaking. So, the least you can do for me is please go and prevent one incident for me and for your customers. Thank you. If you have any questions, yeah. Thank you. That can centralize your governance. Governance can be in the form of cost tracking, rate limit management, and there are other solutions possible. So explore those before you chart on having one central gateway for your entire company. It can be managed by a single team, but I wouldn't recommend deploying it as a single deployment for the entire company, even though it's distributed. With that said, I want to end this talk on a personal note. So it is my son's birthday today, and I'm here talking to strangers about circuit breaking. So the least you can do for me is please go and prevent one incident for me and for your customers. Thank you. If you have any questions, yeah. Thank you.