LLM Routing: Affinity, Caching, and Gateways

How generative AI routing differs from load balancing: KV caching, model affinity, switching budgets, and when to use an inference router, an AI gateway, or both.

LLM Routing

In the era of Generative AI, routing requests is fundamentally different from load-balancing traditional microservices or predictive decision models. Large Language Models (LLMs) are highly stateful during a conversation. They rely on massive context windows, and processing that context repeatedly is both expensive and slow.

To solve this, the industry has standardized around Inference Routers and AI Gateways. These tools sit between your application and the models, using the ubiquitous OpenAI-compatible v1/chat/completions schema to route traffic dynamically while managing state.

The Standard Generative AI Schema

Whether you use a managed cloud router or an open-source gateway like LiteLLM, the integration looks identical to a standard ChatGPT API call. Instead of addressing a specific model, you address the router.

Here is a multi-turn request using advanced routing headers:

curl --location 'https://aispectrum.run/v1/chat/completions' \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $MODEL_ACCESS_KEY" \
  -H "X-Model-Affinity: example-session" \
  -H "X-Routing-Max-Switch-Spend-Pct: 20" \
  -d '{
    "model": "router:test-router",
    "messages": [
      {
        "role": "user",
        "content": "Write a Python function that implements binary search on a sorted array"
      },
      {
        "role": "assistant",
        "content": "<previous-response>"
      },
      {
        "role": "user",
        "content": "Now explain its time complexity"
      }
    ]
  }' | python3 -m json.tool

The model field is set to router:test-router. Behind the scenes, the router evaluates the incoming prompt and directs it to the best model. But how does it make that decision?

How Generative Routing Works

A router is a collection of logically grouped tasks and fallback models.

When a request arrives, the router analyzes the prompt and matches it to a task based on semantic descriptions (for example, a task named math with the description solving, explaining math problems, concepts). Once a task is matched, the router looks at that task's Model Pool and applies a Selection Policy:

  • Cost Efficiency: Selects the model with the lowest token costs.
  • Speed Optimization: Selects the model with the lowest Time-To-First-Token (TTFT).
  • Manual Ranking: Follows a strict, user-defined priority order.
  • Optimal: Uses hybrid benchmarking (public signals like LMSYS Chatbot Arena combined with in-house datasets) to pick the best overall model.

If the request doesn't match any task, or if the chosen model is down, the router sends it to a prioritized list of Fallback Models, so the end user sees no downtime.

Core Concepts of Stateful Routing

Unlike stateless API routing, generative AI routing must account for the heavy computation of text generation. This introduces three concepts.

1. Prompt Caching and the KV Cache

When an LLM processes a prompt, it generates a Key-Value (KV) cache, a numerical representation of the input tokens. In multi-turn conversations or agentic loops, recomputing this cache from scratch on every turn wastes compute and drives up latency.

Modern routers leverage Prompt Caching. The first request populates the cache (often billed at a slight premium for cache creation). Subsequent requests that reuse the cached prefix are billed at a large discount and run much faster.

2. Model Affinity

To take advantage of the KV cache, the router must send later requests from the same conversation to the exact same model instance that holds the cache. This is called Model Affinity.

By passing an identifier like X-Model-Affinity: example-session in your headers, you tell the router to pin the session to the currently selected model. This preserves the KV cache for repeated context, keeps behavior consistent across agentic loops, and keeps latency low.

3. Max Switching Budgets

Because routers continuously evaluate the model pool for better options, they face a dilemma: is it worth breaking Model Affinity to switch models?

Switching models means abandoning the warm KV cache and paying full price to recompute the context on the new model. To govern this, routers use a Max Switching Budget (for example, X-Routing-Max-Switch-Spend-Pct: 20).

The router compares the cached input cost of staying on the current model against the uncached cost of rebuilding the context on a candidate model. If the cost to switch exceeds the budget (for example, 20% more than staying), the router holds the request on the current model. Router analytics can track this directly, letting engineers monitor Requests Held versus Requests Switched.

Performance Trade-offs and Evaluation

Routing is not free. A dynamic routing layer typically adds a small latency overhead (approximately 200ms). To check the trade-off is worth it, modern routing platforms include LLM-as-a-Judge evaluation tools, so you can test routers for Completeness and Correctness against expected outputs before production.

Two Approaches to Generative Routing

The industry relies on two primary architectures:

  1. Managed Cloud Routers (e.g., DigitalOcean Inference Router): Cloud providers offer integrated inference routers that handle KV caching natively. They provide pre-configured templates for tasks like "software engineering," apply optimal selection policies, and manage fallbacks without any infrastructure for you to host.
  2. Universal AI Gateways (e.g., LiteLLM): For teams that need multi-cloud flexibility, open-source AI gateways are the standard. Deployed as a proxy, LiteLLM standardizes over 100 provider APIs into the same schema. You manage the proxy infrastructure, but you get full control over multi-provider load balancing, virtual key generation, strict financial budgets, and centralized spend tracking.

Why Agentic Apps Need a Routing Layer

Agentic applications, where LLMs run in autonomous loops, use tools, and build up large multi-turn context, are fragile and expensive when wired directly to a single model API such as OpenAI or Anthropic.

Without a router or gateway, an agentic app suffers from:

  • Catastrophic failures: If your hardcoded model hits a rate limit or goes down, the agent crashes.
  • Runaway costs: Agents consume huge numbers of tokens. Without a router enforcing Switching Budgets or a gateway enforcing Spend Limits, a rogue agent loop can drain your API budget in hours.
  • Cache thrashing: Without a router managing Model Affinity, your agent can lose its KV cache between turns, forcing the provider to recompute the entire system prompt and tool definitions on every step of the loop.

Doing nothing is off the table. You need a control plane between your agent and the models.

Router or Gateway: A Binary Choice?

At first glance the two look interchangeable, because both put the model behind a single API endpoint (for example, v1/chat/completions). Pick one and you often don't need the other.

  • Choose an AI Gateway (LiteLLM) if you prioritize organizational control. You generate virtual keys for different agent teams, track every fraction of a cent spent across AWS, GCP, and OpenAI, and set hard rate limits. Fallbacks are handled manually in the gateway config.
  • Choose an Inference Router (DigitalOcean) if you prioritize application performance and intelligence. The router reads the prompt semantically, matches it to a task, and calculates KV cache economics (Max Switching Budgets) on the fly.

Composition Over Competition

In larger architectures the two are not mutually exclusive. They can be stacked: an AI Gateway at the edge and an Inference Router at the application layer.

Here is the flow for an agentic app:

  1. The agent sends a request to LiteLLM (the gateway).
  2. LiteLLM checks the agent's Virtual Key, verifies it hasn't exceeded its $50/day budget, and logs the request for observability.
  3. LiteLLM forwards the request to the DigitalOcean Inference Router.
  4. The router reads the X-Model-Affinity header and checks the prompt. It sees a coding task and calculates that staying on the current model is 15% cheaper than switching, which respects the Max Switching Budget. It then routes the prompt to the exact hardware node holding the KV cache.

For smaller teams or specific projects, it is usually a binary choice: pick the tool that solves your most immediate pain, cost tracking (gateway) or cache and latency optimization (router). For large-scale agentic systems, the two are halves of a complete AI control plane.

How the Router Measures Speed

For the Speed Optimization policy, the router relies on continuous in-house benchmarking and real-time telemetry. It does not predict the latency of a single request.

  1. Continuous platform benchmarking. DigitalOcean monitors the performance of every model in its catalog. Its documentation describes a "hybrid evaluation methodology" that includes continuous in-house benchmarking on its own infrastructure, which measures baseline latency.
  2. Real-time telemetry. TTFT fluctuates with network traffic, GPU utilization, and platform load, so static benchmarks aren't enough. The router collects telemetry from live requests and tracks the TTFT each model is actually delivering.
  3. Dynamic selection at runtime. When a request hits a task using Speed Optimization, the router checks the models in your Model Pool against recent telemetry and routes the prompt to the one currently responding fastest.

Cache-Aware Overrides

If you use Model Affinity and Max Switching Budgets, the router may intentionally ignore raw TTFT data.

Say Model A is currently faster than Model B, but Model B holds your warm KV cache. The router weighs the time saved by Model A's speed against the time lost recomputing the entire prompt from scratch. If rebuilding the cache takes longer than the TTFT difference, the router holds the request on Model B to deliver the fastest overall response.

Conclusion

Routing for generative AI balances finding the cheapest or fastest model against keeping the economic benefits of the KV cache. By mastering Task Matching, Model Affinity, and Switching Budgets, engineering teams can build resilient, cost-effective LLM applications that scale in production.