LLM Routing: Affinity, Caching, and Gateways

October 3, 2026 • 8 min read

LLM Routing: Affinity, Caching, and Gateways

How generative AI routing differs from load balancing: KV caching, model affinity, switching budgets, and when to use an inference router, an AI gateway, or both.

The AI-Native Cloud: The Agent Stack

October 2, 2026 • 30 min read

The AI-Native Cloud: The Agent Stack

A map of the AI-native cloud (managed agent runtimes, serverless inference, routing, caching, knowledge bases, evaluations, GPUs) followed by a product-by-product tour of what a full AI cloud offers, with illustrative pricing.

Decision Models: Jev and the End of Prompting for Classification

October 1, 2026 • 6 min read

Decision Models: Jev and the End of Prompting for Classification

Jev from TypeSafe AI understands language at a frontier level but never generates a token. It is a non-autoregressive decision model on DigitalOcean Serverless Inference that returns typed answers with calibrated probabilities. Here is how it differs from an LLM, with working API calls for all three question types.

Time for an Oil Change on Your Coding Harness

August 18, 2026 • 7 min read

Time for an Oil Change on Your Coding Harness

Coding agents like Claude Code load a fixed startup budget of system prompt, tools, memory, and skills before you type a single word. Here's how to audit that budget and prune the biggest offender: bloated memory files.

Loading more articles…