Skip to content

Context Builder & Orchestration

Assembling the right context before the model runs — context compilers, budgets, KV-cache, and MCP federation.

11 articles

Aug 16, 2026·14 min read

What Breaks Prompt Caching: The Stable Prefix Contract

How prompt cache keys are derived across Anthropic, OpenAI, Gemini, vLLM, and SGLang, and the six agent mistakes that silently break the stable prefix.

Jul 21, 2026·24 min read

Memory MCP Servers Compared: 13 Real Options

Memory MCP servers compared: thirteen named options across knowledge-graph, vector, Markdown, and SQL storage, with pricing, and filtered by the local-first vs hosted decision.

Jul 21, 2026·9 min read

Stateless MCP: Where Agent Memory Lives Now

The MCP 2026-07-28 release candidate deletes protocol sessions, the Mcp-Session-Id header, and the handshake. Agent memory moves to an external store.

Jun 19, 2026·11 min read

Context Budgeting: Zones, Allocation & Eviction

Context budgeting allocates finite agent tokens across system, tools, retrieval, history, outputs, and response buffer.

Jun 19, 2026·13 min read

Context Optimizer: Cache, Budget & Placement

Context optimization for AI agents unifies KV-cache hit rate, prefix stability, token budget, latency, cost, and placement into one runtime decision.

Jun 18, 2026·8 min read

Context Compiler vs Orchestration

Where flow control ends and window assembly begins: the boundary between orchestrator and context compiler in LLM agent systems.

Jun 18, 2026·12 min read

Deterministic vs LLM Context Assembly

Deterministic context assembly improves cacheability and auditability; LLM-directed assembly adds adaptivity. Most agent systems need both.

Jun 16, 2026·15 min read

Context Engineering Needs a Compiler

Context engineering is the discipline. The context compiler is the per-turn runtime layer that ranks, budgets, secures, and assembles each model call.

Jun 15, 2026·10 min read

Memory MCP: How to Give AI Agents Persistent Memory

Memory MCP servers explained: what they are, how to choose one by where data lives and what it does, and how to install so an agent remembers across sessions.

Jun 10, 2026·10 min read

Federated MCP: How MCP Federation Works in 2026

MCP federation in 2026: what gateways and the June 2025 spec solved for running multiple MCP servers, and which problems, like auth propagation, remain open.

Jun 6, 2026·19 min read

Prompt-Cache Keys: Stable Prefix, KV-Cache Hit Rate

How prompt-cache keys are canonicalized, why a stable prefix decides your KV-cache hit rate, and what each provider actually discounts on a cache read.

← Back to the Library