Assembling the right context before the model runs — context compilers, budgets, KV-cache, and MCP federation.
16 articles
The canonical Cursor Memory Bank ships fifty five rule files. Exactly one of them is guaranteed to reach a session, and it tells the agent where to write, not what to read. Checked against the repository, the vendor's own pages and the forum threads where it fails.
MCP defines no ordering, conflict detection or merge semantics for memory. Coherence is a property each server chooses, and shipped servers choose differently.
The format has no required fields and one sentence about precedence. Seven tools that list themselves as supporting it describe seven different mechanisms, and one of them says outright that a closer file overrides an earlier one because it comes later in the combined prompt.
A rules file is text placed in a context window. All three vendors say so on their own pages, in their own words. What delivers a rule, what the wrapper around it says, and what does enforce, because something does.
MCP standardises how a client discovers a tool, calls it, and reads the answer. It does not standardise memory. Walking the whole path with the specification and the public reference server open.
How prompt cache keys are derived across Anthropic, OpenAI, Gemini, vLLM, and SGLang, and the six agent mistakes that silently break the stable prefix.
Thirteen memory MCP servers across knowledge-graph, vector, Markdown and SQL storage, with pricing, maintenance reality, and the local-first vs hosted decision.
The final MCP 2026-07-28 specification removes protocol sessions, the Mcp-Session-Id header, and the handshake. Durable agent memory remains external.
Context budgeting allocates finite agent tokens across system, tools, retrieval, history, outputs, and response buffer.
Context optimization for AI agents unifies KV-cache hit rate, prefix stability, token budget, latency, cost, and placement into one runtime decision.
Where flow control ends and window assembly begins: the boundary between orchestrator and context compiler in LLM agent systems.
Deterministic context assembly improves cacheability and auditability; LLM-directed assembly adds adaptivity. Most agent systems need both.
Context engineering is the discipline. The context compiler is the per-turn runtime layer that ranks, budgets, secures, and assembles each model call.
Memory MCP servers explained: what they are, how to choose one by where data lives and what it does, and how to install so an agent remembers across sessions.
MCP federation in 2026: what gateways and the June 2025 spec solved for running multiple MCP servers, and which problems, like auth propagation, remain open.
How prompt-cache keys are canonicalized, why a stable prefix decides your KV-cache hit rate, and what each provider actually discounts on a cache read.