NVIDIA Dynamo now supports a session‑level identifier that lets inference servers treat a chain of LLM calls as a single program. The ID is read from existing headers of popular coding agents such as Claude Code, Codex and OpenCode, and can be added to custom harnesses via a single X‑Dynamo‑Session‑ID header. With this identifier Dynamo can route requests, share KV cache, and pause sessions at tool boundaries, reducing the need to re‑prefill large contexts.
Why it matters
Developers of AI coding assistants can see lower latency because the system reuses cached context instead of repeatedly loading it.