# NVIDIA Dynamo adds session‑aware routing for agentic AI workloads

> The new Dynamo feature uses a stable session ID to keep context cached across multiple turns, improving speed for coding assistants and tool‑using agents.

Oossa · 2026-10-08 · https://oossa.com/en/nvidia-dynamo-adds-session-aware-routing-for-agentic-ai-workloads

NVIDIA Dynamo now supports a session‑level identifier that lets inference servers treat a chain of LLM calls as a single program. The ID is read from existing headers of popular coding agents such as Claude Code, Codex and OpenCode, and can be added to custom harnesses via a single X‑Dynamo‑Session‑ID header. With this identifier Dynamo can route requests, share KV cache, and pause sessions at tool boundaries, reducing the need to re‑prefill large contexts.

## The facts

- Supported agents include Claude Code, Codex and OpenCode.
- Custom harnesses use the header X‑Dynamo‑Session‑ID.

## Why it matters

Developers of AI coding assistants can see lower latency because the system reuses cached context instead of repeatedly loading it.

## Sources & references

1. [Session-Aware Agentic Inference with NVIDIA Dynamo](https://pytorch.org/blog/session-aware-agentic-inference-with-nvidia-dynamo/) – PyTorch

Last updated: 2026-10-08
