Agents

NVIDIA Dynamo introduces session-aware agentic inference

NVIDIA has updated its Dynamo serving stack with session-level tracking to optimize KV cache management, boosting throughput for complex, multi-turn agentic AI workloads.

PyTorch Blog22 hrs agoAgents
Image: PyTorch Blog

NVIDIA has updated its Dynamo inference stack to address the unique traffic patterns of agentic AI workloads. Unlike traditional single-turn chatbots, AI agents often execute multi-step workflows with large initial prompts, parallel subagents, and long pauses during tool execution. To prevent the key-value (KV) cache from being prematurely evicted during these idle periods, Dynamo now utilizes a unified session-level identifier. This session ID maps related requests to a single trajectory, transforming request-level serving infrastructure into a program-aware system that optimizes cache retention.

A key component of this update is a native, Rust-based session-aware scheduler that integrates techniques from the ThunderAgent paper. The scheduler tracks worker utilization and pauses acting programs at tool boundaries under memory pressure, resuming them when capacity clears. In benchmarks running SWE-bench with two TP4 MiniMax-M2 replicas on an 8xH100 node, this program-aware scheduling improved throughput by 12 to 16 percent over standard KV-routing. Additionally, during Uni-Agent SWE-Bench tests on an 8xH20-3e system running Qwen3-Coder-30B-A3B-Instruct, the scheduler achieved 11.0 to 14.6 percent higher model-token throughput than VERL's default Global LB at medium concurrency, while maintaining a prefix-cache hit rate above 94.5 percent at high concurrency.

For practitioners, Dynamo introduces programmatic KV cache management through a proposed KvHint interface for vLLM and SGLang. This allows the router to send soft hints to the inference engine, directing actions like prefetching or pinning cache blocks. To support this, a Session-Prefix Indexer maps block locations to active session lifecycles. Dynamo also features a shared-pool indexer integrated with Mooncake storage to track offloaded cache pages. The system natively recognizes identity headers from popular coding agents like Claude Code, Codex, and OpenCode, and offers plugins for Hermes, OpenClaw, and Pi providers.

This is our own summary of reporting by PyTorch Blog

More in Agents