NVIDIA NeMo Relay Tracks AI Agent Behavior
NVIDIA has introduced NeMo Relay, an observability layer that helps developers trace and debug complex AI agent behaviors to eliminate hidden inefficiencies and token waste.

NVIDIA has detailed NeMo Relay, an observability layer that captures structured trajectories and ordered lifecycle events for AI agents. Integrated natively into the Hermes Agent harness, NeMo Relay helps developers look beyond simple success checks to understand exactly how an agent interacts with models and tools, revealing hidden inefficiencies and token waste.
The system outputs data in three formats. The Agent Trajectory Observability Format (ATOF) provides a JSONL log of lifecycle events, while the Agent Trajectory Interchange Format (ATIF) offers step-by-step JSON records of agent interactions. NeMo Relay also exports OpenTelemetry spans with OpenInference labels to platforms like Arize Phoenix. In a simple terminal-tool task, the runner verified the exact output VALUE=42. In a multi-tool research task using the nvidia/nemotron-3.5-lightning-30b-a3b model, these traces verified complex actions like file reading and web searching to identify the 'COLT 2026' conference.
To show how tracing aids in evaluating harness changes, a Hermes ToolPerf benchmark analyzed 108 runs across nine tasks. The evaluation compared baseline performance against fixed revisions using Claude Sonnet 4.5 and Qwen3 Coder 30B. For Claude Sonnet 4.5, the baseline achieved an 89% success rate (24 of 27 runs) with a 16-second mean duration, while the fixed version achieved 85% (23 of 27 runs) in 22 seconds. For Qwen3 Coder 30B, the fixes improved the success rate from 70% (19 of 27 runs) to 81% (22 of 27 runs). However, the traces revealed that this success came at a cost: mean LLM calls rose from 3.8 to 4.9, tool calls increased from 2.8 to 3.9, tool-result data jumped from 16 KB to 33 KB, and mean duration rose from 27 to 42 seconds.
For AI practitioners, this granularity changes how agent optimization is approached. Instead of guessing why an agent succeeded or failed, developers can pinpoint whether a change actually optimized the workflow or simply created a slower, more talkative agent. This systematic tracing ensures that developers do not mistake a high token-consuming recovery path for a genuinely efficient solution.
This is our own summary of reporting by NVIDIA Developer Blog



