NVIDIA NeMo Relay makes an AI agent’s execution path inspectable: it records lifecycle events around model calls, tool calls, and other work, then projects those events into formats for debugging, trajectory review, or observability backends. A trace can help explain how an agent reached an answer, but it does not prove the task succeeded; that requires a separate verifier. Relay instruments execution boundaries—it does not choose the agent’s plan or replace its framework.
What NeMo Relay does—and what it does not do
NVIDIA describes Relay as a shared runtime for scopes, policy, plugins, and lifecycle events. It sits around execution units such as a session, turn, LLM call, tool call, or subagent run. Through instrumentation, middleware, plugins, and events, it exposes or controls those boundaries while leaving application logic and orchestration with the surrounding application or framework. NVIDIA’s tutorial and the Relay overview describe this division of responsibility.
In its Support and FAQs documentation, NVIDIA states: “NeMo Relay does not choose the next step, schedule a multi-agent workflow, own a planner, or decide which tool an agent should call.” Relay is therefore not an agent orchestrator, model provider, prompt-authoring product, vector database, hosted tracing service, or full agent workbench. NVIDIA’s FAQs explain the product boundaries.
Where an integration fits
The appropriate integration depends on where the work happens and which component owns it. Relay documentation describes a local CLI sidecar, direct SDK instrumentation for application-owned calls, maintained framework integrations, wrappers, and plugins. The overview outlines these choices; Relay does not take over the application’s orchestration in any of them.
Quick wins for a faster PC:
Repair Windows errors before they cause bigger problemsFix Now →Scan for outdated or missing drivers - takes under a minuteDriver Scan →#1 Best Overall
- EVOLUTION AMD RYZEN AI MAX+ 395 MINI PC - GMKtec EVO-X2 is the next evolution in AI mini PC Ryzen Strix Halo series. Thanks to AMD Simultaneous Multithreading (SMT) the core-count is effectively doubled, to 32 threads. Ryzen AI Max+ 395 has 64 MB of L3 cache and can boost up to 5.1 GHz, depending on the workload. The Ryzen AI Max+ 395 is currently rated as the "most powerful x86 APU" on the market for AI computing.
- AI NPU with XDNA 2 ARCHITECTURE - Powered by 16 “Zen 5” CPU cores, 50+ peak AI TOPS XDNA 2 NPU and a truly massive integrated GPU driven by 40 AMD RDNA 3.5 CUs, the Ryzen AI MAX+ 395 is a transformative upgrade and delivers a significant performance boost over the competition. The Ryzen AI Max+ 395 excels in consumer AI workloads like the llama.cpp-powered application: LM Studio. Shaping up to be the must-have app for client LLM workloads, LM Studio allows users to locally run the latest language model without any technical knowledge required and unleash their creativity and productivity.
- AMD RADEON 8090S iGPU GAMING PC - The AMD Radeon RX 8060S offers all 40 CUs with up to 2.9 GHz graphics clock and uses the new RDNA 3.5 architecture. The powerful iGPU is positioned between an RTX 4060 and 4070 laptop GPU and therefore enables gaming in FHD at maximum details in most demanding games. The 8060S can also utilize the full 128GB pool, which is perfect for running LLMs such as Deepseek 70B Q8, which runs comfortably on this machine.
- EIGHT CHANNEL LPDDR5X - LPDDR5X is a new ground breaking memory small form factor installed on-board. With blazing speeds up to to 8000MT/s, it runs 1.5x faster than the DDR5 SODIMMs; 90% better performance over DDR5 SODIMMs in video conferencing and photo editing; 30% better performance in productivity apps; 12% better performance in digital content workloads.
- QUAD SCREEN 8K DISPLAY SUPPORT - EVO-X2 AI Mini PC support 4-screen 4K/8K output via HDMI 2.1 (8K@60Hz), DisplayPort 1.4 (4K@60Hz), and dual USB 4 40Gbps Transfer speed (supporting PD3.0/DP1.4/DATA). Ideal for gaming, video editing, and multitasking, it provides expansive and crisp multi-display support.
What NeMo Relay traces contain
Relay’s canonical event format is ATOF (Agent Trajectory Observability Format) 0.1. It represents work with two kinds of events: scopes and marks. A scope has a start and end and represents timed work, such as an agent, tool, or LLM call. Its start and end pair by UUID, while parent UUIDs preserve the relationship between nested work. A mark is a point-in-time checkpoint, not a timed pair. Relay-generated timestamps are used by default. The events documentation defines these semantics.
Depending on configuration and the exported projection, trace data may include prompts, model responses, tool arguments and results, file paths, timing, identifiers, and error information. The formats are not interchangeable: an export may omit information or represent it differently from the underlying event stream.
Choose an output format for the question you need to answer
| Format | Best suited to | What it represents |
|---|---|---|
| ATOF JSONL | Debugging or auditing individual events | Event-level record with timing, IDs, parent-child relationships, scopes, and marks. |
| ATIF | Reviewing or evaluating an agent’s trajectory | Step-by-step trajectory assembled from lifecycle events. It omits marks because its model is trajectory steps, not independent checkpoints. |
| OpenTelemetry, including OpenInference projection | Sending telemetry to an OTLP-compatible observability system | Spans and related telemetry. In NVIDIA’s tutorial, Phoenix is used to inspect model and tool calls, duration, token use, errors, and available inputs and outputs. The tutorial also names LangSmith as an OTLP-compatible destination. |
Relay’s event-format documentation describes event and projection behavior. For a backend view, NVIDIA’s tutorial demonstrates Phoenix. These destinations are optional; an observability backend is not a prerequisite for Relay.
Rank #2
- Built for Local AI Development: AMD Ryzen AI Halo is designed for local AI development and inference, featuring 128GB unified memory and support for up to 200B parameter models to build and run intensive AI workloads locally.
- 128GB Unified Memory: Features 128GB LPDDR5x unified memory at 8000 MT/s with 256 GB/s memory bandwidth, providing a shared memory pool across the CPU, GPU, and NPU to support larger AI models.
- AMD Ryzen AI Max+ 395 Processor: Features 16 cores, 32 threads, and Zen 5 architecture, paired with AMD Radeon 8060S integrated graphics featuring 40 RDNA 3.5 compute units and an AMD XDNA 2 NPU with up to 50 TOPS.
- Linux AI Developer Platform: Purpose-built for Linux-based AI development with full AMD ROCm software support and preloaded tools, models, and workflows optimized for local AI development.
- Compact, Connected Design: Includes a 2TB M.2 SSD, 10GbE LAN, Wi-Fi 7, Bluetooth 5.4, USB-C connectivity, and HDMI 2.1b.
How to tell whether a tool call succeeded
A trajectory’s tool request shows what the model asked to run; it does not establish that the tool completed successfully. To inspect the recorded outcome, find the corresponding ATOF tool scope and examine its start and end events, along with any error data. The scope UUID pairs the boundaries, and parent UUIDs show how the tool call fits into its parent work. Then compare the recorded execution with a separate task verifier: a tool can run without achieving the requested result, and a trace alone is not a success check.
Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchPC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11What one Relay run can—and cannot—show
In a Hermes Agent tutorial published September 30, 2026, NVIDIA demonstrates a verified terminal-tool run. The runner checks for the exact expected output VALUE=42, confirms completed LLM activity and zero tool errors, and verifies that ATOF and ATIF artifacts exist. NVIDIA reports 74 ATOF events, two completed LLM scopes, 7,239 prompt tokens, 96 completion tokens, 7,335 total tokens, one tool call, zero tool errors, and three ATIF steps for that run. These are results from that tutorial run, not general performance guarantees; NVIDIA notes that token counts, identifiers, and file paths can vary between runs. The tutorial gives the example and its qualifications.
Why compare runs instead of trusting one result
In a larger case study, NVIDIA reports an August 6, 2026 rerun of Hermes ToolPerf across nine tasks, comparing pinned baseline and fixes arms. Each task was run three times per model per arm, for 108 runs total. A task verifier measured completion while Relay ATOF captured model calls, tool calls, errors, retries, result data, and timing. NVIDIA’s tutorial reports these results:
Rank #3
- EVOLUTION RYZEN AI MAX+ 395 MINI PC - GMKtec EVO-X2 is the next evolution in AI mini PC Ryzen Strix Halo series. Thanks to AMD Simultaneous Multithreading (SMT) the core-count is effectively doubled, to 32 threads. Ryzen AI Max+ 395 has 64 MB of L3 cache and can boost up to 5.1 GHz, depending on the workload. The Ryzen AI Max+ 395 is currently rated as the "most powerful x86 APU" on the market for AI computing.
- AI NPU with XDNA 2 ARCHITECTURE - Powered by 16 “Zen 5” CPU cores, 50+ peak AI TOPS XDNA 2 NPU and a truly massive integrated GPU driven by 40 AMD RDNA 3.5 CUs, the Ryzen AI MAX+ 395 is a transformative upgrade and delivers a significant performance boost over the competition. The Ryzen AI Max+ 395 excels in consumer AI workloads like the llama.cpp-powered application: LM Studio. Shaping up to be the must-have app for client LLM workloads, LM Studio allows users to locally run the latest language model without any technical knowledge required and unleash their creativity and productivity.
- AMD RADEON 8090S iGPU GAMING PC - The AMD Radeon RX 8060S offers all 40 CUs with up to 2.9 GHz graphics clock and uses the new RDNA 3.5 architecture. The powerful iGPU is positioned between an RTX 4060 and 4070 laptop GPU and therefore enables gaming in FHD at maximum details in most demanding games. The 8060S can also utilize the full 128GB pool, which is perfect for running LLMs such as Deepseek 70B Q8, which runs comfortably on this machine.
- EIGHT CHANNEL LPDDR5X - LPDDR5X is a new ground breaking memory small form factor installed on-board. With blazing speeds up to to 8000MT/s, it runs 1.5x faster than the DDR5 SODIMMs; 90% better performance over DDR5 SODIMMs in video conferencing and photo editing; 30% better performance in productivity apps; 12% better performance in digital content workloads.
- QUAD SCREEN 8K DISPLAY SUPPORT - EVO-X2 AI Mini PC support 4-screen 4K/8K output via HDMI 2.1 (8K@60Hz), DisplayPort 1.4 (4K@60Hz), and dual USB 4 40Gbps Transfer speed (supporting PD3.0/DP1.4/DATA). Ideal for gaming, video editing, and multitasking, it provides expansive and crisp multi-display support.
| Model and measure | Baseline | Fixes |
|---|---|---|
| Claude Sonnet 4.5: task completion | 24/27 tasks (89%) | 23/27 tasks (85%) |
| Claude Sonnet 4.5: mean duration | 16 s | 22 s |
| Qwen3 Coder 30B: task completion | 19/27 tasks (70%) | 22/27 tasks (81%) |
| Qwen3 Coder 30B: mean LLM calls | 3.8 | 4.9 |
| Qwen3 Coder 30B: mean tool calls | 2.8 | 3.9 |
| Qwen3 Coder 30B: mean tool-result data | 16 KB | 33 KB |
| Qwen3 Coder 30B: mean duration | 27 s | 42 s |
In this sample, the fixes changed little for Sonnet, while Qwen completed three more tasks out of 27 and also used more calls, returned more tool-result data, and took longer on average. Task-level audits showed why aggregate success alone was incomplete: recovering from a blocked command improved completion but required more turns; case-insensitive search triggered extra exploratory searches on some repetitions; and a hidden-file search failure remained unresolved. These results illustrate a particular harness, model set, and workload—not a general prediction for other agents.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.A controlled way to evaluate an agent change
NVIDIA recommends treating task completion as the primary outcome and traces as evidence for diagnosing the path behind it. For a useful comparison:
- Define an exact automated success check. Specify what output or state proves that the task is complete.
- Set a baseline and one focused change. Isolating the change makes it easier to connect an outcome to a prompt, tool, or harness adjustment.
- Hold conditions constant. Keep the model snapshot, provider, task input, execution budget, and timeout the same for both arms.
- Repeat both arms equally. Compare repeated runs rather than relying on a single favorable result.
- Compare verified outcomes first. Then use traces to inspect calls, retries, errors, elapsed time, token use, and cost.
- Repeat across intended workloads or models. A result for one model or task set does not establish that a change will help elsewhere.
A faster run or a lower call count on its own does not establish an optimization: it may also fail more often or behave differently on repeated inputs. The verifier answers whether the task passed; the trace helps explain the behavior that produced that result. NVIDIA’s comparison guidance uses this distinction.
Rank #4
Protect trace data before sharing it
Because traces may contain prompts, responses, tool arguments and results, file paths, and other application data, treat exported artifacts as potentially sensitive. Review and sanitize them before sharing, and account for the selected exporter’s behavior: marks remain in ATOF but are omitted from ATIF, while OpenTelemetry projections may represent event data differently. NVIDIA’s tutorial cautions about trace contents, and the events documentation describes the format distinctions.
Which exporter should you use?
- Use ATOF JSONL when you need the underlying event-level record to investigate timing, IDs, nesting, marks, and recorded outcomes.
- Use ATIF when a step-by-step trajectory is the useful view for review or evaluation, and independent mark events are not needed.
- Use OpenTelemetry or an OpenInference projection when you want spans and related telemetry in an OTLP-compatible observability system; check the projection and backend to understand which payloads and event details are retained.
The integration point is a separate choice: use the path that fits ownership of the real work—CLI sidecar, SDK, framework integration, wrapper, or plugin. Before selecting an export or sharing traces, consider whether it retains the marks and payload details you need, what data can be sanitized, and whether the accompanying evaluation uses fixed inputs and task verification rather than trace volume or one-run speed.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




