Hardware FixRecommendedDevice not working? Your driver may be the problemCheck updates for common hardware issues.Fix DriversOctober DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsWindows FixRecommendedWindows errors stealing your time? Find the fix fastScan stability, cleanup and performance issues.Fix Now×
Skip to content
Santa Clara desk6 min

How NVIDIA NeMo Relay Traces AI Agent Execution

NeMo Relay records AI agent execution boundaries so developers can inspect model and tool work. Learn what its trace formats show, what they omit, and how to compare runs reliably.
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

NVIDIA NeMo Relay makes an AI agent’s execution path inspectable: it records lifecycle events around model calls, tool calls, and other work, then projects those events into formats for debugging, trajectory review, or observability backends. A trace can help explain how an agent reached an answer, but it does not prove the task succeeded; that requires a separate verifier. Relay instruments execution boundaries—it does not choose the agent’s plan or replace its framework.

What NeMo Relay does—and what it does not do

NVIDIA describes Relay as a shared runtime for scopes, policy, plugins, and lifecycle events. It sits around execution units such as a session, turn, LLM call, tool call, or subagent run. Through instrumentation, middleware, plugins, and events, it exposes or controls those boundaries while leaving application logic and orchestration with the surrounding application or framework. NVIDIA’s tutorial and the Relay overview describe this division of responsibility.

In its Support and FAQs documentation, NVIDIA states: “NeMo Relay does not choose the next step, schedule a multi-agent workflow, own a planner, or decide which tool an agent should call.” Relay is therefore not an agent orchestrator, model provider, prompt-authoring product, vector database, hosted tracing service, or full agent workbench. NVIDIA’s FAQs explain the product boundaries.

Where an integration fits

The appropriate integration depends on where the work happens and which component owns it. Relay documentation describes a local CLI sidecar, direct SDK instrumentation for application-owned calls, maintained framework integrations, wrappers, and plugins. The overview outlines these choices; Relay does not take over the application’s orchestration in any of them.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
GMKtec AI Mini PC Ryzen Al Max+ 395 (up to 5.1GHz) Mini Gaming Computers
  • EVOLUTION AMD RYZEN AI MAX+ 395 MINI PC - GMKtec EVO-X2 is the next evolution in AI mini PC Ryzen Strix Halo series. Thanks to AMD Simultaneous Multithreading (SMT) the core-count is effectively doubled, to 32 threads. Ryzen AI Max+ 395 has 64 MB of L3 cache and can boost up to 5.1 GHz, depending on the workload. The Ryzen AI Max+ 395 is currently rated as the "most powerful x86 APU" on the market for AI computing.
  • AI NPU with XDNA 2 ARCHITECTURE - Powered by 16 “Zen 5” CPU cores, 50+ peak AI TOPS XDNA 2 NPU and a truly massive integrated GPU driven by 40 AMD RDNA 3.5 CUs, the Ryzen AI MAX+ 395 is a transformative upgrade and delivers a significant performance boost over the competition. The Ryzen AI Max+ 395 excels in consumer AI workloads like the llama.cpp-powered application: LM Studio. Shaping up to be the must-have app for client LLM workloads, LM Studio allows users to locally run the latest language model without any technical knowledge required and unleash their creativity and productivity.
  • AMD RADEON 8090S iGPU GAMING PC - The AMD Radeon RX 8060S offers all 40 CUs with up to 2.9 GHz graphics clock and uses the new RDNA 3.5 architecture. The powerful iGPU is positioned between an RTX 4060 and 4070 laptop GPU and therefore enables gaming in FHD at maximum details in most demanding games. The 8060S can also utilize the full 128GB pool, which is perfect for running LLMs such as Deepseek 70B Q8, which runs comfortably on this machine.
  • EIGHT CHANNEL LPDDR5X - LPDDR5X is a new ground breaking memory small form factor installed on-board. With blazing speeds up to to 8000MT/s, it runs 1.5x faster than the DDR5 SODIMMs; 90% better performance over DDR5 SODIMMs in video conferencing and photo editing; 30% better performance in productivity apps; 12% better performance in digital content workloads.
  • QUAD SCREEN 8K DISPLAY SUPPORT - EVO-X2 AI Mini PC support 4-screen 4K/8K output via HDMI 2.1 (8K@60Hz), DisplayPort 1.4 (4K@60Hz), and dual USB 4 40Gbps Transfer speed (supporting PD3.0/DP1.4/DATA). Ideal for gaming, video editing, and multitasking, it provides expansive and crisp multi-display support.

What NeMo Relay traces contain

Relay’s canonical event format is ATOF (Agent Trajectory Observability Format) 0.1. It represents work with two kinds of events: scopes and marks. A scope has a start and end and represents timed work, such as an agent, tool, or LLM call. Its start and end pair by UUID, while parent UUIDs preserve the relationship between nested work. A mark is a point-in-time checkpoint, not a timed pair. Relay-generated timestamps are used by default. The events documentation defines these semantics.

Depending on configuration and the exported projection, trace data may include prompts, model responses, tool arguments and results, file paths, timing, identifiers, and error information. The formats are not interchangeable: an export may omit information or represent it differently from the underlying event stream.

Choose an output format for the question you need to answer

Format Best suited to What it represents
ATOF JSONL Debugging or auditing individual events Event-level record with timing, IDs, parent-child relationships, scopes, and marks.
ATIF Reviewing or evaluating an agent’s trajectory Step-by-step trajectory assembled from lifecycle events. It omits marks because its model is trajectory steps, not independent checkpoints.
OpenTelemetry, including OpenInference projection Sending telemetry to an OTLP-compatible observability system Spans and related telemetry. In NVIDIA’s tutorial, Phoenix is used to inspect model and tool calls, duration, token use, errors, and available inputs and outputs. The tutorial also names LangSmith as an OTLP-compatible destination.

Relay’s event-format documentation describes event and projection behavior. For a backend view, NVIDIA’s tutorial demonstrates Phoenix. These destinations are optional; an observability backend is not a prerequisite for Relay.

Rank #2
AMD Ryzen™ AI Halo - Personal AI Desktop Computer - Developer Platform - Linux OS
  • Built for Local AI Development: AMD Ryzen AI Halo is designed for local AI development and inference, featuring 128GB unified memory and support for up to 200B parameter models to build and run intensive AI workloads locally.
  • 128GB Unified Memory: Features 128GB LPDDR5x unified memory at 8000 MT/s with 256 GB/s memory bandwidth, providing a shared memory pool across the CPU, GPU, and NPU to support larger AI models.
  • AMD Ryzen AI Max+ 395 Processor: Features 16 cores, 32 threads, and Zen 5 architecture, paired with AMD Radeon 8060S integrated graphics featuring 40 RDNA 3.5 compute units and an AMD XDNA 2 NPU with up to 50 TOPS.
  • Linux AI Developer Platform: Purpose-built for Linux-based AI development with full AMD ROCm software support and preloaded tools, models, and workflows optimized for local AI development.
  • Compact, Connected Design: Includes a 2TB M.2 SSD, 10GbE LAN, Wi-Fi 7, Bluetooth 5.4, USB-C connectivity, and HDMI 2.1b.

How to tell whether a tool call succeeded

A trajectory’s tool request shows what the model asked to run; it does not establish that the tool completed successfully. To inspect the recorded outcome, find the corresponding ATOF tool scope and examine its start and end events, along with any error data. The scope UUID pairs the boundaries, and parent UUIDs show how the tool call fits into its parent work. Then compare the recorded execution with a separate task verifier: a tool can run without achieving the requested result, and a trace alone is not a success check.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What one Relay run can—and cannot—show

In a Hermes Agent tutorial published September 30, 2026, NVIDIA demonstrates a verified terminal-tool run. The runner checks for the exact expected output VALUE=42, confirms completed LLM activity and zero tool errors, and verifies that ATOF and ATIF artifacts exist. NVIDIA reports 74 ATOF events, two completed LLM scopes, 7,239 prompt tokens, 96 completion tokens, 7,335 total tokens, one tool call, zero tool errors, and three ATIF steps for that run. These are results from that tutorial run, not general performance guarantees; NVIDIA notes that token counts, identifiers, and file paths can vary between runs. The tutorial gives the example and its qualifications.

Why compare runs instead of trusting one result

In a larger case study, NVIDIA reports an August 6, 2026 rerun of Hermes ToolPerf across nine tasks, comparing pinned baseline and fixes arms. Each task was run three times per model per arm, for 108 runs total. A task verifier measured completion while Relay ATOF captured model calls, tool calls, errors, retries, result data, and timing. NVIDIA’s tutorial reports these results:

Rank #3
GMKtec EVO-X2 AI Mini PC Ryzen Al Max+ 395 Superchip 128GB LPDDR5X 2TB SSD
  • EVOLUTION RYZEN AI MAX+ 395 MINI PC - GMKtec EVO-X2 is the next evolution in AI mini PC Ryzen Strix Halo series. Thanks to AMD Simultaneous Multithreading (SMT) the core-count is effectively doubled, to 32 threads. Ryzen AI Max+ 395 has 64 MB of L3 cache and can boost up to 5.1 GHz, depending on the workload. The Ryzen AI Max+ 395 is currently rated as the "most powerful x86 APU" on the market for AI computing.
  • AI NPU with XDNA 2 ARCHITECTURE - Powered by 16 “Zen 5” CPU cores, 50+ peak AI TOPS XDNA 2 NPU and a truly massive integrated GPU driven by 40 AMD RDNA 3.5 CUs, the Ryzen AI MAX+ 395 is a transformative upgrade and delivers a significant performance boost over the competition. The Ryzen AI Max+ 395 excels in consumer AI workloads like the llama.cpp-powered application: LM Studio. Shaping up to be the must-have app for client LLM workloads, LM Studio allows users to locally run the latest language model without any technical knowledge required and unleash their creativity and productivity.
  • AMD RADEON 8090S iGPU GAMING PC - The AMD Radeon RX 8060S offers all 40 CUs with up to 2.9 GHz graphics clock and uses the new RDNA 3.5 architecture. The powerful iGPU is positioned between an RTX 4060 and 4070 laptop GPU and therefore enables gaming in FHD at maximum details in most demanding games. The 8060S can also utilize the full 128GB pool, which is perfect for running LLMs such as Deepseek 70B Q8, which runs comfortably on this machine.
  • EIGHT CHANNEL LPDDR5X - LPDDR5X is a new ground breaking memory small form factor installed on-board. With blazing speeds up to to 8000MT/s, it runs 1.5x faster than the DDR5 SODIMMs; 90% better performance over DDR5 SODIMMs in video conferencing and photo editing; 30% better performance in productivity apps; 12% better performance in digital content workloads.
  • QUAD SCREEN 8K DISPLAY SUPPORT - EVO-X2 AI Mini PC support 4-screen 4K/8K output via HDMI 2.1 (8K@60Hz), DisplayPort 1.4 (4K@60Hz), and dual USB 4 40Gbps Transfer speed (supporting PD3.0/DP1.4/DATA). Ideal for gaming, video editing, and multitasking, it provides expansive and crisp multi-display support.
Model and measure Baseline Fixes
Claude Sonnet 4.5: task completion 24/27 tasks (89%) 23/27 tasks (85%)
Claude Sonnet 4.5: mean duration 16 s 22 s
Qwen3 Coder 30B: task completion 19/27 tasks (70%) 22/27 tasks (81%)
Qwen3 Coder 30B: mean LLM calls 3.8 4.9
Qwen3 Coder 30B: mean tool calls 2.8 3.9
Qwen3 Coder 30B: mean tool-result data 16 KB 33 KB
Qwen3 Coder 30B: mean duration 27 s 42 s

In this sample, the fixes changed little for Sonnet, while Qwen completed three more tasks out of 27 and also used more calls, returned more tool-result data, and took longer on average. Task-level audits showed why aggregate success alone was incomplete: recovering from a blocked command improved completion but required more turns; case-insensitive search triggered extra exploratory searches on some repetitions; and a hidden-file search failure remained unresolved. These results illustrate a particular harness, model set, and workload—not a general prediction for other agents.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

A controlled way to evaluate an agent change

NVIDIA recommends treating task completion as the primary outcome and traces as evidence for diagnosing the path behind it. For a useful comparison:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  1. Define an exact automated success check. Specify what output or state proves that the task is complete.
  2. Set a baseline and one focused change. Isolating the change makes it easier to connect an outcome to a prompt, tool, or harness adjustment.
  3. Hold conditions constant. Keep the model snapshot, provider, task input, execution budget, and timeout the same for both arms.
  4. Repeat both arms equally. Compare repeated runs rather than relying on a single favorable result.
  5. Compare verified outcomes first. Then use traces to inspect calls, retries, errors, elapsed time, token use, and cost.
  6. Repeat across intended workloads or models. A result for one model or task set does not establish that a change will help elsewhere.

A faster run or a lower call count on its own does not establish an optimization: it may also fail more often or behave differently on repeated inputs. The verifier answers whether the task passed; the trace helps explain the behavior that produced that result. NVIDIA’s comparison guidance uses this distinction.

Protect trace data before sharing it

Because traces may contain prompts, responses, tool arguments and results, file paths, and other application data, treat exported artifacts as potentially sensitive. Review and sanitize them before sharing, and account for the selected exporter’s behavior: marks remain in ATOF but are omitted from ATIF, while OpenTelemetry projections may represent event data differently. NVIDIA’s tutorial cautions about trace contents, and the events documentation describes the format distinctions.

Which exporter should you use?

  • Use ATOF JSONL when you need the underlying event-level record to investigate timing, IDs, nesting, marks, and recorded outcomes.
  • Use ATIF when a step-by-step trajectory is the useful view for review or evaluation, and independent mark events are not needed.
  • Use OpenTelemetry or an OpenInference projection when you want spans and related telemetry in an OTLP-compatible observability system; check the projection and backend to understand which payloads and event details are retained.

The integration point is a separate choice: use the path that fits ownership of the real work—CLI sidecar, SDK, framework integration, wrapper, or plugin. Before selecting an export or sharing traces, consider whether it retains the marks and payload details you need, what data can be sanitized, and whether the accompanying evaluation uses fixed inputs and task verification rather than trace volume or one-run speed.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More from the Wire

  1. World desk4 min
    How to Spot an AI Voice Scam Before Sending MoneyDon’t rely on how a caller sounds. Pause, call back through a known number, and verify the emergency with another trusted person before sending money.
  2. Mountain View desk4 min
    Google’s SynthID Detector: How to Check AI-Generated Images, Video and AudioGoogle’s SynthID Detector looks for an embedded watermark in supported images, video and audio. Here is what its results do—and do not—show.
  3. Redmond desk20 min
    How to create a link to File or Folder in Windows 11Windows 11 gives you several ways to point to a file or folder without moving or duplicating it. You can create a desktop shortcut,…
Recommended PC Tool
Recommended PC Tool
Outdated Drivers Are Slowing You DownFree scan - exact matches
PC Slower Than It Used to Be?Free scan - under a minute

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.