🎙️ Daily Podcast (FR) : NViNiO•Podcast™
ADs | ✨ Enhance your Social Media content with NViNiO•AI™ for FREE
An agent can finish a task and still take an inefficient path. A failed search can trigger another search. A truncated file read can lead to a command fetching the same content again. A correct final answer hides those extra steps, even though they increase latency and consume tokens. Inefficiencies create more chances for failure.
To improve an agent’s behavior, developers must understand whether a task succeeded and how the agent completed it. A success check by itself cannot explain why an agent recovered from a tool error, stopped early, or needed extra model calls.
In this tutorial, you’ll run two Hermes Agent examples with NVIDIA NeMo Relay. You’ll use the resulting traces to inspect model and tool calls, errors, retries, duration, and token use, then compare that evidence with each task’s verification result. A Hermes ToolPerf case study shows how to use the same approach to evaluate harness changes across repeated runs.
This tutorial and video explain how to:
Set up an isolated Hermes Agent runtime with its native NeMo Relay integration. Run a simple terminal tool task and inspect its event stream and trajectory. Run a file-and-web research task and explore its OpenTelemetry trace in Arize Phoenix. Combine task verification with trace evidence to evaluate a change to an agent harness.Video 1. A step-by-step code walkthrough for evaluating Hermes Agent traces with NeMo Relay
Prerequisites
Before you begin, make sure you have:
macOS or Linux Git and curl Docker Desktop or Docker Engine installed and running An NVIDIA Build API key for NVIDIA Nemotron 3.5 Lightning. Open the model page and select Generate API Key.Next is an overview of the technologies used for this tutorial and how they work together.
How NeMo Relay works with Hermes Agent harness
NeMo Relay gives agent developers a common way to observe and control model and tool execution. The popular agent harness Hermes Agent includes NeMo Relay natively and represents its sessions, turns, model calls, and tool calls in NeMo Relay’s scope hierarchy. NeMo Relay records lifecycle events as work begins and ends, preserving its timing and parent-child relationships.
Understand the trace outputs
NeMo Relay is used for agent observability. You will work with three representations of agent execution:
| Agent Trajectory Observability Format (ATOF) | A JSONL log of scope starts, scope ends, and point-in-time mark, with IDs and timestamps to reconstruct the agent run. | Use ATOF to debug or audit individual events, timing, and parent-child relationships. |
| Agent Trajectory Interchange Format (ATIF) | A step-by-step JSON record of agent interactions, tool calls, and observations, assembled from lifecycle events. | Use ATIF to review, analyze, or evaluate the agent’s path step by step. |
| OpenTelemetry with OpenInference | OpenTelemetry records the run as parent-child spans. OpenInference labels agent, LLM, and tool spans and defines their attributes. | Use it in OTEL-compliant tools like Phoenix, to inspect model and tool calls, duration, token use, and errors. |
An ATIF tool request shows what the model asked to run, but it does not confirm the outcome. To verify what happened, inspect ATOF for the matching tool start and end events and any recorded errors. Their shared uuid pairs the events, while parent_uuid connects the tool call to its parent.
Review traces before sharing them. Depending on your configuration, they can contain prompts, model responses, tool arguments and results, file paths, and other application data.
For agent safety and security governance, NeMo Relay helps provide the evidence layer: structured traces and trajectories that enterprises, evaluators, and security systems can use to investigate agent behavior, evaluate policies, improve controls, or create specialized security plugins that extend Relay.
Let’s get started with the first agent task run.
The first example is intentionally small so you can verify the complete setup before adding web search and Phoenix. Hermes uses its terminal tool to run the included Python script inside an isolated Docker container. The script prints: VALUE=42.
That fixed output gives the runner an exact success check. A passing run also confirms that Hermes reached the model, invoked the terminal tool in the sandbox, and produced both Relay trace files.
The container cannot access the network, repository checkout, or NVIDIA API key. Hermes also cannot fall back to running terminal commands on the host.
Run the following commands in order. After copying keys.env, add your NVIDIA API key to that file before continuing.
The Hermes Agent and NeMo Relay processes run from this repository’s local environment. Docker is used separately for the terminal tool sandbox and the local Phoenix service.
The setup script in the repository creates a self-contained environment under .tutorial-runtime/ with all the appropriate dependencies like Python 3.11 and Hermes 0.21.1 with NeMo Relay 0.8.3. It does not modify your existing Python or Hermes installation.
When the task finishes, the runner checks the response and both trace files. A passing run prints the verification result, followed by the ATOF and ATIF summaries.
Review the trace summaries
After Hermes completes the task, the runner verifies the expected result and the generated traces. It checks that the terminal command succeeded, that the ATOF trace contains completed LLM activity with token usage and no tool errors, and that a nonempty ATIF trajectory was created. The following output comes from one verified run. Token counts, identifiers, and file paths can vary between runs.
ATOF summary
ATIF summary
Look for the line with Task verified: VALUE=42, which confirms the expected result. The ATOF summary reports the completed model scopes, token usage, tool calls, and tool errors. The ATIF summary presents the same execution as a three-step trajectory.
The path following Artifacts: identifies the run directory containing the complete ATOF event stream and ATIF trajectory. The companion repository explains how to inspect or summarize either file again.
Experiment #2: Run a multi-tool research task and explore its traces
After verifying the basic setup, this second example uses the same Hermes and NeMo Relay environment for a task that requires several tools. Hermes receives a travel record containing clues about an unnamed machine-learning conference. It must read the record, find a conference matching the subject, dates, and location, confirm the answer on the official website, save the verified information in a report, and return the conference name.
Use NeMo Relay’s OpenInference exporter to send OpenTelemetry spans to Arize Phoenix over OTLP (OpenTelemetry Protocol). Phoenix displays the run as an interactive trace, where you can inspect model and tool calls, timing, token usage, errors, and available inputs and outputs.
You can send the same OpenTelemetry trace to other OTLP-compatible backends, such as LangSmith, by changing the endpoint and authentication settings. The NeMo Relay observability guide describes other available exporters and configuration options.
The example reuses the NVIDIA Nemotron model and NVIDIA API key from the first example. Hermes uses its built-in keyless web search, and the runner starts Phoenix in a pinned local container. Run the conference search example:
During the run, NeMo Relay saves the ATOF event stream and ATIF trajectory locally.
Before reporting success, the runner checks that Hermes:
Identified COLT 2026 as a conference name Saved a report containing the expected conference details and official source Successfully completed the read_file, web_search, web_extract, and write_file calls Produced a nonempty ATIF trajectory Sent model and tool spans with positive token usage to PhoenixAfter the checks pass, the terminal output includes a link to the Phoenix project and the local run directory. Open the project in Phoenix to follow the agent run from the initial file read through the web search, source verification, report write, and final response.
Figure 1. Phoenix shows the complete Hermes Agent run on the left and the selected model call on the right, including the conference query, requested read_file call, timing, and token usage
Try the task with another model
To see how another model handles the same conference query, follow the model-profile instructions in the companion repository. Keep the query, available tools, execution limits, and verifier unchanged so you can inspect how the execution path differs.
Because the task uses live web search, use these runs to explore behavior rather than rank models. A controlled comparison requires fixed search responses and repeated runs.
Use traces to evaluate an agent harness change
The two examples show how to verify a result and inspect one run. Evaluating a harness change requires the same checks under controlled, repeated conditions.
Steps for agent harness evaluation
Choose a fixed task with an exact, automated success check. Define a baseline and one focused change to the prompt, tool, configuration, or harness. Keep everything except that change constant, including the model snapshot, provider, task input, execution budget, and timeout. Run the same number of repetitions for the baseline and candidate with NeMo Relay enabled. Compare verified task outcomes first. Then use the traces to examine model calls, tool calls, retries, errors, elapsed time, token usage, and cost. Repeat the evaluation across the models or workloads that the change is expected to support before generalizing the result.Call the candidate an improvement only when it produces a repeatable increase in task completion or preserves completion while improving the reliability, latency, or cost measure you intended to change. One faster run or fewer calls can help explain a result, but neither establishes an optimization by itself.
An unchanged result is useful as well. It can show that an apparent improvement depended on a particular model, environment, or sample.
The Hermes ToolPerf benchmark was developed by Nous Research after analyzing production sessions, auditing tool schemas, and mining production session logs for failure classes. Then, NeMo Relay ATOF traces were used for ground-truth turn accounting. By examining nine failure patterns, the results were turned into deterministic benchmark cases and then used to evaluate a batch of Hermes tool-layer fixes.
The August 6 rerun compared the pinned baseline of Hermes and fixed revisions across nine tasks. Each task ran three times per model per arm, giving 108 runs total, with the same prompts, tools, execution limits, and success checks throughout. A task verifier measured completion, and the NeMo Relay ATOF traces captured model calls, tool calls, errors, retries, tool-result data, and timing.
| Claude Sonnet 4.5 | Baseline | 24/27 (89%) | 2.9 | 2.2 | 17 KB | 16 s |
| Claude Sonnet 4.5 | Fixes | 23/27 (85%) | 2.8 | 2.1 | 17 KB | 22 s |
| Qwen3 Coder 30B | Baseline | 19/27 (70%) | 3.8 | 2.8 | 16 KB | 27 s |
| Qwen3 Coder 30B | Fixes | 22/27 (81%) | 4.9 | 3.9 | 33 KB | 42 s |
In this run, Sonnet showed no real change. It finished 24 of 27 runs on baseline and 23 of 27 on fixes, a difference of one run. Its turns, tool calls, and result data were the same on both arms.
Qwen Coder is where the fixes did something, and the effect cuts two ways. It completed three more tasks successfully, going from 19 of 27 to 22 of 27. It also worked harder for them: mean LLM calls rose from 3.8 to 4.9, tool calls from 2.8 to 3.9, tool-result data from 16 KB to 33 KB, and duration from 27 s to 42 s. The fixes brought success to tasks the baseline abandoned, and the price was a slower, chattier agent.
The task-level audit shows where those trade-offs come from.
On the blocked-command task, baseline Qwen died at the parser block and scored 33%, while the recovery recipes took the fixes arm to 100%. That completion is the reason the turn count went up: recovering costs turns that giving up never spends. The case-insensitive search task moved the other way. The zero-match probe output pushed Qwen into extra exploratory searches on two of three reps, taking it from 3.3 turns to 9.3, a regression worth its own look. The hidden-file search stayed at 0 to 33% on both arms and both models, the same gap the original run found and still open at these SHAs.The success or failure of tasks alone could not have produced this conclusion. The NeMo Relay traces recorded every model call, tool call, error, retry, result payload, and timing for all 108 runs. Qwen’s extra turns were recoveries rather than flailing, and that one task’s regression traced back to a specific probe output. The traces are checked into the results directory, so anyone can unpack them and regenerate these tables from the raw records.
Get started with evaluating agent traces
NeMo Relay gives you a consistent way to capture evidence for evaluating changes to optimize your agent harnesses. ATOF preserves the ordered lifecycle events, while ATIF presents the same work as a readable trajectory. Pairing these traces with a deterministic verifier enables you to compare harness changes without confusing fewer calls with better results.
The companion repository provides the sample artifacts, including the runnable tasks, verifiers, NeMo Relay configuration, and Phoenix setup used in this post.
Learn more:
NeMo Relay documentation for installation, concepts, and configuration. Hermes Agent documentation for installing and using Hermes. Companion tutorial repository for the runnable examples from this post. ATOF and ATIF documentation for the event and trajectory formats. NeMo Relay observability configuration for configuring exporters and telemetry destinations. Supported integrations for using NeMo Relay with other agent frameworks and applications..png)
3 hours ago
English (United States) ·
French (France) ·