The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Use agentevals to score recorded OpenTelemetry traces from kagent agents against a version-controlled golden eval set, then make agreed score thresholds part of CI. This catches changes in recorded tool use or responses; it does not rerun the agent or, by itself, prove that the agent is generally correct.
What agentevals checks—and what it does not
agentevals is a framework-agnostic tool for scoring agent behavior captured in OpenTelemetry traces. It can compare recorded behavior with a golden eval set, run custom evaluators, and apply thresholds for CI/CD. It supports Jaeger JSON and native OTLP trace formats, and can evaluate an existing trace without re-executing its LLM calls.
That makes agentevals useful for regression checks on captured behavior, not a substitute for executing a newly built agent. To test a new kagent version end to end, your pipeline needs to run the agent and capture fresh traces before scoring them. The project describes itself as under active development, so pin the release you use and confirm CLI options and evaluator semantics against that release.
kagent is a Kubernetes-native agent platform. Its project documentation describes testing through public APIs and using task history and traces to diagnose failures. The kagent 1.x overview and observability documentation describe OpenTelemetry traces and structured logs across kagent and Agent Substrate.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Capture traces you can actually evaluate
Start with tasks that matter to users. Include representative success cases, important branches and tool calls, and known failure cases. Generate runs using the kagent version and configuration your regression suite is intended to cover. Make sure tracing is enabled and that your evaluation setup retains those runs.
The kagent 1.x OpenTelemetry stack guide describes an OpenTelemetry Collector and trace backends including Tempo. It says Agent Substrate keeps 1% of traces by default, so a small number of test requests may produce no visible trace. For an evaluation setup, the guide shows otel.traces.samplingRatio=1.0; at that ratio, the router records every forwarded request. The guide cautions that the ratio should be lowered again for production. Treat this as version-specific guidance from the kagent 1.x documentation, not a universal default for every release.
Trace contents can include prompts, tool inputs, and outputs. Decide where those data may be stored, who can access them, and what should be redacted under your organization’s policies; the cited technical documentation does not prescribe a universal retention or redaction policy.
Build a golden eval set around intended behavior
An eval set records reference data that evaluators use for comparison. The agentevals eval-set format documentation says the format follows Google ADK’s EvalSet schema and supports version-controlled test suites. It also describes generating eval sets from golden sessions in the UI.
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Clear out junk files and repair common Windows errorsFree Scan →Begin with a small set of high-value examples, then add cases when incidents, agent changes, or new task variants expose gaps. Make each expectation specific to the failure you want to catch:
- For tool selection, record the expected tool uses or trajectory.
- For user-facing output, include an expected final response or task-specific response criteria.
- For business requirements, consider a custom evaluator that checks the relevant rule.
Keep the references aligned with current product requirements. When a behavior change is intentional, review and update the golden expectation alongside the agent change; otherwise a valid update may be flagged, or an obsolete baseline may reward behavior you no longer want.
Choose evaluators for the behavior at risk
The agentevals README demonstrates tool_trajectory_avg_score against a golden eval set: a trace that calls the expected Helm listing tool passes the example, while one with no matching tool call fails. It also demonstrates response_match_score for comparing a final answer with an expected response. The eval-set guide lists additional choices, including LLM-judge and safety or hallucination evaluators. Check the installed release’s documentation for current names and semantics.
| Evaluation target | What it can help detect | What it cannot establish alone |
|---|---|---|
| Tool trajectory | A changed or missing expected tool-use pattern. | That the final answer was useful, correct, or safe. |
| Response matching | A difference between the final response and an expected answer. | That a similar response is factually sound; valid paraphrases may also score poorly. |
| Safety, hallucination, or custom rules | The specific conditions implemented by the selected evaluator. | General agent quality or correctness beyond those conditions. |
Use a metric that matches the risk, and inspect examples around failures. A deterministic check may be easier to reproduce, while model-based judgments or live agent executions can vary. For important tasks, combine appropriate automated checks with response review or domain-specific criteria rather than treating one score as a complete quality measure.
Run the same checks in CI
The project documents a CLI command in this form for scoring a trace against an eval set:
Rank #4
agentevals run samples/helm.json
--eval-set samples/eval_set_helm.json
-m tool_trajectory_avg_score
It also documents multiple trace inputs, JSON output, and evaluator thresholds in configuration. A practical regression job should keep the eval set and evaluator configuration under version control, pin the agentevals version, supply or capture the trace inputs in a controlled step, and fail on a threshold chosen for your task. The documentation establishes CLI and quality-gating capability, but does not prescribe a particular CI provider or a universal pipeline recipe.
- Check out the agent code, eval set, and evaluator configuration for the change being tested.
- Produce traces. Either score representative traces already captured for the target behavior or run the agent and capture fresh traces if the purpose is to test the newly built version end to end.
- Run agentevals with the selected evaluator or evaluators against those traces and the checked-in expectations.
- Apply your threshold and fail the job when a score falls outside the agreed limit; review the underlying traces rather than assuming every failure is a real regression.
For specialized checks, agentevals documents custom evaluators using a stdin/stdout JSON protocol. They can be written in Python, JavaScript/TypeScript, or another language that can read and write JSON. The custom-evaluator guide shows threshold configuration and an illustrative sample threshold; set your own value based on task requirements and observed behavior rather than copying the example.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Compare evaluation approaches before choosing a gate
Recorded-trace scoring is one part of an evaluation strategy. Compare approaches by the evidence they use and the behavior they measure, rather than assuming that any single score represents overall agent quality.
Recommended Free Tools
Best Value
| Decision axis | Questions to answer |
|---|---|
| Evidence | Are you scoring recorded traces, or rerunning the agent for each test? Recorded traces avoid repeating LLM calls but do not test a new build unless it is run and captured. |
| Behavior dimension | Do you need to check tool trajectory, final responses, safety or hallucination, or task-specific business rules? |
| Reproducibility | Can a deterministic check cover the risk, or does it require model-based judgment or live calls that may vary? |
| Integration effort | Can you import existing traces, or must you wire up OpenTelemetry collection, custom evaluators, and CI steps? |
| Operations | Is local trace inspection enough, or does the team need shared telemetry storage, retention controls, and access management? |
The official project documentation describes agentevals’ capabilities; it does not provide a neutral benchmark comparing it with competing evaluation products.
Triage failures without weakening the baseline
When a CI gate fails, inspect the trace and decide what changed before editing the test. A failure may reflect a genuine regression, a desired behavior update, a flawed fixture, or missing instrumentation. If the new behavior is intended, update the eval set as part of the reviewed agent change and preserve the rationale for changing expectations. If a trace is absent, check sampling and instrumentation before concluding that the agent did not run.
Scores depend on trace quality, eval-set coverage, evaluator behavior, and thresholds. They are evidence for review, not statistically calibrated significance tests or a general guarantee of agent correctness.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




