To evaluate an agent’s tool use, measure two separate things. The first is whether each tool call is appropriate and correctly formed. The second is whether the whole workflow ends in a verified goal state. A well-formed call doesn’t show that the task was done. An agent can pick the right function, pass valid arguments, and still leave the database, ticket, or calendar in the wrong state.
No single benchmark covers both levels, plus user interaction, policy rules, and unreliable tools. This guide explains what each major benchmark measures, how to build a test set from those ideas, and which numbers to report. Every benchmark score below is tied to its paper, model, metric, and year, because none of them generalizes beyond its own setup.
The two levels of tool-use evaluation
The Berkeley Function Calling Leaderboard (BFCL) paper, published in the Proceedings of Machine Learning Research in 2025 by Shishir G. Patil and coauthors, defines the capability this way: “Function calling, also called tool use, refers to an LLM’s ability to invoke external functions, APIs, or user-defined tools in response to user queries—an essential capability for agentic LLM applications.”
That definition covers the call. Evaluating the call and evaluating the outcome are different jobs:
Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Repair Windows errors before they cause bigger problemsFix Now →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →#1 Best Overall
- Call-level (process) evaluation asks whether the agent chose the right tool, filled in the arguments correctly, issued the right number of calls, and declined to call anything when no tool fit. It is cheap, deterministic, and good for debugging.
- Task-level (outcome) evaluation asks whether the end state of the system matches what the task required. It needs an executable environment or a faithful simulation, and it is the right basis for release decisions on workflows that change state.
NVIDIA’s September 2026 article on evaluating tool-calling agents makes a similar split. It is practitioner guidance, not a standard from a standards body. Treat it as a useful checklist rather than a specification.
What the main benchmarks actually measure
BFCL: call form across languages, extended to multi-step settings
BFCL tests serial calls (one after another) and parallel calls (several in one turn) across multiple programming languages. It scores calls with abstract syntax tree (AST) matching, which compares the structure of a generated call to acceptable answers instead of comparing raw strings. The 2025 paper also extends the scope to abstention (not calling when it shouldn’t) and to stateful multi-step agent settings.
The authors conclude that single-turn calling is comparatively strong. Memory, dynamic decision-making, and long-horizon reasoning remain open challenges. For you, that means a strong single-call score is the least informative number for an agent that works over many steps.
τ-bench: conversations, policies, and final database state
τ-bench (2024) simulates conversations between a user and an agent that operates domain APIs under policy constraints. After the conversation, it compares the final database state with an annotated goal state. That checks outcomes directly and doesn’t depend on how the agent got there.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Rank #2
It also proposes pass^k, which captures how reliably an agent succeeds across repeated attempts at the same task. Pass^k asks for success on all k independent trials. It differs from pass@k, which asks for at least one success. In the paper’s experiments, state-of-the-art function-calling agents succeeded on fewer than half of tasks, and retail pass^8 was below 25%. Those figures belong to that paper’s models, tasks, and definitions. They aren’t universal failure rates.
AppWorld-UL: clarification, confirmation, and infeasible requests
AppWorld-UL (2026) adds the user relationship. It reports 516 user-in-the-loop tasks built on nine simulated apps. Some tasks require the agent to ask a clarifying question, get confirmation before acting, or say that an instruction can’t be carried out.
The paper reports these results for Claude Opus 4.7:
- 48.6% success overall;
- 35.7% on the harder compositional subset;
- 21.3% on that compositional subset under a stricter scenario-level metric.
The drop from 48.6% to 21.3% shows how much the metric definition matters. The same model on the same benchmark looks very different depending on the scoring rule. Quote any of these figures with the benchmark, model, metric, and year attached.
ToolBench-X: tools that fail
ToolBench-X, a 2026 preprint, tests whether agents cope with an unreliable tool environment. It defines five hazard types:
- Specification drift: the tool’s documented interface no longer matches its behavior;
- Invocation error: the call is rejected;
- Execution failure: the call is accepted but the tool fails;
- Output drift: the returned data changes shape or meaning;
- Cross-source conflict: two sources disagree.
The hazards are recoverable, and tasks include recovery paths such as retrying, falling back to another tool, verifying a result, and cross-checking sources. Because it is a recent preprint, treat it as new evidence rather than settled consensus. The hazard taxonomy is still useful for designing your own fault-injection tests.
Side-by-side comparison
| Benchmark | Primary question | Horizon / state | Verification | Distinctive element |
|---|---|---|---|---|
| BFCL (PMLR, 2025) | Is the call well formed and appropriate? | Mostly single-turn, extended to stateful multi-step | AST matching for call form | Serial, parallel, multi-language, abstention |
| τ-bench (2024) | Does the conversation end in the right database state under policy? | Multi-turn, stateful, simulated user | Final state vs. annotated goal state | pass^k repeatability |
| AppWorld-UL (2026) | Does the agent handle clarification, confirmation, and infeasibility? | 516 tasks, nine simulated apps | Task and stricter scenario-level metrics | User-in-the-loop interaction |
| ToolBench-X (2026 preprint) | Can the agent diagnose and recover from tool faults? | Tasks with injected hazards | Not stated here | Five hazard types and recovery paths |
How to build your own tool-use evaluation
1. Define success as a state change
Before writing any test, write down what proves a task is done. For a workflow that changes a system, that is a specific record, field value, or message. A transcript that says “Done” isn’t proof. As an illustration, for a refund agent, success could be “the order shows status refunded, the refund amount equals the eligible amount, and no other order was modified.”
2. Assemble a representative test set
Include these categories rather than only happy-path tasks:
- representative tasks drawn from real usage;
- edge cases such as empty results, boundary values, and unusual formats;
- ambiguous requests where the right move is to ask;
- policy-constrained requests where the right move is to refuse or escalate;
- requests no available tool can satisfy, where the right move is to abstain or say so;
- failure and recovery conditions, such as timeouts, malformed outputs, and conflicting sources.
3. Prefer deterministic checks
Use code, not a model judge, wherever the answer can be computed. That includes tool selection, argument values, policy adherence, number of calls, and final state. Where a result truly requires judgment, such as the quality of a clarifying question, write down the rubric and record the judge’s limitations. A judge model’s scores are an estimate, not ground truth.
4. Run call-level diagnostics separately
Score tool selection and argument correctness as separate numbers. A single “tool-call accuracy” figure hides whether the agent picks the wrong tool or picks the right one and fills it in badly, and those two failures need different fixes (tool descriptions and routing versus schemas and argument grounding).
5. Run end-to-end tasks in an executable environment
Use a sandbox with real or faithfully simulated tools and resettable state. Reset state before every trial. Otherwise one run’s side effects contaminate the next and your pass rate means nothing.
6. Repeat trials and report variation
Agents are stochastic, so one run per task overstates reliability. Run each task several times and report how often it succeeds every time, in the spirit of τ-bench’s pass^k, along with the spread across trials. A system that passes a task 6 times out of 8 is a different product from one that passes 8 out of 8, even though both have a nonzero pass rate.
Free tools Windows power users keep installed
One-click scans. No signup required.
Best Value
7. Inject faults on purpose
If your deployment depends on external services, test the ToolBench-X hazard categories against your own tools: change a schema, return an error, return stale or reshaped output, make two sources disagree. Then check whether the agent retries sensibly, falls back, verifies, or cross-checks, and whether it stops when it should instead of inventing a result.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.What to report
| Metric | What it tells you | Use it for |
|---|---|---|
| Task success rate | Share of tasks reaching the verified goal state | Release decisions |
| Variation across independent trials (e.g. pass^k) | Whether success is repeatable | Release decisions |
| Tool-call precision / tool selection accuracy | Whether the right tools are called, and no extras | Debugging |
| Argument accuracy | Whether values and formats are correct | Debugging |
| Steps per successful task | Efficiency of the path | Latency and cost tuning |
| Cost per successful task | Spend divided by successes, not by attempts | Budgeting and model comparison |
Keep process metrics for debugging and outcome metrics for ship or no-ship calls. Cost per successful task matters because a cheap agent that fails often can cost more per useful result than an expensive agent that rarely fails.
Choosing a benchmark for your deployment
Ask seven questions of any benchmark before trusting its score for your use case:
- Is it a single call or a multi-step horizon?
- Is the environment stateless, or does it change state?
- Does it include a simulated user and clarification behavior?
- Are tools actually executed, or are calls only scored for form?
- Is verification deterministic (final state) or reference- or judge-based?
- Does it represent policy, safety, and recovery hazards?
- What are its repeatability, runtime, and cost?
Then match the answers to the action consequences in your product. A read-only search assistant mostly needs call-level accuracy and abstention. An agent that issues refunds, edits records, or sends messages needs state-verified outcomes, policy tests, and confirmation behavior. An agent that depends on flaky third-party APIs needs fault injection. In every case, add internal tests for behavior no public benchmark covers.
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Scan for outdated or missing drivers - takes under a minute3Clear out junk files and repair common Windows errorsCommon mistakes
- Comparing scores across benchmarks. Different task horizons, statefulness, user simulation, and scoring rules make the numbers non-interchangeable. The AppWorld-UL drop from 35.7% to 21.3% under a stricter metric is an example inside a single benchmark.
- Treating a valid call as a finished task. Check the state, not the call log.
- Reporting a single run. Without repeated trials you can’t distinguish a reliable agent from a lucky one.
- Testing only clean tools. Production APIs drift and fail; an agent that has never seen a failure hasn’t been evaluated for production.
- Quoting a headline number without its context. Name the benchmark, model, metric, and year every time.
The evidence base is still moving. AppWorld-UL and ToolBench-X are 2026 work, and a broad survey literature on agent evaluation exists with its own taxonomy of objectives and processes. No one benchmark or survey settles which dimensions matter most for your system, so treat public scores as one instrument among several and rely on your own state-verified tests.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




