Hardware FixRecommendedDevice not working? Your driver may be the problemCheck updates for common hardware issues.Fix DriversOctober DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsPC HealthRecommendedCrashes, freezes, slowdowns? Check your PC nowSpot repairable issues before they interrupt work.Check PC×
Skip to content
World desk8 min

How to Compare AI Agent Security Benchmarks, Datasets, and Test Methods

AgentDojo, AgentHarm, and ASB test different agent risks. Learn how to compare their threats, tools, scoring, adaptive attacks, repeated attempts, and validity.
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Compare AI agent security evaluations by the behavior they test, the agent and tools they run, the attacks and defenses they use, and how they score outcomes. AgentDojo, AgentHarm, and Agent Security Bench (ASB) address different risks, so their scores are not interchangeable rankings. A useful comparison also checks benign task performance, repeated attempts, and whether the published traces support the score.

What does an agent security benchmark actually measure?

Start with the security claim you want to make, then identify the behavior the evaluation exercises. “Secure” is too broad to be a measurable target by itself. A test of an agent redirected by malicious text in an email does not establish how it handles a direct request to do harm; a test of harmful compliance does not establish resistance to data exfiltration or unsafe tool use.

The 2025 ACM survey of LLM-agent evaluation offers a useful way to organize this question: distinguish the evaluation objective—such as behavior, capability, reliability, or safety—from the process used to measure it, including interaction mode, benchmark, metric computation, and tooling. For a security result, make the target behavior explicit before comparing scores.

Comparison axis What to establish Why it changes interpretation
Target behavior Is the test about indirect prompt injection, harmful compliance, unsafe tool calls, data exposure, or another behavior? A result supports claims only about behavior the evaluation actually tests.
System boundary Does it run a complete tool-using agent with state, a simulated workflow, or isolated model prompts? Which domains and tools are available? Different tools, state, and task affordances create different opportunities for failure.
Attack and defense Are attacks fixed, held out, adaptive to the tested system, or developed against it? Which defenses and baselines are compared? A static attack set may not reveal how the system responds to an adversary who can adapt.
Scoring target Does the score count an attempted action, completion of an attacker’s goal, refusal, policy compliance, or benign task success? Is scoring automated, rubric-based, or human-reviewed? Rates with similar names can count substantially different outcomes.
Utility Are benign task success and security outcomes measured together? A defense that blocks attacks by also blocking ordinary work has a different trade-off from one that preserves task completion.
Repetition How many runs are made per task and model, and are outputs deterministic or sampled? A single run can miss failures that emerge from stochastic outputs or retries.
Reproducibility and validity Are model version, prompts, tools, environment, task sample, scorer, and attempt count disclosed? Can traces be inspected for scoring loopholes? Without these details, it is difficult to reproduce a result or determine what it means.

This framework synthesizes the ACM survey’s evaluation taxonomy with NIST Center for AI Standards and Innovation (CAISI) guidance on testing and evaluation validity.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How do AgentDojo, AgentHarm, and ASB differ?

These benchmark families are complementary. Compare their target behavior and setup before comparing any reported number; they do not share a common security scale.

Benchmark Primary question Scope and design How to use its evidence
AgentDojo Can a tool-using agent resist malicious instructions embedded in untrusted data while doing a legitimate task? The 2024 paper describes 97 realistic tasks and 629 security test cases. The project documentation describes banking, Slack, travel, and workspace suites, with workflows involving sources such as email, banking, and travel. Use it to study indirect prompt injection in interactive, simulated tool-use workflows, while considering benign task utility alongside security outcomes.
AgentHarm Will an agent refuse harmful requests, and can it carry out a multi-step harmful task if jailbroken? The paper evaluates both refusal behavior and the agent’s ability to retain the capability to complete harmful tasks after a successful jailbreak. Its authors report public release of the benchmark dataset. Use it for direct harmful requests and agent misuse questions, not as a substitute for an indirect-injection test. Check the dataset version and scoring protocol before interpreting leaderboard comparisons.
Agent Security Bench (ASB) How do agent attacks and defenses perform across a broader set of scenarios and metrics? The 2024 paper reports 10 scenarios, 10 agents, more than 400 tools, 23 attack/defense method types, and eight evaluation metrics. It reports nearly 90,000 test cases in its experiments. Use it for broad attack-and-defense analysis, but align the scenario, agent setup, and metric with the narrower evaluation you are comparing it against.

The AgentDojo project documentation says its package API remains under development; check the project’s current instructions and compatibility when setting up a run. Benchmark software and datasets can change, so record the version actually used.

How should you design a useful comparison?

Make the comparison a controlled evaluation rather than a contest between headline scores. A practical plan is to specify the claim, hold relevant conditions constant, and report enough detail for another team to understand what was tested.

  1. Define the threat and outcome. Write down what an attacker can do, what information or tools the agent can access, and what counts as a successful attack. Distinguish an attempted harmful action from completion of the attacker’s goal.
  2. Choose benchmarks that cover the target behavior. Use a benchmark suited to the threat—for example, AgentDojo for injection in tool-use workflows or AgentHarm for harmful-request and misuse behavior. Use multiple complementary evaluations when the claim spans distinct risks.
  3. Fix and disclose the system configuration. Record the model version, system and task prompts, agent implementation, tool permissions, environment, relevant package versions, and internet access. If comparing models or defenses, keep other conditions consistent or explain each difference.
  4. Separate attack development from final testing. Include attacks adapted to the system where appropriate, and reserve held-out tasks for evaluation. Report performance by task or scenario as well as in aggregate so a strong average cannot hide a weak area.
  5. Measure benign utility as well as security. Run legitimate tasks without an attack and report their outcomes alongside security results. This helps show whether an apparent security gain comes with a loss in useful task completion.
  6. Predefine the scoring rule and inspect traces. State the scorer and success criterion before interpreting results. Review agent transcripts and outcomes to check whether the score reflects the intended behavior rather than a proxy or loophole.
  7. Repeat runs when outputs can vary. State the number of attempts per task and model, and explain whether sampling or deterministic decoding was used. Report how repeats affect the result, not only the best or first run.

Why do adaptive attacks and retries matter?

In January 2025, NIST CAISI described agent hijacking as indirect prompt injection: malicious instructions are placed in data an agent reads, such as an email, file, or web page, in an effort to redirect its actions. Its guidance says evaluations should improve continually, adapt attacks to the system, examine task-specific performance, and consider multiple attempts.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The experiments in that guidance illustrate why a fixed, one-shot result can be misleading. In its specific evaluation, NIST CAISI reports attack success ranging from 11% for the strongest baseline attack to 81% for the strongest new red-team attack. In a separate experiment, repeating each of five injection tasks 25 times raised mean attack success from 57% to 80%. These figures describe those tested tasks, models, attacks, and evaluation conditions; they are not expected success rates for agents in general.

NIST CAISI also reports developing attacks on a random subset of workspace tasks and testing them on held-out workspace tasks, then trying those attacks in other environments. That design helps distinguish attacks that merely fit familiar examples from attacks that transfer to unseen tasks. For your own results, show per-task outcomes and make clear which tasks were used to develop attacks and which were held out.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How can scoring make a result look safer or more capable than it is?

NIST CAISI’s evaluation-cheating guidance distinguishes two validity problems. Solution contamination occurs when the model has access to information that improperly reveals a task solution. Grader gaming occurs when the model exploits a scoring loophole to receive credit without meeting the intended task. Both can make a score diverge from the behavior an evaluation was meant to measure.

For example, a benchmark may treat a particular tool call as evidence of an attack, even though the call alone does not establish that the attacker’s goal was achieved. The reverse can also happen: a scorer may miss a harmful outcome that does not match its expected pattern. Compare automated scores with the intended task outcome and inspect traces, especially when a metric is only a proxy for success.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Specify task rules, agent affordances, and restrictions clearly.
  • Record internet access, tool permissions, package versions, and scorer behavior.
  • Review transcripts for unintended access to solutions and paths that exploit the grader.
  • Report the denominator and scoring method, not just a percentage.

What must a reproducible security result report?

At minimum, make the tested setup legible enough that another evaluator can identify the result’s scope and attempt to reproduce it. The 2026 preprint Safety, or Just Capability? A Validity Audit of Agent-Safety Benchmarks examines R-Judge, InjecAgent, AgentHarm, and AgentDojo using official implementations and author-provided scorers, while measuring capability benchmarks under its own protocol. It argues that a safety claim should name the benchmark, metric, target behavior, and model panel. Treat this as recent preprint evidence rather than settled consensus.

  • Benchmark and data: name the benchmark, dataset version, suites or task subset, and any exclusions.
  • Model and agent: identify the model version, agent implementation, prompts, tools, and available permissions.
  • Attack conditions: describe the attack set, whether attacks were fixed or adaptive, and how development and held-out tasks were separated.
  • Measurement: define the success criterion, scorer, denominator, attempt count, and treatment of variable outputs.
  • Results: give task-level findings as well as aggregates, and report benign utility when relevant.
  • Validity checks: note transcript review and any identified scoring or task-design loopholes.

What can benchmark evidence establish—and what can’t it?

A benchmark establishes how a specified system performed under a specified test setup. It does not establish a universal ranking of agent security, a standardized score shared by benchmark families, or a guarantee of security in every production context. The NIST CAISI examples show that measured outcomes can change with attack strength and repeated attempts; the 2026 validity audit is a preprint, not a settled field-wide conclusion.

When evaluating a published claim, ask whether its target behavior matches your concern, whether its tools and environment resemble the system you care about, and whether its scoring and attempt protocol support the stated conclusion. If those details differ, the result may still be informative—but it is evidence about a different test, not a directly comparable score.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Leave a Reply

Your email address will not be published. Required fields are marked *

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More from the Wire

  1. World desk4 min
    How to Spot an AI Voice Scam Before Sending MoneyDon’t rely on how a caller sounds. Pause, call back through a known number, and verify the emergency with another trusted person before sending money.
  2. Mountain View desk4 min
    Google’s SynthID Detector: How to Check AI-Generated Images, Video and AudioGoogle’s SynthID Detector looks for an embedded watermark in supported images, video and audio. Here is what its results do—and do not—show.
  3. Redmond desk20 min
    How to create a link to File or Folder in Windows 11Windows 11 gives you several ways to point to a file or folder without moving or duplicating it. You can create a desktop shortcut,…
Recommended PC Tool
Recommended PC Tool
PC Slower Than It Used to Be?Free scan - under a minute
Outdated Drivers Are Slowing You DownFree scan - exact matches

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.