DriversRecommendedOutdated drivers can make a good PC feel brokenScan driver issues before chasing fixes manually.Scan NowOctober DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsClean PCRecommendedOne scan can reveal what keeps slowing WindowsLook for cleanup and repair opportunities.Run Scan×
Skip to content
World desk5 min

Read the Approval Split Before Trusting an Agent-Security Benchmark

A security benchmark should separate hard blocks from approval-required outcomes. See how scope, benign controls, replay methods, and dataset independence change what the headline result means.
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

An agent-security benchmark’s headline count can blur two very different outcomes: a tool blocked an action, or it asked a human to decide. In Doberman’s reported RedCode run, 713 of 720 in-scope attack cases were either blocked or routed for approval—but only 589 were hard blocks. The 124 AUTH outcomes still depended on an operator.

What the approval split changes

Security results are meaningful only when their outcome labels stay distinct. A BLOCK means the tested rule stopped the action. AUTH means the action was held for an operator’s approval; it is not evidence that the action was blocked regardless of what the operator chooses. PASS means the action was allowed by the tested rules.

In Alan Fu’s October 1, 2026 account of a Doberman run against RedCode, the recorded test at revision b689a9d included 1,410 attack records. Of those, 690 were outside the declared threat model, leaving 720 in-scope cases. The deterministic rules returned BLOCK for 589, AUTH for 124, and PASS for seven. Thus, 713 of 720 were blocked or required approval, while 589 of 720 were hard-blocked. Those are both accurate summaries, but they describe different security outcomes. Source: Alan Fu, DEV Community, October 1, 2026.

Read the denominator and the outcome counts together

Always ask which cases were counted and what each result means. Here, the 690 excluded records matter: a percentage calculated over all 1,410 records would answer a different question from one calculated over the 720 cases inside the declared threat model.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Measure in the recorded run Reported result What it establishes
Attack records 1,410 total; 690 out of scope; 720 in scope The threat-model boundary and denominator used for the in-scope result.
BLOCK 589 of 720 in-scope cases These cases received a block outcome from the deterministic rules.
AUTH 124 of 720 in-scope cases These cases required an operator decision; approval-dependent cases are not hard blocks.
PASS 7 of 720 in-scope cases These cases passed the tested rules.
BLOCK or AUTH combined 713 of 720 in-scope cases The action was either blocked or made dependent on operator approval; the combined figure does not mean 713 hard blocks.

Keep the threat model beside the counts. An excluded case is not a block, pass, or failure within the reported in-scope denominator; it is outside the question that result answers.

Put benign friction beside attack outcomes

A guardrail can stop unwanted actions and still interrupt legitimate work. The same run included 60 synthetic benign controls: 56 passed, three received AUTH, and one was blocked. That is useful friction evidence, but the controls were synthetic, not production user sessions. They should not be presented as a measured rate of disruption for real users. Doberman run and control counts.

Small, specific subsets also need their own scope. All 30 reverse-shell-listener cases in the run received BLOCK; that describes those 30 cases, not every possible reverse shell. In 60 process-kill cases, all required intervention: 13 were BLOCK and 47 AUTH. The latter result is not equivalent to 60 hard blocks.

Check what kind of evaluation produced the number

The RedCode evaluation replayed mapped tool-call cases through a deterministic engine. It did not drive a live model through a complete attack campaign, nor did it measure the full adaptive layer. Its counts are historical results for the recorded run, not a fresh evaluation of whatever release a reader encounters later. Evaluation details.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Before comparing benchmarks, establish whether the test is a replay, a live workflow, model-free, or adaptive. These methods answer different questions. A replay can make a defined set of cases reproducible; it does not by itself establish how a system responds as a live model adapts across a campaign. Similarly, a guarantee tied to a test is relevant only when that test exercises the property the reader relies on. The source article points to a host parity matrix, underscoring that results may depend on the host as well as the product. Host and test context.

Audit benchmark design, labels, and independence

OASB: useful structure, not a product pass

OASB version 0.4.0 documents 222 standardized attack scenarios with mappings to MITRE ATLAS and OWASP. Its documentation describes running adapters against a suite and marking undeclared capabilities N/A rather than FAIL. The specifications distinguish tool-detection benchmarking from governance auditing. These design details can help readers understand what a benchmark covers, but they do not show that any particular product passed it. OASB overview · OASB specifications, version 0.4.0 · OASB-1 getting started.

Label provenance can make a metric circular

OpenA2A disclosed that it withdrew OASB F1, precision, and false-positive-rate figures after finding that its benign class had been selected using the scanner’s own labels. That selection made the near-zero false-positive result circular. The OASB page reports recall of 223/270 (82.6%) on author-created attack fixtures and 234/495 (47.3%) when self-labeled samples are included; it says it is remeasuring with corpora it neither owns nor labeled. These figures are tied to those datasets and labels, not broad population performance. OASB results and methodology.

Separate maintainer-run results from independent validation

MoorAI reports three scored runs, all executed by its maintainer, and says its repository contains no third-party lab reproductions. Its methodology describes locked test halves intended to check generalization against tuning. A held-out split offers stronger evidence only if it has genuinely remained unseen during development; maintainer-run results and independent reproduction are different levels of validation. MoorAI benchmark methodology and results.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Treat standards proposals as proposals

An IETF Internet-Draft dated July 5, 2026 proposes four first-level dimensions and 55 second-level metrics spanning static, dynamic, attack-defense, compliance, and quantitative evaluation. It is an informational draft—not a certification and not a product result. Cite its status and date rather than treating the framework as an endorsement or passed test. IETF draft, July 5, 2026.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

A practical checklist for evaluating a security claim

Before relying on a benchmark headline, look for these details in the result itself:

  • Product and environment: Name the product version and host tested, and date the result.
  • Scope: State the threat model, included cases, excluded cases, and the denominator for each metric.
  • Outcome definitions: Publish separate counts for BLOCK, AUTH or approval-required, PASS, and detection-only outcomes where applicable.
  • Benign controls: Show friction or false-positive measures alongside attacks, and identify whether benign examples are synthetic or production-derived.
  • Evaluation method: Say whether cases were replayed or run live, whether a model participated, and whether adaptive behavior was tested.
  • Dataset integrity: Disclose who selected and labeled the data, whether a test split was held out during development, and whether it was independently evaluated.
  • Claim-to-test fit: Confirm that the test actually exercises the security property the claim is about.

For a published result, make these details adjacent to the headline rather than leaving readers to infer them: tested product, version and host; corpus and threat model; counts by outcome; benign controls and label provenance; test method; evaluator and held-out-set status; and the evaluation date.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Leave a Reply

Your email address will not be published. Required fields are marked *

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More from the Wire

  1. World desk4 min
    How to Spot an AI Voice Scam Before Sending MoneyDon’t rely on how a caller sounds. Pause, call back through a known number, and verify the emergency with another trusted person before sending money.
  2. Mountain View desk4 min
    Google’s SynthID Detector: How to Check AI-Generated Images, Video and AudioGoogle’s SynthID Detector looks for an embedded watermark in supported images, video and audio. Here is what its results do—and do not—show.
  3. Redmond desk20 min
    How to create a link to File or Folder in Windows 11Windows 11 gives you several ways to point to a file or folder without moving or duplicating it. You can create a desktop shortcut,…
Recommended PC Tool
Recommended PC Tool
Windows Errors? Fix Them Before They SpreadFree repair scan
Outdated Drivers Are Slowing You DownFree scan - exact matches

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.