October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsWindows FixRecommendedWindows errors stealing your time? Find the fix fastScan stability, cleanup and performance issues.Fix NowOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
World desk4 min

AI Security Testing: What Can Be Measured—and What Cannot

AI security is not one test category. AV-Comparatives says meaningful assessments need a defined claim, observable outcome, and clearly stated conditions.
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Yes—but only when the claim and test conditions are specific. AV-Comparatives argues that “AI security” is too broad to serve as one reliable, reproducible test category. Evaluators can measure outcomes such as whether a security product detects malware or whether a defined control blocks a specified agent attack. Those results do not, by themselves, prove that AI security as a whole has been tested.

Why “AI security” is not one test category

“AI” can refer to different technologies and functions, not one discrete security control. In a security product, machine learning may contribute to malware detection, behavioral analysis, phishing protection, anomaly detection, or EDR/XDR. These functions operate alongside other mechanisms, so an external evaluator generally cannot isolate which one caused a detection or prevention result.

As an Amazon Associate I earn from qualifying purchases.

For that reason, AV-Comparatives says it assesses the product’s security outcome rather than attributing that outcome to AI. Its 26 August 2026 article frames the central question as whether the test has a defined subject and whether its results can be measured objectively, reproduced, and judged fairly.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What can be tested meaningfully?

Security outcomes from AI-enabled products

A product can be tested for observable outcomes even when it uses AI internally. Depending on the product and claim, that could mean measuring malware detection, false positives, attack prevention, or useful EDR telemetry. The result describes how the product performed under the test conditions; it does not establish that AI, rather than another component, produced the result.

This outcome-based approach is consistent with the published AV-Comparatives methodologies, which describe real-world protection, performance, anti-phishing, and malware-removal testing. Its Advanced Threat Protection archive also describes evaluations using hacking and penetration techniques against targeted threats such as exploits and fileless attacks.

Specific protections for AI agents

A narrower functional assessment can test a stated control against a defined attack. For example, an evaluator could expose a product that claims to protect an agent from indirect prompt injection to controlled malicious content, then check whether the attack causes an unauthorized action or data disclosure. Other bounded scenarios might test malicious tool use, unauthorized data access, or attempted exfiltration.

The conclusion should say whether the control prevented that attack under the tested conditions. It should not imply that the assessment covers every agent, attack, or meaning of “AI security.”

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Why agent test results can be hard to interpret

An agent’s behavior can depend on its model and version, system prompt, framework, available tools, permissions, memory, context, external information, and configuration. Cloud-hosted systems can also change outside a tester’s control. That means a result may describe one particular setup rather than a stable, general property of the product.

Attribution is another challenge. If an attack fails, the security control may have blocked it—or the model may simply have refused the request, or another environmental factor may have prevented it. Repeat runs and statistical analysis can reduce uncertainty, but they do not eliminate this ambiguity. A precise-looking percentage can still depend heavily on the environment and scenarios selected.

What a credible test should disclose

When comparing claims about AI-security testing, look for enough detail to understand what was tested and what the result actually means:

  • Claim and outcome: Is the test tied to a named security claim and an observable result, such as prevention, detection, or data disclosure?
  • System and configuration: Does it identify the model and version, agent framework, prompts, tools, permissions, memory, and relevant configuration?
  • Attribution: Can the evaluator distinguish the effect of the security control from model refusal or other environmental factors?
  • Repeatability and scope: Were runs repeated, and does the conclusion stay within the system and conditions actually tested?
  • Representativeness and currency: Do the scenarios reflect relevant architectures and attacks, and could changes to the model or threat landscape make the results stale?

These are practical comparison questions derived from the limitations AV-Comparatives identifies, not results of an independent test.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What AV-Comparatives recommends instead

AV-Comparatives says focused functional assessments, individual product reviews, and dedicated research projects are currently more appropriate for emerging AI-security technologies than a broad “AI security test.” Its stated standard is a clearly defined subject, measurable security outcomes, and a methodology that is sufficiently objective, repeatable, representative, and fair.

The organization’s position is not that security involving AI is untestable. It is that the result must be bounded: a test can show what happened in a specified scenario, while a category-wide claim requires evidence that the test cannot supply on its own. As AV-Comparatives CEO and co-founder Andreas Clementi puts it: “Independent testing should measure what can be demonstrated, not what is currently fashionable.”

What the available figures do—and do not—show

AV-Comparatives’ central article does not publish a quantitative result establishing whether AI-security tests are reliable. A separate 2026 security survey announcement reports 1,328 valid responses from 87 countries and says the survey covered AI chatbot use and participants’ perspectives on potential cyber threats. Those are survey participation facts, not evidence that settles the testing-methodology question.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Leave a Reply

Your email address will not be published. Required fields are marked *

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More from the Wire

  1. World desk4 min
    How to Spot an AI Voice Scam Before Sending MoneyDon’t rely on how a caller sounds. Pause, call back through a known number, and verify the emergency with another trusted person before sending money.
  2. Mountain View desk4 min
    Google’s SynthID Detector: How to Check AI-Generated Images, Video and AudioGoogle’s SynthID Detector looks for an embedded watermark in supported images, video and audio. Here is what its results do—and do not—show.
  3. Redmond desk20 min
    How to create a link to File or Folder in Windows 11Windows 11 gives you several ways to point to a file or folder without moving or duplicating it. You can create a desktop shortcut,…
Recommended PC Tool
Recommended PC Tool
Outdated Drivers Are Slowing You DownFree scan - exact matches
Windows Errors? Fix Them Before They SpreadFree repair scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.