Hardware FixRecommendedDevice not working? Your driver may be the problemCheck updates for common hardware issues.Fix DriversOctober DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsWindows FixRecommendedWindows errors stealing your time? Find the fix fastScan stability, cleanup and performance issues.Fix Now×
Skip to content
World desk5 min

A RAG Agent Can Refuse Every Attack and Still Fail Its Users

A safe-sounding refusal is only one part of a RAG agent’s execution. Assess tool actions, data boundaries, attack impact, and legitimate-task completion separately.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Yes. A RAG agent can give a safe-sounding refusal and still fail: it may have already taken an unauthorized action, exposed information, or abandoned the user’s legitimate task. A refusal is one observable response—not proof that the agent was secure, or that it was useful. To assess it, inspect the full execution trace and measure attack impact, task completion, and security boundaries separately.

Why a final refusal is not a security verdict

Retrieved text can carry hostile instructions

Retrieval-augmented generation (RAG) gives a language model material fetched from external documents, databases, or other sources. That material can contain instructions placed there by an attacker, even though the agent did not receive them directly from the user. OWASP describes document poisoning as malicious content entering a retrieval corpus and later reaching the model as context. Its guidance also warns that invisible Unicode or instructions split across chunks can make detection harder. NIST calls this kind of indirect prompt injection agent hijacking: hostile instructions embedded in ingested data can steer an agent toward unintended actions.

As an Amazon Associate I earn from qualifying purchases.

The risk crosses several stages: ingestion, indexing, retrieval, generation, output handling, and any downstream tools. As OWASP’s RAG Security Cheat Sheet puts it, RAG redistributes risk across the data pipeline rather than removing it. A document may be the entry point, but the consequences can occur later—when a model answers, discloses retrieved content, or attempts an action.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The final answer does not show the whole execution

An agent’s final message is not a complete record of what happened. OWASP’s LLM Prompt Injection Prevention Cheat Sheet cautions that a final refusal does not undo an action already taken. For example, a refusal displayed after a tool call cannot reverse a state change that the tool has already committed. Nor does refusal alone establish whether sensitive data was exposed earlier in the run.

#1 Best Overall
Yubico - Security Key C NFC - Basic Compatibility - Multi-Factor authentication (MFA) Security Key and passkey, Connect via USB-C or NFC, FIDO Certified
  • POWERFUL SECURITY KEY: The Security Key C NFC is the essential physical passkey for protecting your digital life from phishing attacks. It ensures only you can access your accounts.
  • WORKS WITH 1000+ ACCOUNTS: Compatible with Google, Microsoft, and Apple. A single Security Key C NFC secures 100 of your favorite accounts, including email, password managers, and more.
  • FAST & CONVENIENT LOGIN: Plug in your Security Key C NFC via USB-C and tap it, or tap it against your phone (NFC) to authenticate. No batteries, no internet connection, and no extra fees required.
  • TRUSTED PASSKEY TECHNOLOGY: Uses the latest passkey standards (FIDO2/WebAuthn & FIDO U2F) but does not support One-Time Passwords. For complex needs, check out the YubiKey 5 Series.
  • BUILT TO LAST: Made from tough, waterproof, and crush-resistant materials. Manufactured in Sweden and programmed in the USA with the highest security standards.

There is a separate failure mode even when no unauthorized action occurs: the agent can block the attack but also fail to complete the user’s benign request. That is a utility failure, not evidence of a successful attack. The title describes possible outcomes that security testing should distinguish; the cited sources do not establish how often this exact sequence occurs in deployed systems.

What published attack results do—and do not—show

Research and evaluations demonstrate that indirect prompt injection can affect agents, but their percentages describe particular setups, not a universal production failure rate. The studies differ in models, tasks, attack sets, environments, and definitions of success, so their numbers should not be compared as if they came from one shared test.

Evaluation Reported result Scope of the figure
InjecAgent, Findings of ACL 2024 24% vulnerability The authors report this for ReAct-prompted GPT-4 on the benchmark’s tested attacks. The benchmark included 1,054 test cases, 17 user tools, and 62 attacker tools; this is not a current model-wide or deployment-wide rate.
NIST CAISI, 2025 Attack success increased from 11% for the strongest baseline to 81% for the strongest novel attack A specific evaluation of agents powered by the upgraded Claude 3.5 Sonnet; the novel attacks were developed with the UK AI Security Institute. The figures are not general agent success rates.
Rag ’n Roll preprint, posted August 9, 2024 About 40% attack success across tested configurations; 60% when ambiguous answers also counted as successful Results apply to the authors’ tested application and their rule for counting ambiguous answers.
WASP, NeurIPS 2025 Up to 86% partial attack success in its end-to-end evaluation Partial success is not full completion of an attacker’s goal; the authors report agents often struggled to complete those goals fully.

These results make the case for end-to-end testing, not for one headline statistic about every RAG agent. None supplies a population estimate for agents that refuse an attack yet fail their users.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #2
Yubico - Security Key NFC - Basic Compatibility - Multi-Factor Authentication (MFA) Key, Connect via USB-A or NFC, FIDO Certified
  • POWERFUL SECURITY KEY: The Security Key NFC is the essential physical passkey for protecting your digital life from phishing attacks. It ensures only you can access your accounts.
  • WORKS WITH 1000+ ACCOUNTS: Compatible with Google, Microsoft, and Apple. A single Security Key NFC secures 100 of your favorite accounts, including email, password managers, and more.
  • FAST & CONVENIENT LOGIN: Plug in your Security Key NFC via USB-A and tap it, or tap it against your phone (NFC) to authenticate. No batteries, no internet connection, and no extra fees required.
  • TRUSTED PASSKEY TECHNOLOGY: Uses the latest passkey standards (FIDO2/WebAuthn & FIDO U2F) but does not support One-Time Passwords. For complex needs, check out the YubiKey 5 Series.
  • BUILT TO LAST: Made from tough, waterproof, and crush-resistant materials. Manufactured in Sweden and programmed in the USA with the highest security standards.

How to evaluate both security and usefulness

Test the agent’s behavior across the whole task, rather than scoring only its final wording. For each run, record the original user goal, what the agent retrieved, which instructions it followed, its outputs and tool calls, and any resulting state changes. Include malicious content in the retrieval path—not only as a direct user prompt—because that is where indirect prompt injection enters.

  • Attack impact: Did retrieved content change the answer, trigger a prohibited action, or expose data?
  • Legitimate utility: Did the agent correctly complete the user’s original task, including when it needed to ignore or safely report hostile content?
  • Boundary integrity: Did retrieval access control, tenant separation, tool permissions, and output constraints hold?

Report these outcomes separately. An agent that refuses more often may reduce some attack opportunities while also blocking benign work. A single refusal rate cannot tell you whether the system protected data, avoided unauthorized state changes, or completed legitimate tasks.

NIST recommends adaptive evaluations, task-specific analysis alongside aggregate performance, and consideration of multiple attempts. For a useful report, describe the tested model and configuration, tasks, attack placement, number of attempts, and precise success definitions. Show results by task as well as in aggregate: a combined score can obscure a weakness that matters for a particular workflow.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How to secure a RAG agent in practice

No single filter makes a RAG system safe. Controls need to protect each boundary between untrusted content, the model, and tools that can change state.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Protect the corpus and retrieval boundary

  • Track document provenance and integrity, and restrict who can add or modify indexed content.
  • Enforce access metadata and tenant isolation at retrieval time; do not rely on the model to disregard documents a user should not see.
  • Inspect retrieved material for suspicious content, while recognizing that text split across chunks or encoded in unusual ways can complicate detection.

A matching document digest can show that a file is consistent with an approved baseline. OWASP cautions that a digest does not prove the document is safe or free of prompt injection.

Keep context bounded and treat it as untrusted

OWASP offers 3–5 retrieved chunks totaling 2,000–4,000 tokens as a reasonable starting point, not a universal safe limit. Models vary in how they attend to context, so test chunk count, size, and position with the model and tasks you actually use. Clearly separate trusted instructions from retrieved material, and avoid granting retrieved text the authority of developer instructions.

Rank #4
Sale
Thetis Nano-A FIDO2 Security Key Hardware Passkey Device with USB Type A, TOTP/HOTP, FIDO2.0 Two Factor Authentication 2FA MFA, Works with Windows/mac/iOS/Android/Linux/Gmail/Facebook/GitHub/Coinbase
  • Ultra-Compact FIDO2 Security Key - Plug-and-stay or carry on a keychain. This USB-A hardware security key offers portable, always-on protection for desktop and mobile use. (Item Size: 0.75 X 0.74 IN x 0.25 IN)
  • USB-A Hardware Key for All Devices - Works with USB-A ports on PC, Mac, Android, and other laptop/notebook device. Enables secure, cross-platform login with FIDO2.0 passkey support.
  • FIDO Certified Security Key - Meets FIDO and FIDO2 standards. Works with Google, Microsoft, GitHub, Dropbox, and more. Please check service compatibility before purchase.
  • Passwordless Login with Passkey - Supports passkey login via WebAuthn and CTAP2. Enjoy password-free sign-ins where supported. Not all websites or services currently support passkeys.
  • Advanced Multi-Factor Authentication - Offers 200 FIDO2 passkey slots and 50 OATH-TOTP slots. Strong, flexible 2FA/MFA support across various apps and authentication platforms.

Constrain tools and validate outputs

  • Give tools only the permissions needed for the task; require explicit authorization for sensitive or irreversible actions.
  • Use allowlisted action schemas and validate arguments and outputs before execution or delivery.
  • Apply output checks for unintended data disclosure or unsafe downstream instructions; a clean-looking final answer is not a substitute for inspecting earlier tool calls.
  • Log retrievals, decisions, tool invocations, and state changes so an incident can be reconstructed.
  • Fail closed when authorization or validation fails: do not execute an unverified action merely because the model proposed it.

What a meaningful security result should say

When reviewing an agent or a vendor’s evaluation, look for evidence across five dimensions rather than a bare claim that it “resists prompt injection”:

  • Coverage: Were ingestion, retrieval, generation, output, and downstream tool execution all tested?
  • Security outcomes: Were attack success, data disclosure, and unauthorized state changes measured distinctly?
  • Utility cost: Were benign task completion, correctness, refusals, and false blocks measured too?
  • Evaluation quality: Did tests include realistic end-to-end tasks, adaptive attacks, multiple attempts, and transparent success definitions?
  • Operational evidence: Were logs, provenance, integrity checks, access controls, and reproducible configurations available?

A refusal is useful evidence about the response the agent produced. It is not, by itself, evidence that no attack succeeded or that the user’s task was completed. Those conclusions require separate measurements of actions, data boundaries, and legitimate outcomes.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More from the Wire

  1. World desk4 min
    How to Spot an AI Voice Scam Before Sending MoneyDon’t rely on how a caller sounds. Pause, call back through a known number, and verify the emergency with another trusted person before sending money.
  2. Mountain View desk4 min
    Google’s SynthID Detector: How to Check AI-Generated Images, Video and AudioGoogle’s SynthID Detector looks for an embedded watermark in supported images, video and audio. Here is what its results do—and do not—show.
  3. Redmond desk20 min
    How to create a link to File or Folder in Windows 11Windows 11 gives you several ways to point to a file or folder without moving or duplicating it. You can create a desktop shortcut,…
Recommended PC Tool
Recommended PC Tool
Crashes, No Sound, or Screen Glitches?Free driver scan
PC Slower Than It Used to Be?Free scan - under a minute

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.