Free tools Windows power users keep installed
One-click scans. No signup required.
Yes. A RAG agent can give a safe-sounding refusal and still fail: it may have already taken an unauthorized action, exposed information, or abandoned the user’s legitimate task. A refusal is one observable response—not proof that the agent was secure, or that it was useful. To assess it, inspect the full execution trace and measure attack impact, task completion, and security boundaries separately.
Why a final refusal is not a security verdict
Retrieved text can carry hostile instructions
Retrieval-augmented generation (RAG) gives a language model material fetched from external documents, databases, or other sources. That material can contain instructions placed there by an attacker, even though the agent did not receive them directly from the user. OWASP describes document poisoning as malicious content entering a retrieval corpus and later reaching the model as context. Its guidance also warns that invisible Unicode or instructions split across chunks can make detection harder. NIST calls this kind of indirect prompt injection agent hijacking: hostile instructions embedded in ingested data can steer an agent toward unintended actions.
As an Amazon Associate I earn from qualifying purchases.
The risk crosses several stages: ingestion, indexing, retrieval, generation, output handling, and any downstream tools. As OWASP’s RAG Security Cheat Sheet puts it, RAG redistributes risk across the data pipeline rather than removing it. A document may be the entry point, but the consequences can occur later—when a model answers, discloses retrieved content, or attempts an action.
Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Repair Windows errors before they cause bigger problemsFix Now →The final answer does not show the whole execution
An agent’s final message is not a complete record of what happened. OWASP’s LLM Prompt Injection Prevention Cheat Sheet cautions that a final refusal does not undo an action already taken. For example, a refusal displayed after a tool call cannot reverse a state change that the tool has already committed. Nor does refusal alone establish whether sensitive data was exposed earlier in the run.
#1 Best Overall
- POWERFUL SECURITY KEY: The Security Key C NFC is the essential physical passkey for protecting your digital life from phishing attacks. It ensures only you can access your accounts.
- WORKS WITH 1000+ ACCOUNTS: Compatible with Google, Microsoft, and Apple. A single Security Key C NFC secures 100 of your favorite accounts, including email, password managers, and more.
- FAST & CONVENIENT LOGIN: Plug in your Security Key C NFC via USB-C and tap it, or tap it against your phone (NFC) to authenticate. No batteries, no internet connection, and no extra fees required.
- TRUSTED PASSKEY TECHNOLOGY: Uses the latest passkey standards (FIDO2/WebAuthn & FIDO U2F) but does not support One-Time Passwords. For complex needs, check out the YubiKey 5 Series.
- BUILT TO LAST: Made from tough, waterproof, and crush-resistant materials. Manufactured in Sweden and programmed in the USA with the highest security standards.
There is a separate failure mode even when no unauthorized action occurs: the agent can block the attack but also fail to complete the user’s benign request. That is a utility failure, not evidence of a successful attack. The title describes possible outcomes that security testing should distinguish; the cited sources do not establish how often this exact sequence occurs in deployed systems.
What published attack results do—and do not—show
Research and evaluations demonstrate that indirect prompt injection can affect agents, but their percentages describe particular setups, not a universal production failure rate. The studies differ in models, tasks, attack sets, environments, and definitions of success, so their numbers should not be compared as if they came from one shared test.
| Evaluation | Reported result | Scope of the figure |
|---|---|---|
| InjecAgent, Findings of ACL 2024 | 24% vulnerability | The authors report this for ReAct-prompted GPT-4 on the benchmark’s tested attacks. The benchmark included 1,054 test cases, 17 user tools, and 62 attacker tools; this is not a current model-wide or deployment-wide rate. |
| NIST CAISI, 2025 | Attack success increased from 11% for the strongest baseline to 81% for the strongest novel attack | A specific evaluation of agents powered by the upgraded Claude 3.5 Sonnet; the novel attacks were developed with the UK AI Security Institute. The figures are not general agent success rates. |
| Rag ’n Roll preprint, posted August 9, 2024 | About 40% attack success across tested configurations; 60% when ambiguous answers also counted as successful | Results apply to the authors’ tested application and their rule for counting ambiguous answers. |
| WASP, NeurIPS 2025 | Up to 86% partial attack success in its end-to-end evaluation | Partial success is not full completion of an attacker’s goal; the authors report agents often struggled to complete those goals fully. |
These results make the case for end-to-end testing, not for one headline statistic about every RAG agent. None supplies a population estimate for agents that refuse an attack yet fail their users.
Rank #2
- POWERFUL SECURITY KEY: The Security Key NFC is the essential physical passkey for protecting your digital life from phishing attacks. It ensures only you can access your accounts.
- WORKS WITH 1000+ ACCOUNTS: Compatible with Google, Microsoft, and Apple. A single Security Key NFC secures 100 of your favorite accounts, including email, password managers, and more.
- FAST & CONVENIENT LOGIN: Plug in your Security Key NFC via USB-A and tap it, or tap it against your phone (NFC) to authenticate. No batteries, no internet connection, and no extra fees required.
- TRUSTED PASSKEY TECHNOLOGY: Uses the latest passkey standards (FIDO2/WebAuthn & FIDO U2F) but does not support One-Time Passwords. For complex needs, check out the YubiKey 5 Series.
- BUILT TO LAST: Made from tough, waterproof, and crush-resistant materials. Manufactured in Sweden and programmed in the USA with the highest security standards.
How to evaluate both security and usefulness
Test the agent’s behavior across the whole task, rather than scoring only its final wording. For each run, record the original user goal, what the agent retrieved, which instructions it followed, its outputs and tool calls, and any resulting state changes. Include malicious content in the retrieval path—not only as a direct user prompt—because that is where indirect prompt injection enters.
- Attack impact: Did retrieved content change the answer, trigger a prohibited action, or expose data?
- Legitimate utility: Did the agent correctly complete the user’s original task, including when it needed to ignore or safely report hostile content?
- Boundary integrity: Did retrieval access control, tenant separation, tool permissions, and output constraints hold?
Report these outcomes separately. An agent that refuses more often may reduce some attack opportunities while also blocking benign work. A single refusal rate cannot tell you whether the system protected data, avoided unauthorized state changes, or completed legitimate tasks.
NIST recommends adaptive evaluations, task-specific analysis alongside aggregate performance, and consideration of multiple attempts. For a useful report, describe the tested model and configuration, tasks, attack placement, number of attempts, and precise success definitions. Show results by task as well as in aggregate: a combined score can obscure a weakness that matters for a particular workflow.
Rank #3
How to secure a RAG agent in practice
No single filter makes a RAG system safe. Controls need to protect each boundary between untrusted content, the model, and tools that can change state.
Protect the corpus and retrieval boundary
- Track document provenance and integrity, and restrict who can add or modify indexed content.
- Enforce access metadata and tenant isolation at retrieval time; do not rely on the model to disregard documents a user should not see.
- Inspect retrieved material for suspicious content, while recognizing that text split across chunks or encoded in unusual ways can complicate detection.
A matching document digest can show that a file is consistent with an approved baseline. OWASP cautions that a digest does not prove the document is safe or free of prompt injection.
Keep context bounded and treat it as untrusted
OWASP offers 3–5 retrieved chunks totaling 2,000–4,000 tokens as a reasonable starting point, not a universal safe limit. Models vary in how they attend to context, so test chunk count, size, and position with the model and tasks you actually use. Clearly separate trusted instructions from retrieved material, and avoid granting retrieved text the authority of developer instructions.
Rank #4
- Ultra-Compact FIDO2 Security Key - Plug-and-stay or carry on a keychain. This USB-A hardware security key offers portable, always-on protection for desktop and mobile use. (Item Size: 0.75 X 0.74 IN x 0.25 IN)
- USB-A Hardware Key for All Devices - Works with USB-A ports on PC, Mac, Android, and other laptop/notebook device. Enables secure, cross-platform login with FIDO2.0 passkey support.
- FIDO Certified Security Key - Meets FIDO and FIDO2 standards. Works with Google, Microsoft, GitHub, Dropbox, and more. Please check service compatibility before purchase.
- Passwordless Login with Passkey - Supports passkey login via WebAuthn and CTAP2. Enjoy password-free sign-ins where supported. Not all websites or services currently support passkeys.
- Advanced Multi-Factor Authentication - Offers 200 FIDO2 passkey slots and 50 OATH-TOTP slots. Strong, flexible 2FA/MFA support across various apps and authentication platforms.
Constrain tools and validate outputs
- Give tools only the permissions needed for the task; require explicit authorization for sensitive or irreversible actions.
- Use allowlisted action schemas and validate arguments and outputs before execution or delivery.
- Apply output checks for unintended data disclosure or unsafe downstream instructions; a clean-looking final answer is not a substitute for inspecting earlier tool calls.
- Log retrievals, decisions, tool invocations, and state changes so an incident can be reconstructed.
- Fail closed when authorization or validation fails: do not execute an unverified action merely because the model proposed it.
What a meaningful security result should say
When reviewing an agent or a vendor’s evaluation, look for evidence across five dimensions rather than a bare claim that it “resists prompt injection”:
- Coverage: Were ingestion, retrieval, generation, output, and downstream tool execution all tested?
- Security outcomes: Were attack success, data disclosure, and unauthorized state changes measured distinctly?
- Utility cost: Were benign task completion, correctness, refusals, and false blocks measured too?
- Evaluation quality: Did tests include realistic end-to-end tasks, adaptive attacks, multiple attempts, and transparent success definitions?
- Operational evidence: Were logs, provenance, integrity checks, access controls, and reproducible configurations available?
A refusal is useful evidence about the response the agent produced. It is not, by itself, evidence that no attack succeeded or that the user’s task was completed. Those conclusions require separate measurements of actions, data boundaries, and legitimate outcomes.
The Tool Desk
Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




