Do not rely on an AI agent’s system prompt to enforce permissions. Enforce authorization in the tools and runtime that execute its actions: give the agent only task-specific access, isolate its environment, and require approval for consequential operations. That way, a prompt-injection failure does not automatically become a security incident.
Why an agent may ignore its written rules
An agent can encounter malicious instructions while doing legitimate work. A webpage, email, or document it was asked to read may contain directions designed to make it misuse an available tool. NIST calls this kind of attack agent hijacking and identifies the difficulty of separating trusted instructions from untrusted data as part of the problem. NIST CAISI’s January 2025 discussion of agent-hijacking evaluations describes the risk.
The important question is not only whether the model recognizes an attack. If the model follows hostile content, the damage depends on the agent’s tools, credentials, orchestration, and execution environment. OpenAI’s March 11, 2026 guidance discusses how contextual manipulation and social engineering make filtering alone insufficient; it recommends constraining what an agent can do if manipulation succeeds. OpenAI: Designing AI agents to resist prompt injection
Marking retrieved content as untrusted can help, but it is not an access-control mechanism. OWASP cautions that labeling alone does not create a security boundary. OWASP’s LLM Prompt Injection Prevention guidance
#1 Best Overall
Where to enforce agent permissions
Enforce permissions in ordinary execution code—at the point where a tool or service checks and performs an operation—not in model-generated text. Each request should be checked against the caller, resource, action, and arguments. The model can propose an action; it must not determine its own authority. OWASP’s AI Agent Security Cheat Sheet
- Scope tools to the task. Expose only necessary tools and resources. Separate read-only access from write-capable operations and avoid broad or wildcard permissions.
- Authorize every side effect. Check each operation and its parameters when the tool executes, including requests routed through an orchestration layer.
- Make approval specific. For sensitive, irreversible, financial, administrative, or externally visible actions, show the reviewer the exact proposed action and arguments. Do not treat a general approval prompt as blanket permission.
- Keep downstream handling safe. Treat model output as untrusted when another system consumes it. For example, use parameterized database queries and safe rendering rather than inserting generated text directly into executable queries or markup.
Contain what an agent can reach
Authorization checks limit actions; runtime isolation limits the agent’s reach if a check or model-level defense fails. Restrict the files, processes, credentials, and network destinations available to the agent. Use filesystem boundaries, appropriate process or container isolation, narrowly scoped credentials, and egress controls. Credentials that are not available to the runtime cannot be retrieved from it through prompt injection. Anthropic’s response to NIST describes security as a property of the whole agent system, not just the model. Anthropic’s NIST RFI on Agentic Security
Do not assume a connector is safe merely because your organization approved it. It may return attacker-controlled content from a website, document, or other source. Treat connector results as untrusted input and apply authorization checks to the resulting proposed action. Anthropic’s containment guidance also discusses restricting what agents can access in their execution environments. Anthropic: How we contain Claude across products
Choose a permission model the team can explain
NIST’s 2025 tool-use taxonomy gives teams useful language for describing agent capabilities: read-only, constrained-write, or write access, in trusted or untrusted environments. NIST presents this as a taxonomy to adapt, not a definitive standard or a ready-made security ranking. NIST: Lessons Learned from the Consortium: Tool Use in Agent Systems
Quick wins for a faster PC:
Repair Windows errors before they cause bigger problemsFix Now →Scan for outdated or missing drivers - takes under a minuteDriver Scan →Clear out junk files and repair common Windows errorsFree Scan →| Design area | Questions to answer |
|---|---|
| Tool authority | Which tools are exposed? Are permissions scoped by operation and resource? Can the agent read, write, or both? |
| Runtime isolation | Which files, processes, credentials, and network destinations can the agent reach? What is outside its environment? |
| Action review | Which operations need approval? Is approval tied to the exact action and arguments, and can it be replayed? |
| Untrusted inputs | Can external data, tool descriptions, or connector results influence tool selection or arguments? |
| Observability and recovery | Are tool calls and policy decisions logged? Can access be revoked and the agent stopped? |
| Evaluation quality | Do tests reflect the deployed agent’s actual tools, data, and task-specific risks? |
Test the deployed boundaries, not just the prompt
Build abuse cases around every external content channel the agent reads and every tool that can change state or send information. Before testing, define the legitimate task, the prohibited result, and what observable evidence would show that the prohibited result occurred.
- Map inputs and side effects. List the web pages, emails, files, and connector results the agent can read, plus every tool that can modify a resource or transmit data.
- Exercise realistic attack paths. Test direct and indirect prompt injection, harmful tool arguments, attempted data exfiltration, privilege escalation, and attempts to bypass review.
- Use safe test fixtures. Run with dummy data and instrumented or sandboxed tools so a successful attack cannot affect production resources.
- Repeat with adaptive attempts. Do not treat a pass against a fixed set of known strings as proof of safety. NIST CAISI recommends adaptive evaluations; its January 2025 experiments used then-current models and AgentDojo-derived scenarios, so their model-specific findings are not a current universal failure rate.
- Verify enforcement and recovery. Confirm that denied actions are blocked at execution, that logs capture decisions and tool calls, and that operators can revoke access or stop an agent.
OWASP notes that its sample smoke tests are illustrative, not a representative security benchmark. Use them as starting points, then evaluate scenarios that match your deployment’s actual tools and data. OWASP AI Agent Security Cheat Sheet
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.What model defenses and benchmark scores can—and cannot—tell you
Vendor defenses and benchmark results can indicate progress on specific systems and evaluations, but they do not guarantee that an agent is safe in your deployment. A result depends on the tested model, attack setup, task, and number of attempts; your tools, data, permissions, and runtime may differ. Judge the security boundary by what the deployed system actually allows and by whether its controls hold when the model makes a bad decision.
Quick Recap
Best Value
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
Free tools Windows power users keep installed
One-click scans. No signup required.




