Driver FixRecommendedSound, Wi-Fi or graphics acting up? Check drivers firstFind missing or outdated drivers fast.Check DriversOctober DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsWindows FixRecommendedWindows errors stealing your time? Find the fix fastScan stability, cleanup and performance issues.Fix Now×
Skip to content
World desk4 min

AI Agents Need Security Boundaries They Cannot Rewrite

A system prompt is not an enforceable permission boundary. Secure AI agents by checking authorization at tool execution, limiting runtime access, and testing realistic attack paths.
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Do not rely on an AI agent’s system prompt to enforce permissions. Enforce authorization in the tools and runtime that execute its actions: give the agent only task-specific access, isolate its environment, and require approval for consequential operations. That way, a prompt-injection failure does not automatically become a security incident.

Why an agent may ignore its written rules

An agent can encounter malicious instructions while doing legitimate work. A webpage, email, or document it was asked to read may contain directions designed to make it misuse an available tool. NIST calls this kind of attack agent hijacking and identifies the difficulty of separating trusted instructions from untrusted data as part of the problem. NIST CAISI’s January 2025 discussion of agent-hijacking evaluations describes the risk.

The important question is not only whether the model recognizes an attack. If the model follows hostile content, the damage depends on the agent’s tools, credentials, orchestration, and execution environment. OpenAI’s March 11, 2026 guidance discusses how contextual manipulation and social engineering make filtering alone insufficient; it recommends constraining what an agent can do if manipulation succeeds. OpenAI: Designing AI agents to resist prompt injection

Marking retrieved content as untrusted can help, but it is not an access-control mechanism. OWASP cautions that labeling alone does not create a security boundary. OWASP’s LLM Prompt Injection Prevention guidance

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall

Where to enforce agent permissions

Enforce permissions in ordinary execution code—at the point where a tool or service checks and performs an operation—not in model-generated text. Each request should be checked against the caller, resource, action, and arguments. The model can propose an action; it must not determine its own authority. OWASP’s AI Agent Security Cheat Sheet

  • Scope tools to the task. Expose only necessary tools and resources. Separate read-only access from write-capable operations and avoid broad or wildcard permissions.
  • Authorize every side effect. Check each operation and its parameters when the tool executes, including requests routed through an orchestration layer.
  • Make approval specific. For sensitive, irreversible, financial, administrative, or externally visible actions, show the reviewer the exact proposed action and arguments. Do not treat a general approval prompt as blanket permission.
  • Keep downstream handling safe. Treat model output as untrusted when another system consumes it. For example, use parameterized database queries and safe rendering rather than inserting generated text directly into executable queries or markup.

Contain what an agent can reach

Authorization checks limit actions; runtime isolation limits the agent’s reach if a check or model-level defense fails. Restrict the files, processes, credentials, and network destinations available to the agent. Use filesystem boundaries, appropriate process or container isolation, narrowly scoped credentials, and egress controls. Credentials that are not available to the runtime cannot be retrieved from it through prompt injection. Anthropic’s response to NIST describes security as a property of the whole agent system, not just the model. Anthropic’s NIST RFI on Agentic Security

Do not assume a connector is safe merely because your organization approved it. It may return attacker-controlled content from a website, document, or other source. Treat connector results as untrusted input and apply authorization checks to the resulting proposed action. Anthropic’s containment guidance also discusses restricting what agents can access in their execution environments. Anthropic: How we contain Claude across products

Choose a permission model the team can explain

NIST’s 2025 tool-use taxonomy gives teams useful language for describing agent capabilities: read-only, constrained-write, or write access, in trusted or untrusted environments. NIST presents this as a taxonomy to adapt, not a definitive standard or a ready-made security ranking. NIST: Lessons Learned from the Consortium: Tool Use in Agent Systems

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Design area Questions to answer
Tool authority Which tools are exposed? Are permissions scoped by operation and resource? Can the agent read, write, or both?
Runtime isolation Which files, processes, credentials, and network destinations can the agent reach? What is outside its environment?
Action review Which operations need approval? Is approval tied to the exact action and arguments, and can it be replayed?
Untrusted inputs Can external data, tool descriptions, or connector results influence tool selection or arguments?
Observability and recovery Are tool calls and policy decisions logged? Can access be revoked and the agent stopped?
Evaluation quality Do tests reflect the deployed agent’s actual tools, data, and task-specific risks?

Test the deployed boundaries, not just the prompt

Build abuse cases around every external content channel the agent reads and every tool that can change state or send information. Before testing, define the legitimate task, the prohibited result, and what observable evidence would show that the prohibited result occurred.

  1. Map inputs and side effects. List the web pages, emails, files, and connector results the agent can read, plus every tool that can modify a resource or transmit data.
  2. Exercise realistic attack paths. Test direct and indirect prompt injection, harmful tool arguments, attempted data exfiltration, privilege escalation, and attempts to bypass review.
  3. Use safe test fixtures. Run with dummy data and instrumented or sandboxed tools so a successful attack cannot affect production resources.
  4. Repeat with adaptive attempts. Do not treat a pass against a fixed set of known strings as proof of safety. NIST CAISI recommends adaptive evaluations; its January 2025 experiments used then-current models and AgentDojo-derived scenarios, so their model-specific findings are not a current universal failure rate.
  5. Verify enforcement and recovery. Confirm that denied actions are blocked at execution, that logs capture decisions and tool calls, and that operators can revoke access or stop an agent.

OWASP notes that its sample smoke tests are illustrative, not a representative security benchmark. Use them as starting points, then evaluate scenarios that match your deployment’s actual tools and data. OWASP AI Agent Security Cheat Sheet

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What model defenses and benchmark scores can—and cannot—tell you

Vendor defenses and benchmark results can indicate progress on specific systems and evaluations, but they do not guarantee that an agent is safe in your deployment. A result depends on the tested model, attack setup, task, and number of attempts; your tools, data, permissions, and runtime may differ. Judge the security boundary by what the deployed system actually allows and by whether its controls hold when the model makes a bad decision.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More from the Wire

  1. World desk4 min
    How to Spot an AI Voice Scam Before Sending MoneyDon’t rely on how a caller sounds. Pause, call back through a known number, and verify the emergency with another trusted person before sending money.
  2. Mountain View desk4 min
    Google’s SynthID Detector: How to Check AI-Generated Images, Video and AudioGoogle’s SynthID Detector looks for an embedded watermark in supported images, video and audio. Here is what its results do—and do not—show.
  3. Redmond desk20 min
    How to create a link to File or Folder in Windows 11Windows 11 gives you several ways to point to a file or folder without moving or duplicating it. You can create a desktop shortcut,…
Recommended PC Tool
Recommended PC Tool
Crashes, No Sound, or Screen Glitches?Free driver scan
PC Slower Than It Used to Be?Free scan - under a minute

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.