To secure an AI agent, limit what it can access and what it can do, keep untrusted content away from privileged instructions, and require a real authorization check before consequential actions. Prompt-injection detection, model training, guardrails, and human review can reduce risk, but none makes an agent safe by itself.
OpenAI’s guidance points to a practical shift: treat agent security as a problem of controlling authority and information flow, not simply spotting suspicious text. That applies whether you are building an agent with OpenAI tools or designing a broader system that reads web pages, documents, messages, or other outside content.
Why prompt injection is an agent security problem
Prompt injection happens when a third party puts malicious instructions in content an agent may encounter—for example, a web page or document. The content can try to redirect the agent, expose information, or induce it to use a tool. Because the agent may also have access to data and capabilities, the issue is not just whether the model recognizes a hostile instruction. It is what that instruction could reach if it influences the agent.
OpenAI’s March 11, 2026 article, Designing AI agents to resist prompt injection, frames the problem in terms of a source and a sink. A source is content that an attacker can influence; a sink is a capability that could cause harm in context, such as sending sensitive information to a third party or invoking a tool. The source might be an ordinary-looking page, while the sink might be a tool with permission to edit records or send messages.
#1 Best Overall
This framing explains why a suspicious-string filter is not a complete defense. OpenAI notes that real-world attempts can resemble social engineering. An agent may be influenced by instructions that are not easy to classify as a familiar attack phrase. Reducing the agent’s authority and limiting data flows can constrain the impact even when detection fails.
OpenAI’s stated design goal is to preserve the expectation that “potentially dangerous actions, or transmissions of potentially sensitive information, should not happen silently or without appropriate safeguards.” The article is by Thomas Shadwell and Adrian Spânu; this is a design goal, not a guarantee that a silent action is impossible.
What to contain: data, execution, and authority
An agent’s risk depends on its full operating environment—not just its system prompt or model. Map the path from outside content to a possible action, then restrict each point where data or authority can cross a boundary.
Rank #2
- Data: What can the agent read, including local files, conversation history, connected accounts, and retrieved documents?
- Execution: What code can it run, and what filesystem, credentials, and network access does that environment expose?
- Tools: Can it only read, or can it also write, send, delete, purchase, or change access?
- Identity: Whose permissions does a tool use, and can the application limit them to this task?
- Side effects: Which actions are reversible, and which need approval before they happen?
- Visibility: Can operators or users inspect what the agent tried to do and what data it proposed to send?
OpenAI’s practical guide recommends assessing tools by factors such as read versus write access, reversibility, account permissions, and financial impact. Use those factors to decide where to add tighter limits, a second authorization check, or human review; do not treat all tools as equally risky.
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Repair Windows errors before they cause bigger problemsFix Now →Scan for outdated or missing drivers - takes under a minuteDriver Scan →How to contain agent execution and credentials
Isolate model-directed work
Run model-directed code in isolated compute and separate workloads or users that must not share data. Set an explicit outbound network policy that allows only the endpoints the task needs, and account for both local and remote tools. OpenAI’s sandbox guidance emphasizes that generated code can access the files, credentials, and network available to its environment. A sandbox is therefore a security boundary only to the extent that you limit what enters it and what it can reach.
Keep credentials outside the agent’s execution environment
Do not place application keys where model-directed code can read them. For third-party credentials, broker access through a trusted server or proxy, or use an applicable documented vault pattern. A secrets manager does not protect a secret once it is handed directly to code running in an untrusted environment. If you suspect exposure, revoke or rotate the affected credential promptly.
Limit data crossing the boundary
Give an agent only the material required for its task. Avoid making an entire filesystem, account, or customer record available when a narrower view will do. Treat retrieved content as untrusted even if it comes from a source your application normally uses; its origin does not make its embedded instructions authoritative.
How to control instructions and tool calls
OpenAI’s agent safety guidance recommends keeping untrusted input in user messages rather than placing it in privileged developer instructions. It also recommends structured outputs—such as enums or validated JSON—to constrain what one workflow stage can pass to another. Structure can reduce uncontrolled instruction flow, but it does not establish that the content is safe or make prompt injection impossible.
Validate the structure and values at the receiving boundary. Before a tool call, check that the requested operation is permitted for the current user, task, and resource. After a tool call, inspect the result before using it to make another decision or send information elsewhere. Keep MCP approvals enabled where applicable, and use input checks, trace graders, and evaluations to find failures. OpenAI cautions that guardrail nodes alone are not foolproof.
Rank #4
Use separate checks for separate jobs
| Control | What it does | What it does not replace |
|---|---|---|
| Guardrail | Automatically checks input, output, or tool behavior against configured criteria. | Authorization, isolation, or a guaranteed block on every harmful action. |
| Human review | Pauses a run so a person or policy can approve or reject a sensitive action before it proceeds. | Least-privilege access or a check that the action is authorized for the affected account. |
| Authentication and authorization | Establishes who is acting and which resources or operations they may use. | Monitoring, review of proposed actions, or containment if an agent is misled. |
| Logging and evaluation | Helps investigate actions and test how the system behaves across scenarios. | Prevention of an unsafe action on its own. |
The Agents SDK guidance distinguishes automatic guardrails from human review: review pauses a run for a person or policy to approve or reject a sensitive step. Put that pause at the action boundary in the application or harness. Do not assume the model will independently ask for approval at the right moment.
When to require approval
Require a blocking approval before actions whose effects matter, especially when they are hard to reverse or affect another person, account, or organization. OpenAI’s examples include cancellations, edits, shell commands, and sensitive MCP actions. A confirmation is useful only if it comes before the side effect and gives the reviewer enough information to judge it.
- Show the exact operation, target, and consequential parameters—not a vague summary such as “continue.”
- Check the acting user’s authorization independently of the model’s proposal.
- Keep the approval separate from the agent’s ability to grant itself more access.
- Record the proposal, decision, actor, and outcome in an audit trail.
- Make safe rejection and recovery possible when an action is denied or interrupted.
These are implementation practices, not a claim that any approval flow eliminates risk. An approved action can still be mistaken, and a compromised or poorly designed approval surface can undermine the check.
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallCrashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteWhat OpenAI says it does in ChatGPT
OpenAI describes layered safeguards for ChatGPT that include training, monitoring, link checks, sandboxing, red-teaming, and user controls. In its March 2026 design article, it says a mechanism called Safe Url can detect a proposed transmission of conversation information to a third party and, in rare cases where the model is convinced, show the information to the user for confirmation or block it. The article also says Canvas and ChatGPT Apps run in a sandbox designed to detect unexpected communications and request consent.
Those descriptions concern OpenAI’s own products and systems. They are not default protections for every agent built with an API, nor evidence that equivalent safeguards exist in a custom application. A developer still needs to design and test the application’s own execution, credential, network, authorization, and review boundaries.
The ChatGPT agent Help Center page describes high-impact-action confirmations, refusal patterns, prompt-injection monitoring, and watch mode requiring supervision on certain sites. It also warns that using websites or apps can expose sensitive material and that safeguards do not remove all risk. For people using ChatGPT agent, OpenAI advises enabling only needed apps, considering the sensitivity of logged-in sites, avoiding unnecessary sensitive inputs, and giving specific rather than broad prompts. Product behavior can change, so check the current Help Center guidance for the controls available to your account.
How to evaluate an agent design
Before deployment, walk through the following checks for each workflow—not only for the agent as a whole:
- Trace an attack path. Identify content an attacker can influence, data the agent can access, and tools that could move or alter that data.
- Reduce access. Remove unnecessary files, account permissions, tools, and network routes; separate workloads that must not share data.
- Secure credentials. Keep secrets out of model-directed execution and broker access through a trusted component with narrowly scoped permissions.
- Validate handoffs. Pass untrusted material as data, use constrained schemas between stages, and validate tool inputs and outputs.
- Gate side effects. Require a blocking review for sensitive or high-impact operations, with details sufficient to approve or reject the exact action.
- Test and observe. Evaluate hostile and ambiguous inputs, inspect traces, and retain enough audit information to investigate failures.
- Plan recovery. Know how to stop a run, revoke credentials, undo reversible changes, and respond to suspected data exposure.
OpenAI’s 2026 design article reports a 50% success rate for one particular test prompt in a reported 2025 prompt-injection example. That result is tied to the specific example described in the article; it is not an estimate of attack prevalence or a general measure of defense effectiveness. The official sources discussed here do not establish a comparable industry-wide attack rate or a general effectiveness percentage for these controls.
The practical rule
Build the system so that an agent influenced by hostile content still lacks the access, credentials, network path, or unreviewed authority needed to cause serious harm. Use model-level safeguards and detection as additional layers, then enforce permissions and approval in the software that actually performs the action.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




