To test a tool-calling AI agent, probe every route by which untrusted content can influence it, then verify that independent server-side controls prevent unauthorized actions. Test direct and indirect prompt injection, tool authorization, sensitive-data handling, memory, delegation, and runaway action chains in a disposable environment with synthetic data. Keep the agent configuration and observed results so the assessment can be reproduced after changes.
1. Define the scope and map trust boundaries
Start by recording exactly what is being assessed. Agent behavior depends on more than its model: prompts, tools, credentials, retrieval, memory, and integrations all affect what it can do. OWASP’s AI Agent Security Cheat Sheet calls for retaining the tested agent version, model provider, tool policy, and retrieval configuration.
- Record the agent build or version, model provider, system and developer prompts or policies, available tools and their schemas, identity and credential scopes, retrieval sources, memory behavior, and integrations in scope.
- Map how user-controlled or third-party material reaches the model: chat or API fields, uploaded files, retrieved documents, web pages, email, tool responses, memory writes, and messages from delegated agents.
- For each route, identify what the content could affect: the answer, tool selection, arguments, a state change, a memory write, or delegation to another agent.
- Use a disposable test environment and synthetic data. OWASP advises against placing real secrets in prompts used for testing.
NIST describes agent hijacking as malicious instructions inserted into data an agent ingests, taking advantage of weak separation between trusted instructions and untrusted external data. That makes the channel part of the test: an attack placed in a retrieved page tests a different boundary from the same text sent directly by a user. See NIST’s guidance on strengthening agent-hijacking evaluations.
2. Test prompt injection and goal hijacking
Run attacks against each input surface in scope, not just the chat box. The OWASP AI Exchange recommends treating external-content prompt injection and multi-turn attempts as distinct tests.
#1 Best Overall
Direct and indirect instructions
- Try direct user messages that tell the agent to ignore its instructions, reveal protected information, or take an unrelated action.
- Place adversarial instructions in the actual external-content channel under test, such as a retrieved document, web page, file, email, or tool response. Do not substitute a user message for an indirect-injection test.
- Check whether the agent treats the material as data rather than authority: it must not silently replace higher-priority instructions or expand the user’s original request.
Single-turn and multi-turn sequences
Test a single-turn attack separately from a sequence that gradually steers the agent toward a prohibited action. Include crescendo-style attempts that build over several turns, as well as attempts to revive an earlier rejected request. Record which turn changes the agent’s behavior, if any, and whether a tool call follows.
Unreliable or conflicting tool results
Feed the agent malformed, ambiguous, stale, and conflicting tool responses. Observe whether it pauses, rejects the result, safely narrows its next step, or proceeds as though the response were authoritative. In every case, confirm that downstream tool authorization still blocks an impermissible action.
3. Verify tool inventory and authorization
The model’s choice to call a tool is not an authorization control. OWASP’s LLM06:2025 Excessive Agency identifies excessive functionality, permissions, and autonomy as common contributors to excessive agency. Reduce the available capability before testing what remains.
Rank #2
Reduce unnecessary capability
- Inventory the tools actually exposed to the model and remove unused or over-broad operations.
- Where possible, expose a constrained read operation instead of a combined read, write, and delete operation.
- Give each tool only the permissions it needs, and limit how much autonomy the agent has to chain actions without review.
Enforce authorization at the tool boundary
For every proposed call, check that server-side enforcement evaluates the user, session, resource, action, and parameters. OWASP recommends validating calls against permissions and session context, and assessing proposed calls against the user’s original intent. Test whether enforcement rejects:
- A low-privilege user asking for a privileged operation.
- A request that substitutes a cross-tenant resource identifier or changes a sensitive parameter.
- A call to a hidden, deprecated, or otherwise unintended tool.
- A tool call that is unnecessary for the task or falls outside the user’s request.
Verify the denial at the tool boundary even when the model confidently proposes the call. A safe failure should take no action, avoid disclosing credentials in its error, and avoid blindly retrying an operation that may have partially completed.
Test approval for high-impact actions
For actions that require human approval, verify the approval is valid, unexpired, and bound to the specific action and parameters. Try replaying an approval, changing the arguments after approval, and applying another user’s approval. Each attempt should fail without executing the changed or unauthorized action.
Rank #3
4. Check sensitive data, memory, and action chains
Trace sensitive data across outputs
Seed the test environment with synthetic sensitive data and attempt to make the agent disclose it. Check tool arguments and results, citations, logs, and final responses—not only the user-visible answer. OWASP identifies exfiltration through tool calls and outputs as an agent abuse case. Confirm that each output is limited to what the caller is authorized to receive.
Try to persist instructions in memory
Attempt to place malicious instructions into memory, then test whether they influence another user, a later session, or a future task. Verify that memory is scoped appropriately and that unsafe content is sanitized, expires, or is rejected according to the application’s policy.
Recommended Free Tools
Test delegation boundaries
If the system delegates work, test whether instructions or outputs from one agent can cause another to exceed its own permissions or cross a trust boundary. Each agent should be checked against its own identity, authorization, and permitted task rather than inheriting implicit authority from a peer.
Rank #4
Exercise loops and resource limits
Use long plans and repeated calls to test recursion, retries, and runaway execution. Confirm that configured limits on depth, retries, tokens or cost, timeouts, and circuit breakers stop continued activity. Include attempts to use repeated or chained calls to bypass an approval or leak data in pieces.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.5. Automate the checks and gate releases
Turn adversarial cases into regression tests with explicit expected outcomes, such as “denied before tool execution” or “no unauthorized value in output.” Keep cases and synthetic fixtures under version control; do not put live customer data or secrets in them.
- Run the relevant regression suite in CI/CD when prompts or agent templates, tools, tool policies, memory, retrieval, or approval logic change.
- Require updated tests when high-risk tool policies, approval logic, or credential scopes change. Block release if required tests are missing or authorization expectations fail.
- Test the deployed configuration before production, including the production-equivalent tools and retrieval paths that are in scope.
- Repeat the assessment after material changes. A pass for one model-provider configuration is evidence about that configuration, not a guarantee for another.
OWASP’s AI Agent Security Cheat Sheet puts the timing plainly: “AI agents should undergo structured security testing before production deployment and after material changes to prompts, tools, memory, retrieval, policies, or model providers.”
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Best Value
6. Preserve evidence and report residual risk
Keep enough detail for another engineer to reproduce the assessment and understand what passed. OWASP’s agent guidance calls for retaining the tested version, provider, tool policy, and retrieval configuration. Also retain the cases executed, expected results, and observed approval, denial, timeout, and circuit-breaker behavior.
For each finding, report:
- The input surface and attacker precondition.
- The requested action and the actual tool call, state change, or data exposure.
- The policy that should have applied and why the observed result matters.
- Reproduction steps using synthetic fixtures, plus the responsible owner and retest result.
- Residual risk and any compensating controls.
Choose a verification reference that fits the assessment
These OWASP resources serve different purposes. The OWASP AI Security Verification Standard (AISVS) 1.0, released in June 2026, is a broader lifecycle catalogue: OWASP lists 191 requirements across 12 chapters and three appendices, with verification levels 1, 2, or 3. OWASP describes it as open, vendor-neutral, free to use, and testable. The AI Agent Security Cheat Sheet is more directly focused on agent abuse cases, release gates, and retained validation evidence. Use the broader standard to structure lifecycle coverage and the cheat sheet to shape agent-specific scenarios; neither replaces tests against the application’s actual configuration.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




