Short answer: “Models are wrong 60% of the time” is not a universal accuracy rate for generative AI. The figure comes from a Columbia Journalism Review/Tow Center test in which eight AI search tools returned incorrect results for more than 60% of 1,600 article-identification queries. Forrester’s broader warning is still important: AI systems and agents can produce confident errors, act at machine speed, and turn a bad recommendation into a privileged action unless enterprises add verification, identity controls, approval gates and rollback.
The phrase “chaos agent” came from Forrester analyst Allie Mellen’s remarks at the firm’s 2025 Security and Risk Summit, as reported by VentureBeat on November 13, 2025.
What Forrester meant by “chaos agent”
Forrester was using a metaphor, not naming a formal technical classification. The concern is that generative AI combines several difficult properties:
- It can produce plausible but false text.
- It can repeat or amplify errors at very high speed and scale.
- Attackers can weaponize models, prompts, documents and tools.
- Deployments create nonhuman identities, credentials and permission paths.
- False positives can multiply during security investigation and response.
- A system can sound certain even when its evidence is weak or absent.
That combination changes the risk calculation. A wrong chatbot answer is a quality problem. A wrong agent action can edit a record, send a message, change code, call an inappropriate API or use authority beyond the task.
#1 Best Overall
Forrester’s later guidance treats agents as engineered systems requiring orchestration, unique identities, least privilege, logging, named ownership, staged rollout, approval gates and rollback paths. The firm says three-quarters of enterprise leaders report adopting agentic AI, while meaningful production deployments beyond “agentish” chatbots remain uncommon. It also reports that 60% of enterprise generative-AI decision-makers see agentic sprawl as a challenge. See Forrester’s 2026 agentic-AI assessment and its agentic architecture guidance.
Where the “60% wrong” number came from
The number comes from a specific retrieval-and-citation test, not a general benchmark of every large language model. The Columbia Journalism Review/Tow Center study tested eight generative-search products:
- ChatGPT Search
- Perplexity and Perplexity Pro
- DeepSeek Search
- Microsoft Copilot
- xAI’s Grok-2 and Grok-3 beta
- Google Gemini
Researchers selected excerpts from 200 news articles—10 articles from each of 20 publishers—and ran 1,600 queries. They checked whether a system identified the correct article, publisher and URL. Across the test, the tools answered incorrectly more than 60% of the time.
Results varied sharply. Perplexity was incorrect on 37% of queries, while Grok 3 reached a 94% error rate in this test set. The study also encountered practical complications including crawler access, publisher blocking, syndicated copies and fabricated links. Those conditions are part of what was measured; they are not evidence that every model performs identically on every task.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Scan for outdated or missing drivers - takes under a minute3Repair Windows errors before they cause bigger problemsWhat the statistic does not say
It does not establish that all generative AI is wrong 60% of the time. A model may be useful for rewriting text and still be unsafe for identifying a legal source, modifying production code or executing a payment. The useful question is: accurate for which task, with what evidence, under which conditions, and with what consequences if it fails?
Why confident errors matter more than obvious mistakes
The Tow Center researchers found that systems often supplied an incorrect answer instead of declining when reliable evidence was unavailable. That behavior is more dangerous than a visible failure because polished prose, a citation-shaped link or an authoritative tone can be mistaken for verification.
Rank #2
In a research workflow, a wrong URL can send an analyst to the wrong article. In an operational workflow, an agent can use that wrong conclusion to select a customer record, change a setting or trigger another tool. Users cannot provide meaningful oversight if the system hides uncertainty and the evidence trail is incomplete.
How chatbot errors differ from agent failures
Chatbot or search error
The system returns a wrong answer, summary, citation or URL. The user may catch it before anything changes outside the conversation.
Agent failure
An agent must plan and execute a sequence across browsers, code, programs, data and people. It may edit the wrong record, call the wrong API, escalate privileges, stop before completing the objective, duplicate work, enter a loop or leave an unrecoverable state. The failure can therefore involve both reasoning and control of external systems.
What the other studies show—and what they do not
| Study | What was tested | Reported result | Proper interpretation |
|---|---|---|---|
| Tow Center/CJR | AI-search article identification and citation | More than 60% incorrect overall | Task-specific retrieval and citation failure, not a universal AI accuracy rate |
| AgentCompany | Agents completing professional tasks in a simulated software company | VentureBeat reported about 24% autonomous completion by top systems; failure reached 70%–90% as complexity increased | Long, multi-tool tasks remain difficult under the benchmark’s task design and agent setup |
| Salesforce-related research | CRM-oriented agent tasks | 62% baseline-task failure in the cited research | A result for that agent and task configuration, not all enterprise agents |
| Veracode | Security of generated Java, Python, C and JavaScript | 45% of tested samples introduced an OWASP Top 10 vulnerability | A security-testing result, not a population estimate for all production code |
AgentCompany evaluated 175 professional tasks involving web browsing, coding, programs and coworkers. Its results support caution about long-horizon work, but do not prove that every commercial agent fails 70%–90% of the time, that every task is equally hard, or that humans are always more accurate or cheaper. A simulated software company is not a direct forecast for healthcare, finance, government or manufacturing.
Veracode’s program covered 80 coding tasks, four languages and more than 100 models, with testing against OWASP Top 10 categories. VentureBeat reported language-specific security pass rates of 28.5% for Java, 55.3% for Python, 57.3% for C and 61.7% for JavaScript. These figures describe Veracode’s test design and vulnerability definitions, not all AI-written software.
Why guardrails can reduce task performance
Permissions, data filters, output constraints and approval rules can block information or actions an agent needs. Completion may therefore fall when safety controls are added, as illustrated by the Salesforce-related result. That does not show that guardrails are ineffective. It shows that task design, context, tool schemas and recovery paths must be designed together.
Free tools Windows power users keep installed
One-click scans. No signup required.
Rank #3
The right response is to make constraints explicit and testable: define what the agent may read, which tools it may call, which actions require approval, what evidence it must provide and how it should recover when a safe operation is blocked.
Why agents are an identity-security problem
An agent may hold API keys, OAuth tokens, certificates or service-account permissions. It may read internal data, invoke tools, create or modify records, or delegate work to another agent. The central security question is therefore not only whether the model gives a bad answer; it is whether that answer can become a privileged action.
- Give every agent a distinct identity—never a shared administrator credential.
- Use least privilege and short-lived tokens where possible.
- Assign a named human owner, purpose, start date and retirement date.
- Log credentials, data access, tool calls and delegated actions.
- Maintain an inventory of agents, models, vendors, tools and environments.
A practical deployment framework
1. Classify the consequence of failure
- Low: brainstorming, summarization and draft generation.
- Moderate: internal recommendations, code suggestions and customer-service drafts.
- High: payments, access changes, production releases, legal or medical decisions, safety operations and externally binding communications.
As consequence and irreversibility rise, reduce autonomy and increase evidence and approval requirements.
2. Start with bounded workflows
Define a narrow objective, approved inputs, limited tools, measurable success criteria and a human approval step before an external or irreversible action. Read-only search with mandatory citations is a safer starting point than unrestricted write access.
3. Instrument the complete chain
Record the task request, retrieved context, model and version, tools called, data accessed, output, approvals, final action, errors and rollback events. Logging only the final answer leaves the most important evidence missing.
4. Evaluate the system, not just the model
- Retrieval and citation accuracy
- Refusal and uncertainty behavior
- Prompt-injection and indirect-injection resistance
- Tool-selection accuracy and permission boundaries
- Long-horizon completion and recovery after tool failure
- Data leakage and regression after model, prompt, index or tool changes
5. Add approval and recovery controls
Use deterministic policy or an authorized person to approve payments, access changes, production deployments and legally significant communications. Add execution budgets, kill switches, checkpoints and rollback paths so a bad action is containable.
Rank #4
6. Red-team the model-plus-tools system
Test malicious documents, prompt injection, data exfiltration, privilege escalation, tool abuse, fabricated citations and cascading multiagent failure. A model safety test without its real tools and permissions is incomplete.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Where AI is a reasonable fit—and where it is not
Good initial candidates
- Drafting internal documents and summarizing low-risk material
- Classifying or routing work with human review
- Generating test cases and code suggestions followed by review and security scanning
- Searching a controlled knowledge base when citations are mandatory
- Recommending operational actions without executing them
Poor initial candidates for unsupervised deployment
- Moving money or changing access permissions
- Direct production deployment
- Medical, legal, employment, credit, insurance or safety decisions
- Deleting or altering records without recovery
- Broad access to confidential data with no clearly bounded purpose
- Multi-tool chains with no execution budget or approval boundary
What enterprises should measure before scaling
Compare model and infrastructure cost with evaluation, monitoring, security tooling, human review, integration, data-cleaning and incident-response costs. Also price false positives, false negatives and the operational cost of blocking unsafe automation. A low per-token price does not mean a low total cost of ownership.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
For buying decisions, match the control gap to the product category. Application-security tools such as Veracode or GitHub Advanced Security can scan code, but they do not establish factual accuracy or agent authorization. Salesforce Agentforce may suit a Salesforce-centered CRM workflow, while cross-vendor deployments need identity, authorization, evaluation, logging and runtime controls. Forrester’s AEGIS guidance can help define governance and architecture requirements, but advisory research is not a substitute for enforcement, scanning or red-team testing. Public prices and promotional terms change, so verify them before purchase.
The defensible conclusion
Forrester’s warning is directionally right but the headline compresses unrelated measurements. The Tow Center’s more-than-60% result concerns article-identification and citation queries; the agent, CRM and code-security percentages measure different tasks, systems and definitions of failure.
Do not ask whether “AI” has one fixed error rate. Ask whether a particular system can produce evidence, stay within least-privilege boundaries, survive adversarial input, expose its actions and recover when it fails. Generative AI can be useful when autonomy is bounded and outputs are verifiable. It becomes a security liability when confident guesses are connected directly to sensitive data, powerful tools and irreversible actions.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Repair Windows errors before they cause bigger problemsFix Now →




