Driver FixRecommendedSound, Wi-Fi or graphics acting up? Check drivers firstFind missing or outdated drivers fast.Check DriversOctober DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsPC HealthRecommendedCrashes, freezes, slowdowns? Check your PC nowSpot repairable issues before they interrupt work.Check PC×
Skip to content
agentic AI

Forrester Calls Generative AI a “Chaos Agent”—What the 60% Error Claim Really Means

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Short answer: “Models are wrong 60% of the time” is not a universal accuracy rate for generative AI. The figure comes from a Columbia Journalism Review/Tow Center test in which eight AI search tools returned incorrect results for more than 60% of 1,600 article-identification queries. Forrester’s broader warning is still important: AI systems and agents can produce confident errors, act at machine speed, and turn a bad recommendation into a privileged action unless enterprises add verification, identity controls, approval gates and rollback.

The phrase “chaos agent” came from Forrester analyst Allie Mellen’s remarks at the firm’s 2025 Security and Risk Summit, as reported by VentureBeat on November 13, 2025.

What Forrester meant by “chaos agent”

Forrester was using a metaphor, not naming a formal technical classification. The concern is that generative AI combines several difficult properties:

  • It can produce plausible but false text.
  • It can repeat or amplify errors at very high speed and scale.
  • Attackers can weaponize models, prompts, documents and tools.
  • Deployments create nonhuman identities, credentials and permission paths.
  • False positives can multiply during security investigation and response.
  • A system can sound certain even when its evidence is weak or absent.

That combination changes the risk calculation. A wrong chatbot answer is a quality problem. A wrong agent action can edit a record, send a message, change code, call an inappropriate API or use authority beyond the task.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Forrester’s later guidance treats agents as engineered systems requiring orchestration, unique identities, least privilege, logging, named ownership, staged rollout, approval gates and rollback paths. The firm says three-quarters of enterprise leaders report adopting agentic AI, while meaningful production deployments beyond “agentish” chatbots remain uncommon. It also reports that 60% of enterprise generative-AI decision-makers see agentic sprawl as a challenge. See Forrester’s 2026 agentic-AI assessment and its agentic architecture guidance.

Where the “60% wrong” number came from

The number comes from a specific retrieval-and-citation test, not a general benchmark of every large language model. The Columbia Journalism Review/Tow Center study tested eight generative-search products:

  • ChatGPT Search
  • Perplexity and Perplexity Pro
  • DeepSeek Search
  • Microsoft Copilot
  • xAI’s Grok-2 and Grok-3 beta
  • Google Gemini

Researchers selected excerpts from 200 news articles—10 articles from each of 20 publishers—and ran 1,600 queries. They checked whether a system identified the correct article, publisher and URL. Across the test, the tools answered incorrectly more than 60% of the time.

Results varied sharply. Perplexity was incorrect on 37% of queries, while Grok 3 reached a 94% error rate in this test set. The study also encountered practical complications including crawler access, publisher blocking, syndicated copies and fabricated links. Those conditions are part of what was measured; they are not evidence that every model performs identically on every task.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What the statistic does not say

It does not establish that all generative AI is wrong 60% of the time. A model may be useful for rewriting text and still be unsafe for identifying a legal source, modifying production code or executing a payment. The useful question is: accurate for which task, with what evidence, under which conditions, and with what consequences if it fails?

Why confident errors matter more than obvious mistakes

The Tow Center researchers found that systems often supplied an incorrect answer instead of declining when reliable evidence was unavailable. That behavior is more dangerous than a visible failure because polished prose, a citation-shaped link or an authoritative tone can be mistaken for verification.

In a research workflow, a wrong URL can send an analyst to the wrong article. In an operational workflow, an agent can use that wrong conclusion to select a customer record, change a setting or trigger another tool. Users cannot provide meaningful oversight if the system hides uncertainty and the evidence trail is incomplete.

How chatbot errors differ from agent failures

Chatbot or search error

The system returns a wrong answer, summary, citation or URL. The user may catch it before anything changes outside the conversation.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Agent failure

An agent must plan and execute a sequence across browsers, code, programs, data and people. It may edit the wrong record, call the wrong API, escalate privileges, stop before completing the objective, duplicate work, enter a loop or leave an unrecoverable state. The failure can therefore involve both reasoning and control of external systems.

What the other studies show—and what they do not

Study What was tested Reported result Proper interpretation
Tow Center/CJR AI-search article identification and citation More than 60% incorrect overall Task-specific retrieval and citation failure, not a universal AI accuracy rate
AgentCompany Agents completing professional tasks in a simulated software company VentureBeat reported about 24% autonomous completion by top systems; failure reached 70%–90% as complexity increased Long, multi-tool tasks remain difficult under the benchmark’s task design and agent setup
Salesforce-related research CRM-oriented agent tasks 62% baseline-task failure in the cited research A result for that agent and task configuration, not all enterprise agents
Veracode Security of generated Java, Python, C and JavaScript 45% of tested samples introduced an OWASP Top 10 vulnerability A security-testing result, not a population estimate for all production code

AgentCompany evaluated 175 professional tasks involving web browsing, coding, programs and coworkers. Its results support caution about long-horizon work, but do not prove that every commercial agent fails 70%–90% of the time, that every task is equally hard, or that humans are always more accurate or cheaper. A simulated software company is not a direct forecast for healthcare, finance, government or manufacturing.

Veracode’s program covered 80 coding tasks, four languages and more than 100 models, with testing against OWASP Top 10 categories. VentureBeat reported language-specific security pass rates of 28.5% for Java, 55.3% for Python, 57.3% for C and 61.7% for JavaScript. These figures describe Veracode’s test design and vulnerability definitions, not all AI-written software.

Why guardrails can reduce task performance

Permissions, data filters, output constraints and approval rules can block information or actions an agent needs. Completion may therefore fall when safety controls are added, as illustrated by the Salesforce-related result. That does not show that guardrails are ineffective. It shows that task design, context, tool schemas and recovery paths must be designed together.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The right response is to make constraints explicit and testable: define what the agent may read, which tools it may call, which actions require approval, what evidence it must provide and how it should recover when a safe operation is blocked.

Why agents are an identity-security problem

An agent may hold API keys, OAuth tokens, certificates or service-account permissions. It may read internal data, invoke tools, create or modify records, or delegate work to another agent. The central security question is therefore not only whether the model gives a bad answer; it is whether that answer can become a privileged action.

  • Give every agent a distinct identity—never a shared administrator credential.
  • Use least privilege and short-lived tokens where possible.
  • Assign a named human owner, purpose, start date and retirement date.
  • Log credentials, data access, tool calls and delegated actions.
  • Maintain an inventory of agents, models, vendors, tools and environments.

A practical deployment framework

1. Classify the consequence of failure

  • Low: brainstorming, summarization and draft generation.
  • Moderate: internal recommendations, code suggestions and customer-service drafts.
  • High: payments, access changes, production releases, legal or medical decisions, safety operations and externally binding communications.

As consequence and irreversibility rise, reduce autonomy and increase evidence and approval requirements.

2. Start with bounded workflows

Define a narrow objective, approved inputs, limited tools, measurable success criteria and a human approval step before an external or irreversible action. Read-only search with mandatory citations is a safer starting point than unrestricted write access.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

3. Instrument the complete chain

Record the task request, retrieved context, model and version, tools called, data accessed, output, approvals, final action, errors and rollback events. Logging only the final answer leaves the most important evidence missing.

4. Evaluate the system, not just the model

  • Retrieval and citation accuracy
  • Refusal and uncertainty behavior
  • Prompt-injection and indirect-injection resistance
  • Tool-selection accuracy and permission boundaries
  • Long-horizon completion and recovery after tool failure
  • Data leakage and regression after model, prompt, index or tool changes

5. Add approval and recovery controls

Use deterministic policy or an authorized person to approve payments, access changes, production deployments and legally significant communications. Add execution budgets, kill switches, checkpoints and rollback paths so a bad action is containable.

6. Red-team the model-plus-tools system

Test malicious documents, prompt injection, data exfiltration, privilege escalation, tool abuse, fabricated citations and cascading multiagent failure. A model safety test without its real tools and permissions is incomplete.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Where AI is a reasonable fit—and where it is not

Good initial candidates

  • Drafting internal documents and summarizing low-risk material
  • Classifying or routing work with human review
  • Generating test cases and code suggestions followed by review and security scanning
  • Searching a controlled knowledge base when citations are mandatory
  • Recommending operational actions without executing them

Poor initial candidates for unsupervised deployment

  • Moving money or changing access permissions
  • Direct production deployment
  • Medical, legal, employment, credit, insurance or safety decisions
  • Deleting or altering records without recovery
  • Broad access to confidential data with no clearly bounded purpose
  • Multi-tool chains with no execution budget or approval boundary

What enterprises should measure before scaling

Compare model and infrastructure cost with evaluation, monitoring, security tooling, human review, integration, data-cleaning and incident-response costs. Also price false positives, false negatives and the operational cost of blocking unsafe automation. A low per-token price does not mean a low total cost of ownership.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For buying decisions, match the control gap to the product category. Application-security tools such as Veracode or GitHub Advanced Security can scan code, but they do not establish factual accuracy or agent authorization. Salesforce Agentforce may suit a Salesforce-centered CRM workflow, while cross-vendor deployments need identity, authorization, evaluation, logging and runtime controls. Forrester’s AEGIS guidance can help define governance and architecture requirements, but advisory research is not a substitute for enforcement, scanning or red-team testing. Public prices and promotional terms change, so verify them before purchase.

The defensible conclusion

Forrester’s warning is directionally right but the headline compresses unrelated measurements. The Tow Center’s more-than-60% result concerns article-identification and citation queries; the agent, CRM and code-security percentages measure different tasks, systems and definitions of failure.

Do not ask whether “AI” has one fixed error rate. Ask whether a particular system can produce evidence, stay within least-privilege boundaries, survive adversarial input, expose its actions and recover when it fails. Generative AI can be useful when autonomy is bounded and outputs are verifiable. It becomes a security liability when confident guesses are connected directly to sensitive data, powerful tools and irreversible actions.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Read next

Recommended PC Tool
Recommended PC Tool
Outdated Drivers Are Slowing You DownFree scan - exact matches
Windows Errors? Fix Them Before They SpreadFree repair scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.