Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

A July 2023 study showed that an algorithm could generate odd-looking strings that, when appended to certain harmful prompts, made several aligned chatbots more likely to answer instead of refuse. The striking finding was not a magic phrase anyone could invent by guesswork: it was that automated search could produce adversarial suffixes that transferred from open models to selected commercial chatbots.

That is a meaningful weakness in relying on a chatbot’s refusal as a security boundary. It does not show that every user can reliably defeat every chatbot, that the providers’ systems were breached, or that the original strings still work on today’s models.

What did the researchers discover?

The headline refers to “Universal and Transferable Adversarial Attacks on Aligned Language Models,” a paper posted in July 2023 by researchers associated with Carnegie Mellon University, the Center for AI Safety, Google DeepMind, and the Bosch Center for AI. The study described an adversarial suffix: a sequence of tokens added to the end of a request that was designed to steer a model away from its trained refusal behavior. Read the paper.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The strings could look nonsensical to a person. They were not hand-written magic words. Researchers used a combination of greedy and gradient-based search against open models to find token sequences that could encourage a model to produce a response to requests it would otherwise refuse. Carnegie Mellon’s announcement describes the optimization approach.

The researchers reported that some suffixes transferred to other models, including public interfaces for ChatGPT, Google Bard and Claude, as well as multiple open models. “Universal” in the paper’s title refers to suffixes that could work across multiple prompts or targets in the tested setting—not every request, every model, or every version. The CMU research overview summarizes the finding.

The work tested whether models would produce content they were meant to refuse, including dangerous or otherwise harmful assistance. That demonstrates a failure of policy compliance, not that a chatbot could independently carry out an act. Generated text may also be inaccurate, incomplete or nonsensical even when a refusal is bypassed. This article does not reproduce operational suffixes or harmful instructions.

Why could a string found on one model affect another?

Language models break incoming text into tokens and use learned patterns to predict what comes next. The researchers optimized candidate token sequences against an open model, searching for suffixes that shifted the model toward a more compliant continuation. Some learned patterns and design characteristics are shared across language models, which can help an input found on one model generalize to another.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

That is a plausible explanation for transfer, not proof that the models are identical or that every suffix works equally well. A provider can change the model, tokenizer, system instructions, moderation layers or interface, any of which may alter the result. Performance can also vary with the prompt and target behavior.

The reported transfer was significant because the researchers did not need access to the internal weights of each commercial system to demonstrate that some suffixes found elsewhere could affect its public interface. It did not establish that the same attack would work through every API, consumer interface or later model release.

How “easy” was the attack?

It helps to separate finding a suffix from using one. Once someone has a candidate string, adding it to a prompt may be straightforward. Discovering effective strings was a technical optimization task that required model access, computation and expertise; it was not simply a matter of casually trying a few phrases.

  • Researchers selected the target behaviors, models and evaluation process and conducted the safety research.
  • The optimization procedure searched candidate token sequences automatically rather than relying on a person to invent every suffix.
  • A potential attacker could reuse or adapt a discovered suffix, although success would depend on the target system and could change after updates.
  • A provider could respond with model changes, filtering, monitoring or other controls, but patching one visible string would not necessarily address the broader search technique.

CMU described the method as enabling a “virtually unlimited number” of attacks. That is the researchers’ characterization of the method’s ability to generate variants, not a measured count of attacks that work on every service.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What does “guardrail” mean?

A chatbot’s safety controls are usually a set of layers rather than one switch. Depending on the product, they may include:

  • Post-training alignment: Training intended to make the model refuse certain requests or follow desired behavior.
  • System instructions: Higher-priority directions that shape how the model should respond.
  • Input moderation: Screening prompts before the model processes them.
  • Output moderation: Checking generated responses before they reach the user.
  • Runtime controls: Rate limits, abuse monitoring, account restrictions and human review.
  • Application controls: Rules added by the developer using a model, such as limits on data access or tool use.

A model may refuse ordinary harmful requests and still be vulnerable to adversarial input. Conversely, a failure in one layer does not mean every other layer has failed. Safety depends on how these controls work together in a particular product.

Is a jailbreak the same as hacking a chatbot?

Usually, no. A jailbreak manipulates a model into violating its intended behavior—for example, through adversarial tokens, conflicting instructions, role-play or pressure across multiple turns. Related risks include prompt injection, where instructions hidden in a document or webpage influence a model processing that material, and tool-use manipulation, where an agent is coaxed into misusing an available capability.

That is different from gaining unauthorized access to a provider’s servers, stealing an account or modifying the model’s weights. Unless there is evidence of such a compromise, “jailbreak” or “adversarial prompt attack” is more precise than “hack.”

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Why does the issue matter beyond a chatbot’s answer?

A jailbreak that produces harmful text is a policy failure, but text generation alone does not mean a system has acted in the world. The stakes can rise when a model is connected to tools, private information, code execution, email, financial systems or automated workflows. A manipulated model with permission to take consequential actions can create risks that a text-only assistant does not have.

That is why the application’s permissions and supervision matter alongside the model’s alignment. Useful safeguards include limiting a model to the data and tools it needs, sandboxing execution, monitoring activity and requiring human confirmation for high-impact actions. A refusal policy cannot substitute for access control.

There are trade-offs. Aggressive filters can block legitimate medical, educational, cybersecurity or research requests, while permissive systems can leave room for abuse. Evaluation should consider the application’s actual uses, the consequences of a failure and whether controls can detect or contain one.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Does the original 2023 attack still work on chatbots in 2026?

The available evidence does not establish that the exact 2023 suffixes remain effective against current ChatGPT, Gemini or Claude versions. The paper tested models and interfaces available at the time; providers have since had opportunities to change models and safety layers. Treat the named products as historical test targets, not a current compatibility list.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Jailbreak research has continued. A 2025 Microsoft Research project studied methods for improving automated jailbreak generation against more strongly aligned models, and a separate 2025 paper proposed a defense aimed at adversarial suffixes. Their results concern the methods and evaluation conditions in those papers; neither makes any one attack or defense universally effective. See the Microsoft Research paper and the suffix-defense proposal.

A 2026 study also reported broad bypasses across contemporary open models under its tested conditions. That is evidence of an ongoing research problem, not proof that every consumer chatbot can be bypassed in every context. Read the 2026 study.

What should AI developers and organizations do?

No single filter, model update or vendor product should be treated as a permanent fix. A practical defense combines model-level measures with controls around the application:

  • Test for adversarial prompts and prompt injection across the models and interfaces actually deployed, then repeat testing after significant updates.
  • Use input and output checks as additional layers, not as substitutes for access controls.
  • Apply least privilege: give a model only the tools, records and permissions needed for its task.
  • Sandbox code and other tool actions, and require human approval before consequential or hard-to-reverse operations.
  • Monitor for abuse and unusual behavior, with a process to investigate incidents and adjust controls.
  • Measure false positives as well as attack resistance so that legitimate use is not needlessly blocked.

Organizations evaluating security products should ask which model families and attack types are covered, where scanning happens, how false positives are handled, what activity can be logged, and how the system integrates with existing identity, monitoring and data-protection controls. Vendor claims and benchmark scores should be considered in light of the tested setup rather than as guarantees.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.