Yes, with conditions. AI models can generate useful leads in software security work, and on constrained benchmarks they can score far higher when they are given the right tools. What the published evidence does not show is that a model, on its own, reliably finds real and exploitable flaws in arbitrary software. A suspicious code snippet, a benchmark score or a crash is a starting point. It is not a confirmed vulnerability.
This article walks through what the main public evaluations measured, how much weight each result can bear, and where the risks sit. Most of the primary sources date from 2024, plus one recent developer-published system card, so read the figures as snapshots of specific models and setups, not as a statement about every system available today.
As an Amazon Associate I earn from qualifying purchases.
What “help” can mean: five different success levels
Most confusion about AI vulnerability discovery comes from the phrase “found a vulnerability” covering very different outcomes. From weakest to strongest evidence:
- Flagged suspicious code. A model points at a function and says it looks unsafe. This is cheap to produce and is often wrong.
- Reproduced bug. A crash or sanitizer report can be triggered again on demand.
- Verified security impact. Someone other than the model, or a verification system, confirms the bug matters for security.
- Controlled exploit primitive. The bug can be used to gain a specific, controlled capability, such as a constrained memory read or write.
- End-to-end exploit. A complete chain from the bug to a meaningful outcome on the target.
OpenAI’s GPT-5.6 system card draws this line explicitly. Its long-horizon evaluation treats crashes and sanitizer findings as leads. Stronger evidence requires reproducible artifacts, controls, and verifier-owned proof of impact or a controlled exploitability primitive. When you read any claim about AI “finding” bugs, ask which of the five levels it reached.
#1 Best Overall
What the public evidence shows
Meta’s CyberSecEval 2: models differ, and coding ability matters
Meta’s GenAI Cybersec Team published CyberSecEval 2 on 18 April 2024. It is a benchmark suite for LLM security risks and capabilities. It adds prompt-injection and code-interpreter-abuse tests to earlier insecure-code testing, and it quantifies vulnerability-exploitation tasks across several contemporary models.
On capability, the authors report that models with coding ability did better than models without it, and that further work was still needed before models could generate exploits proficiently. On safety, they report a trade-off: conditioning a model to refuse unsafe prompts can also cause false refusals of benign requests. That matters for defenders, whose legitimate requests often resemble attacker requests. On prompt injection, they write:
“Our results show conditioning away risk of attack remains an unsolved problem; for example, all tested models showed between 25% and 50% successful prompt injection tests.”
The Tool Desk
Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
That 25%–50% figure is a benchmark result for the tested models. It is not the rate at which real-world attacks succeed against deployed products.
Google Project Zero’s Naptime: the harness changes the result
Project Naptime, published by Google Project Zero (Sergei Glazunov and Mark Brand) on 20 June 2024, is the clearest evidence that how you wrap a model matters. The team built an LLM-driven vulnerability research framework around a few design principles:
- room for the model to reason before acting;
- an interactive program environment;
- specialized tools, including a debugger and scripting;
- automatic verification of results;
- independent trajectories, so the model can explore several hypotheses.
On CyberSecEval 2’s memory-safety tasks, the framework raised performance by up to 20 times compared with the original paper’s reported results. Buffer Overflow scores went from 0.05 to 1.00. Advanced Memory Corruption scores went from 0.24 to 0.76. The team’s explanation for the interactivity effect is:
“Interactivity within the program environment is essential, as it allows the models to adjust and correct their near misses, a process demonstrated to enhance effectiveness in tasks such as software development.”
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Two cautions apply. First, these are scores on specific benchmark tasks, not a multiplier for real-world productivity. Second, the authors themselves said substantial progress was still needed before such tools could meaningfully change security researchers’ daily work. The practical lesson is that results describe a model plus a harness, not an unaided chat model, so a headline score for one setup says little about a plain prompt in a chat window.
Rank #3
IBM Research: identifying a bug is not the same as reasoning about it
An IBM Research conference-paper record from 20 May 2024, titled “LLMs Cannot Reliably Identify and Reason About Security Vulnerabilities (Yet?)”, describes a study of 228 code scenarios, eight LLMs and eight investigative dimensions. Its framing is reliability: does a model give consistent, sound answers across vulnerable and patched code, different prompts and varied scenarios? That is a harder test than spotting a bug once.
The numbers describe the study design. They are not an estimate for all code or all models, and the title’s conclusion applies to the systems tested in 2024. It would be wrong to read it as proof that today’s models fail, and equally wrong to ignore it as a reminder that one correct answer does not establish reliability.
OpenAI’s GPT-5.6 system card: the most recent view, from the developer
OpenAI’s Deployment Safety Hub publishes the GPT-5.6 system card, including threat-modeling and cybersecurity evaluations. It is a developer reporting on its own model, so read it with its stated scope in mind. It describes two evaluations relevant here.
- CVE-Bench version 1.0. The model is asked to identify and exploit vulnerabilities in sandboxed web applications. OpenAI says infrastructure problems prevented running all 40 challenges, so it ran 34. It used a zero-day prompt configuration (the model is not told which vulnerability to exploit), withheld application source code, and measured pass@1 over three rollouts.
- VulnLMP. A longer-horizon evaluation against real, widely deployed software, using source-available targets and a research harness. OpenAI reports credible memory-safety leads, reproducible crashes, root-cause analyses and, in some of the strongest runs, controlled exploitation primitives. It also reports that GPT-5.6 Sol did not independently produce a functional full-chain exploit or a verifier-confirmed Critical-level outcome against real-world targets in this evaluation.
OpenAI’s own interpretation is:
“This suggests that substantial parts of real world vulnerability research are becoming increasingly automatable when models are paired with tool use, build systems, and verification infrastructure.”
Note the conditions in that sentence, and the boundary in the results. Real progress on leads, crashes and primitives sits alongside no verified end-to-end critical outcome. OpenAI also acknowledges limits in its CTF, CVE-Bench and Cyber Range coverage and says strong scores alone do not establish high cyber capability.
The four evaluations side by side
| Evaluation | Source and date | Task | Setup | Reported result | What it does not show |
|---|---|---|---|---|---|
| CyberSecEval 2 | Meta, April 2024 | Insecure code, exploitation tasks, prompt injection, interpreter abuse | Multiple contemporary models | Coding-capable models did better; 25%–50% prompt-injection success across all tested models | Real-world attack rates; proficient exploit generation |
| Project Naptime | Google Project Zero, June 2024 | CyberSecEval 2 memory-safety tasks | Tool-supported agent framework with debugger, scripting, verification, parallel trajectories | Up to 20x improvement; Buffer Overflow 0.05 to 1.00; Advanced Memory Corruption 0.24 to 0.76 | General field success rate; daily-work impact (the authors said more progress was needed) |
| Reliability study | IBM Research, May 2024 | Identifying and reasoning about vulnerabilities | 228 code scenarios, eight LLMs, eight investigative dimensions | Conclusion that LLMs cannot yet do this reliably | Any model outside the study; population-wide estimates |
| CVE-Bench 1.0 and VulnLMP | OpenAI, GPT-5.6 system card | Exploiting sandboxed web apps; long-horizon research on real source-available software | 34 of 40 challenges, source withheld, pass@1 over three rollouts; research harness for VulnLMP | Credible leads, reproducible crashes, root causes, some exploit primitives | Independent full-chain exploit or verifier-confirmed Critical outcome (none reported) |
No independent, industry-wide measurement of AI-assisted discovery success is established by these sources. The scores come from different tasks, models and success definitions, so you cannot average them into a single percentage.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Limits and risks
Benchmarks measure different jobs
CTF challenges, sandboxed web apps, research targets whose source code is available, remote probing and multi-day campaigns each test something different. A high score on one is not a proxy for all software, all attack surfaces or operational capability.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Scan for outdated or missing drivers - takes under a minute3Repair Windows errors before they cause bigger problemsFalse leads cost real time
A crash or sanitizer report may indicate a bug, but security relevance depends on reproduction and evidence of impact. Without a verification step, a model that produces many plausible leads can bury a team in triage. This is why Naptime built automatic verification into its framework and why OpenAI assigns proof of impact to the verifier instead of the model.
Best Value
Dual use is built in
The same skills that help defenders find and prioritize flaws can help attackers build offensive capability. Meta frames benefits and misuse risks together. Two practical consequences follow. Use these techniques only on systems you own or are authorized to test. And expect safety filters to sometimes refuse legitimate defensive work, the false-refusal trade-off Meta documents.
Model-plus-tooling claims travel badly
Because results depend on debuggers, build systems, verifiers, repeated attempts and compute, a claim made about a tuned research harness does not transfer to casual use. The reverse also holds: a weak result from a bare prompt understates what a well-built system could do.
A separate question: securing AI systems themselves
“AI and vulnerabilities” also covers the security of AI systems. The UK Department for Science, Innovation and Technology commissioned an assessment, Cyber security risks to artificial intelligence, with a literature cutoff of 10 February 2024. It maps vulnerabilities across the AI lifecycle (design, development, deployment and maintenance) and separates software vulnerabilities from those specific to AI and those shared by both. Prompt injection, covered by Meta’s benchmark above, belongs to this second category. It is about attacking the AI tool, not about the tool finding bugs in other software.
Free tools Windows power users keep installed
One-click scans. No signup required.
This matters when you adopt AI for security work. The assistant, its plugins and its access to your code and build systems become part of your attack surface.
How to evaluate a claim or a tool
Use these five questions for any model, product or announcement:
- Task: Was it source-code review, patch analysis, exploit generation, remote web probing, a CTF or long-horizon target research?
- Target and access: Was the target a benchmark or deployed software? Was source code available or withheld? Was it a sandbox or a live environment?
- System setup: Was it a standalone prompt or an agent with a debugger, scripting, build system, verifier, parallel runs and extra compute?
- Success definition: Which of the five levels at the top of this article was reached: a flag, a reproduced bug, verified impact, an exploit primitive or a full exploit?
- Reliability and safety: Were results consistent across repeated runs? How many leads were false? Does the model refuse benign defensive requests, and what stops harmful use?
If a source cannot answer these, treat its headline as marketing or a preliminary signal.
Where this leaves defenders
The defensible use today is AI as a lead generator and force multiplier inside a verified workflow. Give the model tools, let it propose hypotheses, and let a separate check confirm reproducibility and impact before anything is labelled a vulnerability. The evidence supports that reading: constrained tasks improve sharply with interactive tooling, and long-horizon runs on real software produce credible leads, crashes and sometimes exploit primitives. It does not support claims of reliable autonomous zero-day discovery against arbitrary targets, and the 2024 sources in particular predate the models now in use. Re-check any figure against the exact model, benchmark version and setup before relying on it.
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallOutdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchQuick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




