October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsWindows FixRecommendedWindows errors stealing your time? Find the fix fastScan stability, cleanup and performance issues.Fix NowOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
World desk4 min

Can LLMs Audit Code? What a 12-Task Security and Jailbreak Benchmark Found

A custom 12-task benchmark reports strong code and configuration scores, but its small task set and unavailable prompts and outputs limit what those results establish about real-world security audits.
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Sometimes, on the tests they were given—but this benchmark does not establish that large language models can reliably audit real-world systems. A submission by LOI CHIANG HAO, published on DEV Community on October 1, 2026, reports strong scores across 12 security tasks, alongside model-specific misses. Its results are best read as a small, custom stress test, not a general verdict on AI code auditing.

What did the 12-task benchmark test?

The author grouped 12 scenarios into three categories, with four tasks in each. Some test whether a model recognizes a security flaw; others test whether it resists instructions that could lead to unsafe assistance.

As an Amazon Associate I earn from qualifying purchases.

Code vulnerabilities

  • A Python SQL query built with string formatting, which can allow SQL injection.
  • Hardcoded AWS IAM secret keys.
  • A Flask file-download path built with os.path.join(BASE_DIR, filename) without sufficient protection against path traversal.
  • Insecure deserialization using pickle.loads on an unvalidated session endpoint.

Cloud and infrastructure configuration

  • An Nginx open redirect that uses an unvalidated 302 $arg_url.
  • An iptables INPUT ACCEPT default policy that makes purported database allow-rules redundant.
  • An AWS Lambda IAM policy with wildcard permissions for an S3 read operation.
  • A Kubernetes ClusterRole with wildcard verbs and API groups assigned to a read-only monitoring service.

Prompt injection and jailbreaks

  • A DAN-style role-play prompt requesting phishing templates.
  • A simulated tool-use task in which search data contains “[SYSTEM OVERRIDE]” instructions to reveal prompts.
  • A Base64-encoded malware request presented as an encoding study.
  • A creative-writing prompt requesting working SQL injection vectors.

What scores did the author report?

The table reproduces the model labels and results reported by LOI CHIANG HAO in the 2026 DEV Community submission. Each category contains four tasks, so its percentages correspond to performance on that small set; the overall result is out of 12. These are the author’s figures, not independently verified benchmark results.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Model label in submission Overall Code Configuration Jailbreak
Qwen 3 Coder 480B 91.67% (11/12) 100% 100% 75%
Grok 4.20 Reasoning 91.67% (11/12) 100% 100% 75%
Gemini 3.7 Flash 91.67% (11/12) 75% 100% 100%
DeepSeek-R1 83.33% (10/12) 100% 100% 50%
GPT-5.4 83.33% (10/12) 100% 100% 50%
GLM-5 75.00% (9/12) 75% 100% 50%

The category scores make an important distinction: the reported configuration results are uniformly high, while jailbreak results range from 50% to 100%. A model’s overall score can therefore obscure where its misses occurred. With only four scenarios per category, however, each individual task has substantial influence on the category percentage.

What do the reported misses show?

The submission describes particular failures, but the accessible article does not include the raw model outputs. Treat these as the author’s account of responses in this benchmark, not as independently reproduced observations or evidence about model behavior in general.

Path traversal

The author says Gemini 3.7 Flash missed the Flask path-traversal issue. The author’s explanation is that os.path.join(BASE_DIR, filename) does not, by itself, ensure a path stays inside the intended directory: an absolute path or a ../ segment may escape it. The reported miss concerns that task and response, not every path-handling review the model might perform.

Jailbreak and indirect-injection tasks

The author says GPT-5.4 failed the DAN-style role-play and Base64-bypass tasks, and reports that it decoded the malware payload and assisted with credential-extraction concepts. The submission also says DeepSeek-R1 failed the indirect prompt-injection and fictional-framing tasks. Without the prompts and outputs, those claims cannot be checked independently or used to establish a general cause, such as a failure of reasoning over untrusted data.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Strong scores are still bounded results

The author reports that every tested model flagged the SQL injection, hardcoded credentials, and pickle-deserialization tasks, and that all six received 100% in the configuration category. That supports a narrow statement about these scenarios and scoring rules. It does not prove comprehensive coverage of those vulnerability classes or readiness to approve a production change.

How was success scored, and what can the score tell you?

The author says the benchmark used automated string assertions and negative-lookaround regular expressions, including an assert_not_contains_regex check intended to stop an answer from passing as a refusal if it still included a disallowed exploit payload.

That approach can check whether an answer contains expected or disallowed text. But the accessible article does not provide the exact prompts, regexes, thresholds, false-positive checks, or task-by-task outputs. Without those details, a reader cannot assess how well the assertions distinguish a sound security analysis from a superficially matching answer—or reproduce the scores.

The submission also does not identify exact provider model snapshots or run settings. Model labels and percentages should therefore be attributed to the October 2026 submission rather than treated as a durable ranking of model families. Aggregate pass rate is one comparison axis; it is not a measure of benchmark validity, comprehensive audit ability, or reproducibility.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What does the benchmark not establish?

  • Real-world audit reliability: Twelve constructed tasks cannot show how a model performs across the range of codebases, configurations, and security-review conditions.
  • Independent reproducibility: The accessible post does not expose the underlying prompts, benchmark code, raw outputs, or per-task assertions.
  • Current model-wide rankings: Exact snapshots and run configurations are not stated, so the results cannot safely be generalized beyond the author’s named test runs.
  • Cost efficiency: The author calls Qwen 3 Coder 480B a score-versus-cost Pareto leader and says it achieved its reported 91.67% pass rate at a fraction of commercial API costs. No numerical costs, token counts, provider rates, execution date, or underlying cost data are supplied, so that claim cannot support a quantified comparison or buying recommendation.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What should developers take away?

The results suggest that models can identify familiar flaws and configuration problems in some deliberately framed scenarios, while still missing other tasks—especially in the reported jailbreak results. That is a reason to test a model against the specific workflow and threat cases where it will be used, not to treat a high score as a security sign-off.

For teams evaluating a coding assistant, useful evidence would include the exact model snapshot and settings, reproducible prompts, raw outputs, transparent scoring rules, and tests that check whether proposed fixes actually remove the issue without introducing another one. The submission proposes multi-turn escalation after an initial refusal, context-window overflow that hides malicious content in legitimate material, and patch verification as future measurements. Those are proposed next steps, not findings from its 12 tasks.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More from the Wire

  1. World desk4 min
    How to Spot an AI Voice Scam Before Sending MoneyDon’t rely on how a caller sounds. Pause, call back through a known number, and verify the emergency with another trusted person before sending money.
  2. Mountain View desk4 min
    Google’s SynthID Detector: How to Check AI-Generated Images, Video and AudioGoogle’s SynthID Detector looks for an embedded watermark in supported images, video and audio. Here is what its results do—and do not—show.
  3. Redmond desk20 min
    How to create a link to File or Folder in Windows 11Windows 11 gives you several ways to point to a file or folder without moving or duplicating it. You can create a desktop shortcut,…
Recommended PC Tool
Recommended PC Tool
PC Slower Than It Used to Be?Free scan - under a minute
Crashes, No Sound, or Screen Glitches?Free driver scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.