The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →You can build an AI-powered code vulnerability scanner, but the model shouldn’t be the scanner. The workable design uses an established static-analysis engine, such as CodeQL or Semgrep, to find candidate issues. A language model then handles one narrow, clearly defined job, such as judging candidate findings in context or checking code against your organization’s own security instructions. The results go back to developers as reviewable alerts. An LLM on its own can’t guarantee that it finds every vulnerability, and nothing in the available evidence supports claiming otherwise.
This guide walks through that workflow in order: scope, analysis engine, AI role, reporting, evaluation, and securing the scanner itself.
Can AI find vulnerabilities in source code?
It can help, within limits. Static application security testing (SAST) analyzes source code for vulnerabilities, and tools like CodeQL and Semgrep are established ways to do it. Adding AI to that pipeline is a design choice. It isn’t a replacement for it. No comparable, published performance figures exist for an “LLM plus static analysis” scanner of the kind described here, so any detection rate or false-positive rate you quote has to come from your own documented evaluation (covered below).
Two practical consequences follow:
- Give the model a bounded task whose output you can check, not an open-ended “find all the bugs” instruction.
- Treat the AI layer’s output as a reviewable opinion attached to a finding, not as ground truth.
The workflow at a glance
- Define scope: languages, frameworks, vulnerability classes, and what gets scanned (full repositories, pull requests, or selected code).
- Run a static-analysis engine to produce candidate findings.
- Apply an AI layer for one specific contextual task.
- Report inside the developer workflow, ideally as SARIF so existing tooling can display it.
- Evaluate against a documented corpus, and re-evaluate whenever the rules, prompts or model change.
Step 1: Define scope before choosing tools
Decide what the scanner promises and what it doesn’t. Write down:
#1 Best Overall
- Target languages and frameworks. Support differs per engine, and framework-specific behavior (routing, templating, ORM calls) often decides whether a finding is real.
- Vulnerability classes. Injection, authentication flaws, unsafe deserialization and so on. A scanner that claims “all vulnerabilities” is making a claim nobody can verify.
- Scan unit. Full repository on a schedule, pull-request diffs, or selected files. This affects cost, latency and how much context the AI layer sees.
- Build requirements. CodeQL documents its supported languages and systems, and analysis of compiled languages may require a successful build. Confirm the requirements against your real repositories, not a sample project.
Step 2: Choose the analysis engine
The engine produces the findings the rest of the system depends on, so compare options on concrete axes rather than reputation.
| Axis | CodeQL | Semgrep | AI-assisted layer |
|---|---|---|---|
| What it is | GitHub’s code analysis engine for automating security checks; treats code as data | Static analysis engine for bugs, vulnerabilities and code standards (OWASP’s description) | An LLM applied to a specific task on top of an engine’s output or a repository |
| Customization | Supports custom queries | Rule-based; customization depth not detailed in the sources reviewed | Natural-language instructions, such as organization-specific security guidelines |
| Build needs | Compiled languages may require a successful build | Not stated in the sources reviewed | None inherent, but depends on how code is gathered |
| Output into GitHub | Native code scanning alerts | Can be brought in as third-party results via SARIF | Must be converted to SARIF or another reportable format |
| Validated performance | Evaluate on your own corpus | Evaluate on your own corpus | No comparable published figures for this architecture |
GitHub Docs describes the relationship plainly: “CodeQL is the code analysis engine developed by GitHub to automate security checks.” Because GitHub code scanning accepts third-party results in SARIF, you aren’t forced to pick a single engine. You can run more than one and merge results into the same alert view.
Step 3: Give the AI one explicit job
The biggest design mistake is letting the model’s role stay vague. Pick a task you can describe, constrain and measure. Two reasonable options:
Contextual review of candidate findings
The static engine flags a code path. The model receives the finding, the relevant code and its immediate surroundings, then returns a structured judgment: does the flagged data flow look reachable and unsanitized, and why? Its answer is attached to the finding. It doesn’t delete it.
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallOutdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchRepository checks against custom instructions
OWASP’s AGHAST project is an example of this approach: an LLM examines a repository against organization-specific security instructions. Its hybrid and static modes require Semgrep Community Edition. Treat it as a demonstration of a pattern, not as evidence of a particular detection rate.
Design rules for the AI layer
- Require structured output (verdict, confidence, rationale, cited lines) so results can be validated and diffed.
- Never silently suppress. If the model marks a finding as likely benign, keep the original alert visible with the reasoning, so a human can overrule it.
- Treat scanned code as untrusted input. Comments or strings in the repository can contain text aimed at the model. Keep instructions and code clearly separated, and don’t give the model tools it doesn’t need.
- Record the model identifier, prompt version and rule versions with every result, so a change in output can be traced.
Here is the shape of a result record the AI layer might return for each candidate:
Rank #3
{
"finding_id": "engine-rule-123:src/api/user.py:48",
"verdict": "likely_true_positive",
"confidence": "medium",
"rationale": "User-supplied 'id' reaches the query string without parameterization.",
"evidence_lines": [44, 46, 48],
"model": "<model id>",
"prompt_version": "triage-v3"
}
Step 4: Report where developers already work
A scanner nobody sees is a scanner nobody uses. GitHub code scanning presents potential vulnerabilities as repository alerts, can run on schedules or on repository events such as pushes and pull requests, and accepts results from third-party tools in SARIF (Static Analysis Results Interchange Format). Emitting SARIF therefore gives you alerts, history and review flow without building your own interface.
Both common engines can produce SARIF from the command line. For example, in a Semgrep-based setup:
Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Repair Windows errors before they cause bigger problemsFix Now →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →semgrep scan --config auto --sarif --output results.sarif
CodeQL’s CLI follows a create-database-then-analyze pattern, and its analysis step can write SARIF as well. Check the current CLI documentation for the exact flags in your version. Your AI layer should then read this file, add its verdicts, and write an enriched SARIF file for upload with GitHub’s SARIF upload action or API. If you target a different platform, confirm it accepts SARIF before relying on the same design.
Rank #4
For pull requests, scan the changed code with enough surrounding context to be meaningful, and keep comments to findings a developer can act on. A noisy bot trains people to ignore it.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Step 5: Evaluate before you make any claim
Build a documented test corpus of vulnerable and non-vulnerable examples in the languages and frameworks you actually support. Include the non-vulnerable cases deliberately. They’re what reveal false positives, and they’re where an overconfident model tends to fail.
Track these dimensions:
- Missed issues (false negatives) per vulnerability class.
- False positives, both from the engine alone and after the AI layer.
- Severity usefulness: do the assigned severities help reviewers prioritize?
- Reproducibility: run the same input several times and compare. LLM output can vary.
- Drift: re-run the whole corpus after any model, prompt or rule change.
Compare three configurations on the same corpus: engine only, AI only, and engine plus AI. That’s the only way to show whether the AI layer adds anything. If it doesn’t improve results for a given task, remove it. Publish a performance number only if it comes from this documented evaluation, with the corpus and conditions stated.
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Repair Windows errors before they cause bigger problems3Scan for outdated or missing drivers - takes under a minuteBest Value
Securing the scanner itself
If your scanner includes an LLM, it’s an LLM application and inherits LLM-specific risks. OWASP warns that failures in LLM applications include issues that conventional SAST, DAST and SCA weren’t designed to find. If the scanner will also evaluate LLM-based applications, plan testing beyond those conventional categories. OWASP points readers to dedicated LLM application security and red-team guidance for that work.
Practical safeguards include limiting what the model can access, keeping credentials out of its context, running it in an isolated environment, and reviewing what leaves your network if you use a hosted model. Source code is sensitive, so decide up front whether it can be sent to an external service.
Quick Recap
Common failure modes
- Overclaiming coverage. Saying the tool finds “all vulnerabilities” when it supports a handful of languages and classes.
- Build failures in compiled projects producing empty or partial results that look like a clean scan. Report scan completeness alongside findings.
- Model-only scanning with no deterministic engine underneath, which makes results hard to reproduce or audit.
- Hidden suppression, where AI triage removes alerts developers never see.
- No re-evaluation after a model upgrade, so quality changes unnoticed.
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




