Hardware FixRecommendedDevice not working? Your driver may be the problemCheck updates for common hardware issues.Fix DriversOctober DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsClean PCRecommendedOne scan can reveal what keeps slowing WindowsLook for cleanup and repair opportunities.Run Scan×
Skip to content
World desk4 min

How Well Can AI Models Reconstruct Reported Cyber Attacks?

Cyber Autopsy scores AI models on timelines, evidence citations, event relationships and uncertainty when reconstructing reported cyber incidents. The early leaderboard is informative, but its single-run results are not a stable ranking.
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

AI models can turn a documented cyber incident into a structured timeline, but the scores from Cyber Autopsy are an early snapshot—not proof that one model is reliably best. In a leaderboard snapshot reported on 2 October 2026, Gemma 4 scored 83.22 EGRS overall, ahead of GPT-5.6 Luna at 81.06 and Grok 4.20 at 80.50. The benchmark’s author says each model was run once, so those numbers should be read as exploratory results rather than a stable ranking.

What Cyber Autopsy asks a model to do

Cyber Autopsy evaluates whether a model can reconstruct a reported incident from an evidence packet. The task is not to carry out an intrusion or simulate live attacker behavior. A model must organize reported details into events, connect events in a timeline or relationship graph, cite the supporting evidence, and distinguish what is confirmed from what is inferred, attempted, failed, or unknown.

That last distinction matters: a coherent-sounding attack narrative can still overstate what the evidence establishes. As benchmark author ujja puts it, “A plausible attack story is not enough; unsupported certainty should count against it.”

How the score is built

The benchmark uses a deterministic matching process. Event matching is one-to-one; text similarity proposes candidate matches, shared evidence IDs add a bonus, and a threshold filters weak matches. The EGRS score combines event recall and precision with relationship quality, evidence attribution, status accuracy, uncertainty calibration, and recognition of failed actions. It also penalizes hallucinated events.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The published formula is: EGRS = 100 × max(0, 0.25 × recall + 0.20 × precision + 0.15 × link F1 + 0.15 × evidence attribution + 0.10 × status accuracy + 0.10 × unknown calibration + 0.05 × failed recognition − 0.25 × hallucination rate).

What the seven initial tasks cover

The initial evaluation contains seven task rows built from four public reports. Some reuse the same incident evidence in different conditions, so the seven rows are not seven independent incidents.

Incident and task IDs What the report describes Evidence and scope caveat
RansomHub intrusion: CASE-001 and CASE-004 The DFIR Report account describes password spraying, RDP access, credential access, Rclone exfiltration, and RansomHub deployment. CASE-004 uses only first-day evidence and has a 15-event reference graph; the full-case reference has 28 events. The benchmark descriptions identify this case as based on host and network telemetry described by The DFIR Report.
GTG-1002 espionage campaign: CASE-002, CASE-011, and CASE-012 Anthropic’s incident and technical reports describe an alleged AI-orchestrated campaign against roughly 30 targets. The campaign details and attribution are vendor-reported, not independently verified victim-side telemetry. CASE-011 and CASE-012 use identical evidence but different human-versus-AI-agent framing.
GTG-2002 extortion operation: CASE-003 Anthropic’s August 2025 misuse report describes a Claude Code-assisted data-extortion operation affecting at least 17 organisations. The task’s reference reconstruction has eight events. The report’s ransom-note images were simulated recreations and were excluded from the benchmark evidence.
AI-enabled credential harvesting: CASE-013 Google GTIG/Mandiant’s September 2026 report describes a campaign that reportedly harvested thousands of credentials in under six hours. The victim and model are undisclosed, and the claims are vendor-reported. The reference reconstruction contains seven events.

These cases differ in how much detail the reports provide, the kind of evidence they contain, and the size of their reference graphs. Scores across them are therefore not a clean comparison of incident difficulty.

What the reported scores show—and what they do not

For a leaderboard snapshot fetched on 2 October 2026, the article reports these overall EGRS scores: Gemma 4, 83.22; GPT-5.6 Luna, 81.06; and Grok 4.20, 80.50. The overall figure is an equal-weight mean across seven task rows, including related variants. The article says duplicate and failing task attachments were removed and earlier evaluated versions restored for that snapshot.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Results also vary by case. Gemma 4 scored 92.11 on CASE-003, the shorter extortion task. On CASE-013, Gemini 3.7 Flash scored 89.33 while Claude Opus 5 scored 52.47—a 36.86-point spread calculated from those two results. Gemini led two case rows, Gemma led three, Grok led one, and GPT-5.6 Luna led one. Those figures describe performance on these particular task versions and evidence packets, not general cybersecurity ability.

For RansomHub, Gemini scored 79.57 on the first-day task and 70.55 on the full-case task, a 9.02-point difference. The reference graphs differ in size, so this does not establish that providing less evidence makes reconstruction easier.

Why the leaderboard is only a snapshot

  • Each model was run once; the article reports no repeated-trial confidence intervals.
  • The seven rows include related task variants, so they are not independent samples.
  • Task versions differ: CASE-001 through CASE-011 use task version 3, while CASE-012 and CASE-013 use republished version 1. A benchmark row pinned to one task version does not automatically inherit scores from another.
  • Task creation status and whether a particular model completed a task are separate matters.

The benchmark author reported that seven more cases, CASE-014 through CASE-020, had been added after the leaderboard snapshot, with their gold graphs still undergoing independent review at the time of writing. They broaden the incident types and source material, but do not create a controlled human-versus-AI experiment.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What the framing comparison can tell us

CASE-011 and CASE-012 keep the evidence identical while changing whether the activity is framed as human-led or AI-agent-led. The reported difference between the human-framed and AI-agent-framed scores ranges from +9.25 points for Grok to −4.61 for Claude Opus 5; five models scored higher under each framing condition.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

This is an exploratory indication that wording may affect reconstruction scores. It cannot establish who actually conducted the reported campaign, and it is not a controlled comparison of human and AI attackers.

How to interpret the results usefully

The most informative comparison is not just the overall leaderboard number. For any particular case, consider the evidence source and depth, task version, reference graph size, and how a model handled event relationships, citations, uncertainty, and failed actions. A high score on a short graph does not automatically mean stronger performance on a larger or differently documented incident.

Cyber Autopsy is best read as a pilot evaluation of evidence-grounded reconstruction: it makes explicit whether a model accounts for reported events and how carefully it ties claims to evidence. Its current results offer a starting point for comparing structured outputs, but the one-run design and varied cases limit how confidently they can be generalized.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More from the Wire

  1. World desk4 min
    How to Spot an AI Voice Scam Before Sending MoneyDon’t rely on how a caller sounds. Pause, call back through a known number, and verify the emergency with another trusted person before sending money.
  2. Mountain View desk4 min
    Google’s SynthID Detector: How to Check AI-Generated Images, Video and AudioGoogle’s SynthID Detector looks for an embedded watermark in supported images, video and audio. Here is what its results do—and do not—show.
  3. Redmond desk20 min
    How to create a link to File or Folder in Windows 11Windows 11 gives you several ways to point to a file or folder without moving or duplicating it. You can create a desktop shortcut,…
Recommended PC Tool
Recommended PC Tool
Windows Errors? Fix Them Before They SpreadFree repair scan
Crashes, No Sound, or Screen Glitches?Free driver scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.