Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Repair Windows errors before they cause bigger problems3Fix the driver behind crashes, sound loss and screen glitchesAI models can turn a documented cyber incident into a structured timeline, but the scores from Cyber Autopsy are an early snapshot—not proof that one model is reliably best. In a leaderboard snapshot reported on 2 October 2026, Gemma 4 scored 83.22 EGRS overall, ahead of GPT-5.6 Luna at 81.06 and Grok 4.20 at 80.50. The benchmark’s author says each model was run once, so those numbers should be read as exploratory results rather than a stable ranking.
What Cyber Autopsy asks a model to do
Cyber Autopsy evaluates whether a model can reconstruct a reported incident from an evidence packet. The task is not to carry out an intrusion or simulate live attacker behavior. A model must organize reported details into events, connect events in a timeline or relationship graph, cite the supporting evidence, and distinguish what is confirmed from what is inferred, attempted, failed, or unknown.
That last distinction matters: a coherent-sounding attack narrative can still overstate what the evidence establishes. As benchmark author ujja puts it, “A plausible attack story is not enough; unsupported certainty should count against it.”
How the score is built
The benchmark uses a deterministic matching process. Event matching is one-to-one; text similarity proposes candidate matches, shared evidence IDs add a bonus, and a threshold filters weak matches. The EGRS score combines event recall and precision with relationship quality, evidence attribution, status accuracy, uncertainty calibration, and recognition of failed actions. It also penalizes hallucinated events.
Recommended Free Tools
#1 Best Overall
The published formula is: EGRS = 100 × max(0, 0.25 × recall + 0.20 × precision + 0.15 × link F1 + 0.15 × evidence attribution + 0.10 × status accuracy + 0.10 × unknown calibration + 0.05 × failed recognition − 0.25 × hallucination rate).
What the seven initial tasks cover
The initial evaluation contains seven task rows built from four public reports. Some reuse the same incident evidence in different conditions, so the seven rows are not seven independent incidents.
| Incident and task IDs | What the report describes | Evidence and scope caveat |
|---|---|---|
| RansomHub intrusion: CASE-001 and CASE-004 | The DFIR Report account describes password spraying, RDP access, credential access, Rclone exfiltration, and RansomHub deployment. | CASE-004 uses only first-day evidence and has a 15-event reference graph; the full-case reference has 28 events. The benchmark descriptions identify this case as based on host and network telemetry described by The DFIR Report. |
| GTG-1002 espionage campaign: CASE-002, CASE-011, and CASE-012 | Anthropic’s incident and technical reports describe an alleged AI-orchestrated campaign against roughly 30 targets. | The campaign details and attribution are vendor-reported, not independently verified victim-side telemetry. CASE-011 and CASE-012 use identical evidence but different human-versus-AI-agent framing. |
| GTG-2002 extortion operation: CASE-003 | Anthropic’s August 2025 misuse report describes a Claude Code-assisted data-extortion operation affecting at least 17 organisations. | The task’s reference reconstruction has eight events. The report’s ransom-note images were simulated recreations and were excluded from the benchmark evidence. |
| AI-enabled credential harvesting: CASE-013 | Google GTIG/Mandiant’s September 2026 report describes a campaign that reportedly harvested thousands of credentials in under six hours. | The victim and model are undisclosed, and the claims are vendor-reported. The reference reconstruction contains seven events. |
These cases differ in how much detail the reports provide, the kind of evidence they contain, and the size of their reference graphs. Scores across them are therefore not a clean comparison of incident difficulty.
What the reported scores show—and what they do not
For a leaderboard snapshot fetched on 2 October 2026, the article reports these overall EGRS scores: Gemma 4, 83.22; GPT-5.6 Luna, 81.06; and Grok 4.20, 80.50. The overall figure is an equal-weight mean across seven task rows, including related variants. The article says duplicate and failing task attachments were removed and earlier evaluated versions restored for that snapshot.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Rank #3
Results also vary by case. Gemma 4 scored 92.11 on CASE-003, the shorter extortion task. On CASE-013, Gemini 3.7 Flash scored 89.33 while Claude Opus 5 scored 52.47—a 36.86-point spread calculated from those two results. Gemini led two case rows, Gemma led three, Grok led one, and GPT-5.6 Luna led one. Those figures describe performance on these particular task versions and evidence packets, not general cybersecurity ability.
For RansomHub, Gemini scored 79.57 on the first-day task and 70.55 on the full-case task, a 9.02-point difference. The reference graphs differ in size, so this does not establish that providing less evidence makes reconstruction easier.
Rank #4
Why the leaderboard is only a snapshot
- Each model was run once; the article reports no repeated-trial confidence intervals.
- The seven rows include related task variants, so they are not independent samples.
- Task versions differ: CASE-001 through CASE-011 use task version 3, while CASE-012 and CASE-013 use republished version 1. A benchmark row pinned to one task version does not automatically inherit scores from another.
- Task creation status and whether a particular model completed a task are separate matters.
The benchmark author reported that seven more cases, CASE-014 through CASE-020, had been added after the leaderboard snapshot, with their gold graphs still undergoing independent review at the time of writing. They broaden the incident types and source material, but do not create a controlled human-versus-AI experiment.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.What the framing comparison can tell us
CASE-011 and CASE-012 keep the evidence identical while changing whether the activity is framed as human-led or AI-agent-led. The reported difference between the human-framed and AI-agent-framed scores ranges from +9.25 points for Grok to −4.61 for Claude Opus 5; five models scored higher under each framing condition.
Best Value
This is an exploratory indication that wording may affect reconstruction scores. It cannot establish who actually conducted the reported campaign, and it is not a controlled comparison of human and AI attackers.
How to interpret the results usefully
The most informative comparison is not just the overall leaderboard number. For any particular case, consider the evidence source and depth, task version, reference graph size, and how a model handled event relationships, citations, uncertainty, and failed actions. A high score on a short graph does not automatically mean stronger performance on a larger or differently documented incident.
Cyber Autopsy is best read as a pilot evaluation of evidence-grounded reconstruction: it makes explicit whether a model accounts for reported events and how carefully it ties claims to evidence. Its current results offer a starting point for comparing structured outputs, but the one-run design and varied cases limit how confidently they can be generalized.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




