October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsClean PCRecommendedOne scan can reveal what keeps slowing WindowsLook for cleanup and repair opportunities.Run ScanOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
World desk4 min

Day 3: The Benchmark Caught Me Too

A benchmark's average can conceal a risky weak spot. Sean Campbell's Day 3 report examines escalation errors, uncertain rankings, repeat runs, and how ambiguous AI-assisted notes became a published grade.
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

An aggregate benchmark score can hide a dangerous weak spot—and an evaluator can make the same mistake by turning ambiguous evidence into a confident claim. In Day 3 of his Kaggle benchmarking challenge, Sean Campbell reports that his model results raised that warning, then finds it in his own AI-assisted writing workflow.

What the Day 3 benchmark measures

Campbell says his benchmark contains 200 invented items divided among four task shapes: route, classify, judge, and ground. One in five items is answerable only with ESCALATE. Rather than relying on one combined score, he tracks task performance and false-confidence rate separately, then examines each model’s weakest shape using Wilson intervals.

As an Amazon Associate I earn from qualifying purchases.

The results below are Campbell’s reported measurements, not an independently replicated evaluation. The post reports results for 12 hosted models; its weakest-shape summary is:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Weakest reported task shape Models
Ground Gemini 3.7 Flash, Gemini 3.1 Pro, Claude Sonnet 5, Claude Opus 5, Gemini 3.8 Flash, GPT-5.5, GPT-5.4 nano
Classify Qwen3 235B Instruct, Claude Haiku 4.5, Gemma 4 26B, gpt-oss-20b, DeepSeek-R1

Claude Haiku 4.5 was measured on only three shapes because all of its route calls failed. A weakest-shape label is a useful pointer, not a complete model ranking: it says where a model struggled most in this benchmark, not how it will perform across other tasks or settings.

Why false confidence matters more than an average

The starkest example in Campbell’s account is Haiku 4.5 on judge items. It answered 9 of the 10 items that should have been escalated, which the author describes as a 90% false-confidence rate for that shape. Across its three measured shapes, it answered anyway on 10 of 28 unanswerable items. The problem was concentrated in judge rather than evenly distributed across tasks.

That distinction is why a single aggregate score can mislead. A model may do adequately on common or answerable items while failing precisely where it should recognize that the evidence is insufficient. Looking at the worst task shape and the response to unanswerable cases separately makes that failure visible.

A zero observed rate is not proof of zero risk

For the top six rows in the author’s table, each shape had only 8 to 12 unanswerable items. Campbell estimates that even zero observed false-confidence errors in samples that small leaves an upper bound of roughly 24% to 32%. The intervals overlap, so the post does not establish a confident ranking among the apparent leaders.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

In practical terms, “no failures observed” means no failures appeared in that limited sample. It does not establish that the underlying failure rate is zero. The number of unanswerable examples and the uncertainty around the rate belong beside the rate itself.

What two runs say about repeatability

Campbell reports running four frontier models twice over all 200 items and comparing whether each gave the same answer. These are agreement counts, not a ranking of quality:

Model Same answer across two runs Reported agreement
Claude Opus 5 199 of 200 items 99.5%
Claude Sonnet 5 195 of 200 items 97.5%
Gemini 3.1 Pro 195 of 200 items 97.5%
GPT-5.5 194 of 200 items 97.0%

The intervals overlap, so these small differences do not justify declaring one model the most consistent. Nor does agreement alone show that an answer is correct: a model can repeat the same wrong answer.

When a scoring difference may be a parsing difference

Campbell says Gemini’s five verdict flips came from replies that hit an output-length cap and parsed successfully in only one run, rather than from substantively different answers. The scorer counted an error as its own verdict. That is a consequential design choice: if malformed or incomplete output is scored as a distinct answer, the agreement measure combines model behavior with the parser’s handling of output failures.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Generation settings were not identical

Only Gemini ran at temperature 0. The two Claude 5 models rejected that setting, while GPT-5.5 used its default. Campbell also says a third run for classify, judge, and ground had hit Kaggle’s daily spend cap and was expected the following day, so the figures in the post were not yet final. These qualifications limit how directly the reported repeat-run figures can be compared.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Operational problems that can look like model behavior

The post also describes execution issues that affect how benchmark results should be interpreted. These are Campbell’s observations and recommendations, not independently verified statements about Kaggle’s general behavior.

  • Run timer versus call time: Campbell says repeat runs appeared to take 2–5 seconds for 40–60 items, even though downloads contained all expected items. He cautions against reading Kaggle’s run timer as a direct measure of time spent on model calls.
  • Retries and duplicate spend: A retrying sandbox task ran for five minutes, was killed at 300 seconds, and resubmitted paid runs, causing duplicate spend.

His proposed safeguard is to separate submission from collection: submit in one short task, collect in another, and make paid actions refuse duplicate runs. This addresses two different failure modes—long-running retries and accidental resubmission—without treating a platform timer as proof that model calls were missing.

The personal correction: do not turn an ambiguous note into a grade

Campbell’s title refers to an earlier AI-assisted writing session. He says a terse note was misread as a grade; he had not graded anything, but the session recorded a grade in his voice and he published it without noticing.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

His change is simple: preserve wording that might be a grade as words, and ask for clarification rather than silently converting it into a factual claim. The lesson matches the benchmark’s central warning. When the evidence is unclear, a confident answer can be an evaluation failure—even when the system producing it is the writer’s own workflow.

A practical way to read a model benchmark

For a comparison like this one, inspect more than the headline score. The Day 3 results point to a focused checklist:

  • Find each model’s weakest task shape, not just its average.
  • Check false-confidence rates on cases that should be escalated, alongside the number of such cases.
  • Read the uncertainty interval; a zero observed rate in a small sample is not a zero-risk guarantee.
  • Separate repeatability from correctness, and avoid ranking models when intervals overlap.
  • Check whether output caps, parser behavior, generation settings, timers, or retries could explain an apparent failure or difference.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More from the Wire

  1. World desk4 min
    How to Spot an AI Voice Scam Before Sending MoneyDon’t rely on how a caller sounds. Pause, call back through a known number, and verify the emergency with another trusted person before sending money.
  2. Mountain View desk4 min
    Google’s SynthID Detector: How to Check AI-Generated Images, Video and AudioGoogle’s SynthID Detector looks for an embedded watermark in supported images, video and audio. Here is what its results do—and do not—show.
  3. Redmond desk20 min
    How to create a link to File or Folder in Windows 11Windows 11 gives you several ways to point to a file or folder without moving or duplicating it. You can create a desktop shortcut,…
Recommended PC Tool
Recommended PC Tool
Crashes, No Sound, or Screen Glitches?Free driver scan
Windows Errors? Fix Them Before They SpreadFree repair scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.