An aggregate benchmark score can hide a dangerous weak spot—and an evaluator can make the same mistake by turning ambiguous evidence into a confident claim. In Day 3 of his Kaggle benchmarking challenge, Sean Campbell reports that his model results raised that warning, then finds it in his own AI-assisted writing workflow.
What the Day 3 benchmark measures
Campbell says his benchmark contains 200 invented items divided among four task shapes: route, classify, judge, and ground. One in five items is answerable only with ESCALATE. Rather than relying on one combined score, he tracks task performance and false-confidence rate separately, then examines each model’s weakest shape using Wilson intervals.
As an Amazon Associate I earn from qualifying purchases.
The results below are Campbell’s reported measurements, not an independently replicated evaluation. The post reports results for 12 hosted models; its weakest-shape summary is:
| Weakest reported task shape | Models |
|---|---|
| Ground | Gemini 3.7 Flash, Gemini 3.1 Pro, Claude Sonnet 5, Claude Opus 5, Gemini 3.8 Flash, GPT-5.5, GPT-5.4 nano |
| Classify | Qwen3 235B Instruct, Claude Haiku 4.5, Gemma 4 26B, gpt-oss-20b, DeepSeek-R1 |
Claude Haiku 4.5 was measured on only three shapes because all of its route calls failed. A weakest-shape label is a useful pointer, not a complete model ranking: it says where a model struggled most in this benchmark, not how it will perform across other tasks or settings.
#1 Best Overall
Why false confidence matters more than an average
The starkest example in Campbell’s account is Haiku 4.5 on judge items. It answered 9 of the 10 items that should have been escalated, which the author describes as a 90% false-confidence rate for that shape. Across its three measured shapes, it answered anyway on 10 of 28 unanswerable items. The problem was concentrated in judge rather than evenly distributed across tasks.
That distinction is why a single aggregate score can mislead. A model may do adequately on common or answerable items while failing precisely where it should recognize that the evidence is insufficient. Looking at the worst task shape and the response to unanswerable cases separately makes that failure visible.
A zero observed rate is not proof of zero risk
For the top six rows in the author’s table, each shape had only 8 to 12 unanswerable items. Campbell estimates that even zero observed false-confidence errors in samples that small leaves an upper bound of roughly 24% to 32%. The intervals overlap, so the post does not establish a confident ranking among the apparent leaders.
Recommended Free Tools
In practical terms, “no failures observed” means no failures appeared in that limited sample. It does not establish that the underlying failure rate is zero. The number of unanswerable examples and the uncertainty around the rate belong beside the rate itself.
What two runs say about repeatability
Campbell reports running four frontier models twice over all 200 items and comparing whether each gave the same answer. These are agreement counts, not a ranking of quality:
| Model | Same answer across two runs | Reported agreement |
|---|---|---|
| Claude Opus 5 | 199 of 200 items | 99.5% |
| Claude Sonnet 5 | 195 of 200 items | 97.5% |
| Gemini 3.1 Pro | 195 of 200 items | 97.5% |
| GPT-5.5 | 194 of 200 items | 97.0% |
The intervals overlap, so these small differences do not justify declaring one model the most consistent. Nor does agreement alone show that an answer is correct: a model can repeat the same wrong answer.
Rank #4
When a scoring difference may be a parsing difference
Campbell says Gemini’s five verdict flips came from replies that hit an output-length cap and parsed successfully in only one run, rather than from substantively different answers. The scorer counted an error as its own verdict. That is a consequential design choice: if malformed or incomplete output is scored as a distinct answer, the agreement measure combines model behavior with the parser’s handling of output failures.
Generation settings were not identical
Only Gemini ran at temperature 0. The two Claude 5 models rejected that setting, while GPT-5.5 used its default. Campbell also says a third run for classify, judge, and ground had hit Kaggle’s daily spend cap and was expected the following day, so the figures in the post were not yet final. These qualifications limit how directly the reported repeat-run figures can be compared.
Best Value
Operational problems that can look like model behavior
The post also describes execution issues that affect how benchmark results should be interpreted. These are Campbell’s observations and recommendations, not independently verified statements about Kaggle’s general behavior.
- Run timer versus call time: Campbell says repeat runs appeared to take 2–5 seconds for 40–60 items, even though downloads contained all expected items. He cautions against reading Kaggle’s run timer as a direct measure of time spent on model calls.
- Retries and duplicate spend: A retrying sandbox task ran for five minutes, was killed at 300 seconds, and resubmitted paid runs, causing duplicate spend.
His proposed safeguard is to separate submission from collection: submit in one short task, collect in another, and make paid actions refuse duplicate runs. This addresses two different failure modes—long-running retries and accidental resubmission—without treating a platform timer as proof that model calls were missing.
The personal correction: do not turn an ambiguous note into a grade
Campbell’s title refers to an earlier AI-assisted writing session. He says a terse note was misread as a grade; he had not graded anything, but the session recorded a grade in his voice and he published it without noticing.
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Fix the driver behind crashes, sound loss and screen glitches3Clear out junk files and repair common Windows errorsHis change is simple: preserve wording that might be a grade as words, and ask for clarification rather than silently converting it into a factual claim. The lesson matches the benchmark’s central warning. When the evidence is unclear, a confident answer can be an evaluation failure—even when the system producing it is the writer’s own workflow.
A practical way to read a model benchmark
For a comparison like this one, inspect more than the headline score. The Day 3 results point to a focused checklist:
Quick Recap
- Find each model’s weakest task shape, not just its average.
- Check false-confidence rates on cases that should be escalated, alongside the number of such cases.
- Read the uncertainty interval; a zero observed rate in a small sample is not a zero-risk guarantee.
- Separate repeatability from correctness, and avoid ranking models when intervals overlap.
- Check whether output caps, parser behavior, generation settings, timers, or retries could explain an apparent failure or difference.
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




