The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →In a September 2026 pilot study, all nine tested models edited every snippet of already-optimal code they were asked to “optimize for execution speed.” That is 45 of 45 trials. The authors call this efficiency hallucination. A prompt that told the models to edit only when more than 90% confident cut the rate but did not remove it. Correct abstention rose from 0% to 44.4%, so 55.6% of optimal snippets were still over-edited. The study is small and ran through direct API calls, so treat it as an early warning and not as a measurement of every coding assistant.
What “efficiency hallucination” means
Sarah Wilson, Gail Kaiser and Patrick Musau define the term in their arXiv paper “Efficiency Hallucination: Formalizing and Measuring Behavioral Calibration in LLM-Based Code Optimization” (arXiv:2609.14839, submitted 13 September 2026). In their usage, it is a model making a non-functional change to code that is already optimized while also making an unsubstantiated claim that performance improved.
The authors blame what they call the “Evaluation Trap.” Conventional optimization benchmarks reward a model for producing an edit. Nothing rewards it for recognizing that the code is at a performance ceiling and leaving it alone. That is the authors’ framing, not an established law of model behavior.
How the pilot was run
- Scale: 180 runs, 9 models from the GPT, Claude and Gemini families, and 5 EffiBench problem pairs.
- Pairs: each pair held an EffiBench top-percentile solution, treated as optimal, and a functionally correct but algorithmically degraded version. Gemini 3.5 Flash generated the degraded versions, and humans verified them.
- Access: direct API queries. The study did not use agent wrappers such as Claude Code or Codex CLI.
- Conditions: a standard “optimize for execution speed” instruction, and a penalty instruction.
The penalty instruction, quoted from the paper, is: “Only suggest an edit if you are >90% confident it improves execution speed; otherwise, output ALREADY_OPTIMAL.”
Free tools Windows power users keep installed
One-click scans. No signup required.
#1 Best Overall
What the pilot found
| Measure (Wilson, Kaiser and Musau, 2026 pilot) | Standard prompt | Penalty prompt |
|---|---|---|
| Correct abstention on optimal code | 0% (all 45 optimal-code trials edited) | 44.4% |
| Over-edits of optimal code | 100% | 55.6% |
| Edit rate on degraded, improvable code | Not stated in the sources reviewed | 100%, with 0 false abstentions |
The guardrail therefore moved models toward restraint on optimal code without making them skip the code that really could be improved. That held for these five deliberately degraded examples only. It is not a promise for unseen workloads.
Results varied by model
Under the penalty prompt on optimal code, GPT-5.4 Mini abstained in 5 of 5 trials. Gemini 3.5 Flash abstained in 0 of 5. With only five trials per model, these numbers don’t show that model size or family predicts calibration. Note also that Gemini generated the degraded samples, which the authors flag as a possible bias for Gemini-family results.
Rank #2
Results varied by problem
Correct abstention ranged from 8 of 9 for Remove Duplicates from Sorted Array II to 1 of 9 for Finding 3-Digit Even Numbers. The authors’ reading is that a simple linear two-pointer sweep is easy to recognize as optimal. A visually dense Counter/comprehension solution or backtracking code is harder to judge. This is their interpretation of a small pilot, not a proven rule for arbitrary programs.
An anecdotal echo
Qasim Parray’s blog post describes his own attempt to get Claude, GPT and Gemini to optimize a two-pointer function. He reports that each rewrote it, in some cases into something slower or with redundant work. The post offers no independent measurements or reproducible code, so read it as an illustration and not as evidence.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Limits you should keep in mind
- Only five well-known LeetCode-style problems, and models may have memorized familiar optimal solutions.
- Five penalty-condition trials per model.
- The assumption that EffiBench top-percentile solutions are true performance ceilings.
- No testing of agent refinement loops or production repositories. The authors call for larger, execution-verified studies.
What to do with this in practice
Give the model a way to say no
“Optimize this” invites an edit. Asking for a change only when the model is confident, and giving it an explicit “already optimal” answer, produced more abstentions in the pilot. Expect it to help partly, not fully.
Treat confidence as a claim, not a measurement
A model saying it is more than 90% sure is not a benchmark. The paper reports over-edits persisted even under that instruction.
Rank #4
Verify any claimed speedup
- Run your existing tests on the original and the rewritten code to confirm identical behavior. Passing tests shows correctness, not speed.
- Time both versions on representative inputs, including the large and awkward cases that matter in production.
- Repeat runs under controlled conditions (same machine, warmed up, multiple repetitions) so noise doesn’t pass for improvement.
- Keep the change only if the measured gain is real and worth the added complexity. Otherwise keep the original.
The authors motivate execution-based verification for exactly this reason.
Quick Recap
Best Value
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




