Free tools Windows power users keep installed
One-click scans. No signup required.
A model can return perfectly parseable JSON and still apply the wrong update. A Kaggle benchmark reported on October 1, 2026, tests both sides of that problem: whether a response obeys a strict JSON interface and whether its values match the exact expected state. Its results are useful as a diagnostic, not as a general ranking of models.
What “valid JSON” leaves unanswered
Parsing checks syntax: can software read the response as JSON? Schema checks shape: does it have the required fields and types? Neither check, by itself, establishes that the response changed the state correctly.
As an Amazon Associate I earn from qualifying purchases.
For a patch task, the model must interpret instructions, apply them to an initial state, and return the resulting values in the required format. A response can therefore fail in at least two different ways:
Recommended Free Tools
- Wrong state, valid structure: the output parses and has the right keys and types, but one or more values are wrong.
- Right state, unusable presentation: the values are correct, but the response includes extra material—such as Markdown fences—so a strict consumer cannot parse the complete response as the required JSON document.
The benchmark author summarizes the first distinction this way: “A JSON response can parse successfully and still change the wrong state.”
#1 Best Overall
What the Kaggle benchmark tests
Bilingual Patch Contracts consists of 12 handcrafted state-update scenarios. Each scenario appears with an English, Chinese, and code-switched instruction body, for 36 prompts total. Within each three-prompt group, the initial state and expected answer are shared. The contract prefix and canonical output keys remain in English, so this is not a fully Chinese interaction benchmark.
The scenarios probe different ways instructions can go wrong, including:
- applying a later correction rather than an earlier value;
- respecting negation and distinguishing
nullfrom an empty value; - preserving tag order and case sensitivity;
- converting hours to minutes and applying sequential conditions;
- treating instruction-like text as literal data rather than as a new command; and
- copying Unicode, backslashes, quotation marks, and a newline exactly.
What counts as a pass
A response passes only if the complete response is one JSON object with exactly five keys, the required value types, and every expected value. The scorer does not remove Markdown, repair the output, or ask a second model to judge it.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Some representational differences are accepted: whitespace, key order, and equivalent Unicode escapes do not change the result. Others fail: duplicate keys, extra fields, nonfinite values, booleans or floats where an integer is required, and arrays whose elements are in the wrong order. This strictness makes the score about exact contract compliance, not merely whether a human can infer the intended answer.
How the run was made
For the October 1, 2026 run, the author used ordinary text generation, requested temperature 0 and seed 0 through the SDK, and started a fresh isolated conversation for each case. The run did not use constrained JSON decoding, schema enforcement, or tools. The author reports checking all 36 case IDs against frozen prompts and answers and independently recalculating saved scores. Provider behavior can vary across runs, so these results describe that run rather than a guarantee of repeatability.
Reported results from the October 1, 2026 run
The following are the benchmark author’s reported results for version 2. Exact match means meeting the full strict pass rule; valid JSON and valid schema are separate measurements.
Rank #3
| Model | Strict exact match | Valid JSON | Valid schema |
|---|---|---|---|
| Gemini 3.7 Flash | 36/36 (100%) | 36/36 | 36/36 |
| GPT-5.4 nano | 24/36 (66.7%) | 36/36 | 36/36 |
| Claude Haiku 4.5 | 0/36 (0%) | 0/36 | 0/36 |
| Qwen3-Next-80B-A3B-Instruct | No complete score | No complete score | No complete score |
Qwen3-Next-80B-A3B-Instruct was attempted in both the pilot and version 2, but the attempts stopped with HTTP 429 and a provider heavy-load message. It was excluded rather than scored as zero. The report does not provide a complete per-language score table for all listed models; the figures above are aggregate results across the 36 prompts.
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallCrashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteWhy the two failure types matter
GPT-5.4 nano: valid output, incorrect updates
GPT-5.4 nano produced valid JSON and valid field types on all 36 cases, yet only 24 outputs matched the expected state exactly. In the case-sensitive tags scenario, for example, it retained lowercase beta even though the instruction required removing it. A parser and a type checker would accept that response; an application relying on the updated state would receive the wrong value.
Claude Haiku 4.5: Markdown breaks a strict interface
Claude Haiku 4.5 wrapped every answer in a Markdown code fence despite an explicit instruction not to. Under the benchmark rule, the full response is not a JSON document, even if the JSON text inside the fence contains correct values.
Rank #4
The author reports a separate counterfactual: stripping only complete outer fences would make 33 of 36 responses pass value checks. That diagnostic is not the benchmark score, and it does not change the reported results. In a real system, silently stripping output is a design choice with its own consequences; it should not be confused with the model having followed a raw-JSON contract.
What the language comparisons do—and do not—show
For GPT-5.4 nano, the mixed-language total was two cases higher than the English total. Looking at matched scenarios gives a narrower picture: seven passed in both English and mixed-language forms, three failed in both, and two passed only in the mixed form. English-versus-Chinese comparisons were also mixed.
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Fix the driver behind crashes, sound loss and screen glitches3Clear out junk files and repair common Windows errorsThese paired results point to particular cases worth inspecting; they do not establish that the model is generally stronger in Chinese or code-switching. The instruction bodies were hand-authored, and their phrasing and token lengths were not perfectly controlled. Also, while the bodies vary by language, the shared contract prefix and output keys remain English.
Best Value
- Great extension activities for science and biology
- Correlated to standards
- Comprehensive biology vocabulary study
- Fascinating true-to-life illustrations
How to interpret the benchmark
The author describes Bilingual Patch Contracts as “a small diagnostic benchmark, not a general model ranking.” That is the right scale for its conclusions:
- Twelve underlying scenarios, not 36 independent semantic problems. The three language versions are paired observations of each scenario.
- One run is not a production reliability estimate. The report does not establish how often a model will succeed across different prompts, deployments, or future runs.
- A perfect score is a ceiling, not proof of broader reliability. Gemini 3.7 Flash passed all 36 prompts, but this suite cannot distinguish its performance beyond the examples it contains.
- Language coverage is bounded. English contract instructions and output keys constrain what can be inferred about fully Chinese interactions.
- Some operational dimensions are outside scope. Latency, cost, and tool calling were not benchmarked.
Version 2 also corrected task registration so Kaggle selects the whole-suite aggregate rather than a helper function. The prompts, fixtures, and scorer were unchanged. One numeric task scores strict exact matches divided by 36, so the overall score equals that task score; infrastructure errors abort the suite instead of silently reducing its denominator.
What a useful patch evaluation should report
This benchmark illustrates why a single “JSON success” number is inadequate for state-update systems. Evaluations should distinguish whether the response is parseable, whether it satisfies the schema, and whether the resulting state is exactly right. They should also make clear whether presentation failures count, how retries or repairs are handled, and whether repeated language variants are being treated as independent cases.
For developers consuming model output, the practical implication is to validate each layer the application depends on: parse the complete response, check its keys and types, then verify the state transition against explicit expected values or invariants. A successful parse is an interface checkpoint—not evidence that the requested update was applied correctly.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




