October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsClean PCRecommendedOne scan can reveal what keeps slowing WindowsLook for cleanup and repair opportunities.Run ScanOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
World desk5 min

Valid JSON Is Not Enough: What Bilingual Patch Contracts Test on Kaggle

Bilingual Patch Contracts tests both strict JSON output and exact state updates across 12 scenarios in English, Chinese and code-switched instructions. Its October 2026 results are a diagnostic, not a general model ranking.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A model can return perfectly parseable JSON and still apply the wrong update. A Kaggle benchmark reported on October 1, 2026, tests both sides of that problem: whether a response obeys a strict JSON interface and whether its values match the exact expected state. Its results are useful as a diagnostic, not as a general ranking of models.

What “valid JSON” leaves unanswered

Parsing checks syntax: can software read the response as JSON? Schema checks shape: does it have the required fields and types? Neither check, by itself, establishes that the response changed the state correctly.

As an Amazon Associate I earn from qualifying purchases.

For a patch task, the model must interpret instructions, apply them to an initial state, and return the resulting values in the required format. A response can therefore fail in at least two different ways:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Wrong state, valid structure: the output parses and has the right keys and types, but one or more values are wrong.
  • Right state, unusable presentation: the values are correct, but the response includes extra material—such as Markdown fences—so a strict consumer cannot parse the complete response as the required JSON document.

The benchmark author summarizes the first distinction this way: “A JSON response can parse successfully and still change the wrong state.”

What the Kaggle benchmark tests

Bilingual Patch Contracts consists of 12 handcrafted state-update scenarios. Each scenario appears with an English, Chinese, and code-switched instruction body, for 36 prompts total. Within each three-prompt group, the initial state and expected answer are shared. The contract prefix and canonical output keys remain in English, so this is not a fully Chinese interaction benchmark.

The scenarios probe different ways instructions can go wrong, including:

  • applying a later correction rather than an earlier value;
  • respecting negation and distinguishing null from an empty value;
  • preserving tag order and case sensitivity;
  • converting hours to minutes and applying sequential conditions;
  • treating instruction-like text as literal data rather than as a new command; and
  • copying Unicode, backslashes, quotation marks, and a newline exactly.

What counts as a pass

A response passes only if the complete response is one JSON object with exactly five keys, the required value types, and every expected value. The scorer does not remove Markdown, repair the output, or ask a second model to judge it.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some representational differences are accepted: whitespace, key order, and equivalent Unicode escapes do not change the result. Others fail: duplicate keys, extra fields, nonfinite values, booleans or floats where an integer is required, and arrays whose elements are in the wrong order. This strictness makes the score about exact contract compliance, not merely whether a human can infer the intended answer.

How the run was made

For the October 1, 2026 run, the author used ordinary text generation, requested temperature 0 and seed 0 through the SDK, and started a fresh isolated conversation for each case. The run did not use constrained JSON decoding, schema enforcement, or tools. The author reports checking all 36 case IDs against frozen prompts and answers and independently recalculating saved scores. Provider behavior can vary across runs, so these results describe that run rather than a guarantee of repeatability.

Reported results from the October 1, 2026 run

The following are the benchmark author’s reported results for version 2. Exact match means meeting the full strict pass rule; valid JSON and valid schema are separate measurements.

Model Strict exact match Valid JSON Valid schema
Gemini 3.7 Flash 36/36 (100%) 36/36 36/36
GPT-5.4 nano 24/36 (66.7%) 36/36 36/36
Claude Haiku 4.5 0/36 (0%) 0/36 0/36
Qwen3-Next-80B-A3B-Instruct No complete score No complete score No complete score

Qwen3-Next-80B-A3B-Instruct was attempted in both the pilot and version 2, but the attempts stopped with HTTP 429 and a provider heavy-load message. It was excluded rather than scored as zero. The report does not provide a complete per-language score table for all listed models; the figures above are aggregate results across the 36 prompts.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Why the two failure types matter

GPT-5.4 nano: valid output, incorrect updates

GPT-5.4 nano produced valid JSON and valid field types on all 36 cases, yet only 24 outputs matched the expected state exactly. In the case-sensitive tags scenario, for example, it retained lowercase beta even though the instruction required removing it. A parser and a type checker would accept that response; an application relying on the updated state would receive the wrong value.

Claude Haiku 4.5: Markdown breaks a strict interface

Claude Haiku 4.5 wrapped every answer in a Markdown code fence despite an explicit instruction not to. Under the benchmark rule, the full response is not a JSON document, even if the JSON text inside the fence contains correct values.

The author reports a separate counterfactual: stripping only complete outer fences would make 33 of 36 responses pass value checks. That diagnostic is not the benchmark score, and it does not change the reported results. In a real system, silently stripping output is a design choice with its own consequences; it should not be confused with the model having followed a raw-JSON contract.

What the language comparisons do—and do not—show

For GPT-5.4 nano, the mixed-language total was two cases higher than the English total. Looking at matched scenarios gives a narrower picture: seven passed in both English and mixed-language forms, three failed in both, and two passed only in the mixed form. English-versus-Chinese comparisons were also mixed.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

These paired results point to particular cases worth inspecting; they do not establish that the model is generally stronger in Chinese or code-switching. The instruction bodies were hand-authored, and their phrasing and token lengths were not perfectly controlled. Also, while the bodies vary by language, the shared contract prefix and output keys remain English.

Best Value
Carson Dellosa The 100 Series: Biology Workbook—Grades 6-12 Science, Matter, Atoms, Cells, Genetics, Elements, Bonds, Classroom or Homeschool Curriculum (128 pgs)
  • Great extension activities for science and biology
  • Correlated to standards
  • Comprehensive biology vocabulary study
  • Fascinating true-to-life illustrations
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How to interpret the benchmark

The author describes Bilingual Patch Contracts as “a small diagnostic benchmark, not a general model ranking.” That is the right scale for its conclusions:

  • Twelve underlying scenarios, not 36 independent semantic problems. The three language versions are paired observations of each scenario.
  • One run is not a production reliability estimate. The report does not establish how often a model will succeed across different prompts, deployments, or future runs.
  • A perfect score is a ceiling, not proof of broader reliability. Gemini 3.7 Flash passed all 36 prompts, but this suite cannot distinguish its performance beyond the examples it contains.
  • Language coverage is bounded. English contract instructions and output keys constrain what can be inferred about fully Chinese interactions.
  • Some operational dimensions are outside scope. Latency, cost, and tool calling were not benchmarked.

Version 2 also corrected task registration so Kaggle selects the whole-suite aggregate rather than a helper function. The prompts, fixtures, and scorer were unchanged. One numeric task scores strict exact matches divided by 36, so the overall score equals that task score; infrastructure errors abort the suite instead of silently reducing its denominator.

What a useful patch evaluation should report

This benchmark illustrates why a single “JSON success” number is inadequate for state-update systems. Evaluations should distinguish whether the response is parseable, whether it satisfies the schema, and whether the resulting state is exactly right. They should also make clear whether presentation failures count, how retries or repairs are handled, and whether repeated language variants are being treated as independent cases.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For developers consuming model output, the practical implication is to validate each layer the application depends on: parse the complete response, check its keys and types, then verify the state transition against explicit expected values or invariants. A successful parse is an interface checkpoint—not evidence that the requested update was applied correctly.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More from the Wire

  1. World desk4 min
    How to Spot an AI Voice Scam Before Sending MoneyDon’t rely on how a caller sounds. Pause, call back through a known number, and verify the emergency with another trusted person before sending money.
  2. Mountain View desk4 min
    Google’s SynthID Detector: How to Check AI-Generated Images, Video and AudioGoogle’s SynthID Detector looks for an embedded watermark in supported images, video and audio. Here is what its results do—and do not—show.
  3. Redmond desk20 min
    How to create a link to File or Folder in Windows 11Windows 11 gives you several ways to point to a file or folder without moving or duplicating it. You can create a desktop shortcut,…
Recommended PC Tool
Recommended PC Tool
PC Slower Than It Used to Be?Free scan - under a minute
Crashes, No Sound, or Screen Glitches?Free driver scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.