DriversRecommendedOutdated drivers can make a good PC feel brokenScan driver issues before chasing fixes manually.Scan NowOctober DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsPC HealthRecommendedCrashes, freezes, slowdowns? Check your PC nowSpot repairable issues before they interrupt work.Check PC×
Skip to content
World desk6 min

Extracting Reliable Structured Data from LLMs

Schema-constrained output can make LLM responses easier to parse, but reliable extraction also requires checking every field against the source and testing difficult cases.
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Reliable LLM extraction requires two separate checks: confirm the response matches the required structure, then verify that every value is supported by the source. JSON mode or schema-constrained output can reduce formatting errors, but a valid, schema-shaped response can still contain omissions, wrong values, or invented facts.

What “reliable structured extraction” actually requires

Suppose you ask a model to read an invoice, manual, or support ticket and return fields such as a date, part number, and status. There are two different questions to answer:

  • Is the response structurally usable? Can your application parse it, and does it have the required keys and value types?
  • Is the extracted information correct? Does each value accurately reflect the source, with no unsupported details or incorrect associations?

Schema controls address the first question. They do not, by themselves, answer the second. OpenAI describes JSON mode as producing valid JSON without guaranteeing conformity to a particular schema; its Structured Outputs feature is designed to enforce schema adherence. Neither guarantee establishes factual accuracy. See the OpenAI Structured Outputs guide and the August 6, 2024 announcement.

Choose the output mechanism for the job

Use the mechanism that matches what the model is supposed to do. Provider names and supported schema features can change, so check the current documentation for the model and API version you plan to deploy.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Mechanism Use it when What it establishes
Prompting for “JSON only” You have no suitable structured-output feature or are testing a simple prototype. The instruction alone does not ensure parseable JSON or conformance to an exact schema.
JSON mode You need valid JSON, but do not need the API to enforce a particular schema. Valid JSON is not the same as exact schema adherence. OpenAI makes this distinction in its Structured Outputs announcement.
Schema-constrained response formatting The assistant’s response is itself the structured result your application will consume. Where supported and correctly configured, it constrains output to a specified schema; it does not verify that values are true or source-grounded. See OpenAI’s guide and Anthropic’s documentation.
Tool or function calling The model needs to invoke a function or pass arguments to a tool. The arguments can be constrained by a tool schema, but successful argument formatting does not prove the arguments are appropriate or correct.

In short: use tool or function calling to trigger an action; use a structured response format when the answer itself should be schema-shaped. Anthropic’s platform documentation describes structured outputs as constraining Claude’s responses to follow a schema for valid, parseable downstream output. That is a structural guarantee, not a claim of semantic correctness.

Define the destination contract before prompting

Write down what the receiving application can accept before asking a model to extract anything. A schema is useful only if it reflects the actual downstream contract.

  • Fields and types: specify every required key and whether its value is a string, number, boolean, array, object, or another supported type.
  • Missing information: decide whether an absent value should be represented as null, an empty collection, a status such as “unknown,” or an omitted key. Make the rule explicit; do not leave the model to choose.
  • Allowed values: use an explicit set of permitted values where the application requires one, and define what to return when the source does not support a choice.
  • Extra keys: decide whether unexpected fields are permitted or rejected.
  • Meaning and evidence: describe ambiguous fields and, where useful, ask for a source quotation or location alongside the normalized value.

OpenAI recommends clear, intuitive key names, descriptions for important keys, and evaluations tailored to the use case in its Structured Outputs guide. Descriptions should settle interpretation, not merely repeat a field name. For example, define whether “date” means the date printed on a document, its due date, or the date it was received.

Build the extraction pipeline in stages

  1. Prepare the input. Pass the relevant source text to the model and retain a stable reference to the original document or passage. If the source is incomplete or unreadable, the extraction process needs a way to report that rather than fill gaps.
  2. Request only the contract’s fields. Tell the model how to handle absent or ambiguous information and require the selected structured-output feature or tool schema where appropriate. Do not rely on a “JSON only” prompt as a substitute for schema enforcement.
  3. Check the response ending. Treat a refusal or incomplete response, such as one cut off at an output limit, as an unsuccessful extraction—not as a completed record. OpenAI’s guide documents refusal and incomplete-output cases that can leave the expected result absent or incomplete.
  4. Validate structure in your application. Parse the response and check required fields, types, allowed values, and extra-key policy. Reject or quarantine results that do not meet the contract.
  5. Validate meaning against the source. Check that each value is present in, inferable from, or correctly normalized from the input according to your rules. Flag unsupported claims, missing values, and values attached to the wrong field.
  6. Apply downstream safeguards. For consequential actions, route uncertain or failed records to a review path instead of treating schema validity as authorization to act.

A schema can catch a date returned in the wrong type; it cannot necessarily tell whether the model selected the due date instead of the issue date. That second check needs source-grounded rules and, for evaluation, expected answers.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #3
Sale
Storytelling with Data: A Data Visualization Guide for Business Professionals
  • Wiley
  • Language: english
  • Book - storytelling with data: a data visualization guide for business professionals

Measure structure and semantic accuracy separately

Do not use “valid JSON” as the success metric for an extraction system. Evaluate at least two layers independently:

  • Structural measures: parse success, schema adherence, required-field coverage, type validity, and behavior on refusal or truncation.
  • Semantic measures: field-level correctness against source-grounded expected values, omissions, unsupported values, incorrect normalization, and values assigned to the wrong fields.

Build an evaluation set from representative inputs, including difficult cases: missing fields, conflicting or ambiguous statements, unusual formatting, and values that should not be inferred. Keep expected outputs grounded in the source. Test not only whether a field is filled, but whether the chosen value and its association with that field are right.

Include schema changes in the evaluation process. When you add, remove, rename, or redefine fields—or change provider or model versions—rerun the relevant tests. The 2026 StructHallu-Drift study examines schema evolution and reports model- and output-format-specific error patterns in its tested setting; it is a reason to test changes, not a universal prediction of failure.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What published results do—and do not—show

Published benchmarks illustrate why format compliance and factual extraction must remain separate:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • OpenAI reported 100% adherence on its complex JSON Schema evaluation for GPT-4o-2024-08-06 with Structured Outputs, compared with less than 40% for GPT-4-0613. These are provider-reported results for that evaluation and those models, not factual extraction accuracy or a universal guarantee. See the August 6, 2024 announcement.
  • The January 2025 JSONSchemaBench paper includes 10,000 real-world JSON schemas and evaluates constrained decoding on efficiency, coverage of constraint types, and output quality. Its dimensions are useful when assessing whether a method supports the schema features and operating conditions your application needs.
  • The StructHallu-Drift paper, published in the ACL workshop proceedings in July 2026, reports at least one semantic hallucination in 39–54% of structured outputs across its tested 1,200 schema-model evaluation instances, four models, and three tasks. It also reports approximately 85% semantic validity for SQL and 7–24% for schema-grounded record generation in that study’s setup. These task-specific findings are not universal rates and should not be generalized into a general comparison of SQL and record extraction.

The practical lesson is not that one format or provider always wins. It is that syntactic constraints can improve output shape while leaving semantic errors to be measured and handled separately.

Compare providers and approaches on the same task

There is no basis in these sources for declaring one current provider or constrained-decoding framework the overall winner. A useful comparison uses the same representative inputs, destination schema, and expected answers, then checks:

  • How often the output parses and conforms to the schema.
  • How accurately each field reflects the source, including missing and ambiguous cases.
  • Whether the implementation supports the schema features your contract actually uses.
  • What happens with refusals, truncation, invalid inputs, and absent information.
  • Latency, efficiency, and integration overhead for your workload.

JSONSchemaBench explicitly considers efficiency, constraint coverage, and output quality; semantic evaluation and task-specific behavior also matter, as the StructHallu-Drift findings illustrate. Provider documentation accessed October 5, 2026 may change, so verify current syntax, supported schema subsets, model availability, and refusal or truncation behavior before deployment.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More from the Wire

  1. World desk4 min
    How to Spot an AI Voice Scam Before Sending MoneyDon’t rely on how a caller sounds. Pause, call back through a known number, and verify the emergency with another trusted person before sending money.
  2. Mountain View desk4 min
    Google’s SynthID Detector: How to Check AI-Generated Images, Video and AudioGoogle’s SynthID Detector looks for an embedded watermark in supported images, video and audio. Here is what its results do—and do not—show.
  3. Redmond desk20 min
    How to create a link to File or Folder in Windows 11Windows 11 gives you several ways to point to a file or folder without moving or duplicating it. You can create a desktop shortcut,…
Recommended PC Tool
Recommended PC Tool
PC Slower Than It Used to Be?Free scan - under a minute
Outdated Drivers Are Slowing You DownFree scan - exact matches

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.