October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsClean PCRecommendedOne scan can reveal what keeps slowing WindowsLook for cleanup and repair opportunities.Run ScanOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
data extraction

How to Automatically Extract Structured Information from Unstructured Text

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

To automatically extract structured information from unstructured text, first define the fields you need, then choose an extraction method that fits the source, and finally validate each returned value against the original. JSON that matches a schema is easier to parse; it is not proof that the values are accurate.

Start with the record you need

“Unstructured text” can mean anything from a paragraph in an email to a scanned contract. The reliable way to turn it into records is to decide what a record means before picking a model or service. For example, a support-message record might contain a customer name, product, issue, requested action, and date. A contract record might contain parties, effective date, renewal terms, and governing jurisdiction.

Write down the fields and their rules. For each field, specify its type, whether it is required, whether multiple values are allowed, and what to do when the text does not contain a clear answer. Distinguish “not present” from “unclear” if your workflow needs to handle those cases differently. Do not make the extractor guess a required value simply because your downstream database expects one.

  • Required: must be present for a usable record, or trigger a review/error state.
  • Optional: may be omitted or represented as null when absent, according to your chosen output convention.
  • Repeated: represented as an array, even if a particular document contains only one item.
  • Constrained: limited to a defined set of values, such as a status enum.

Also decide whether the output must retain evidence: for example, a text span, page number, or source location supporting an important value. Auditability can be essential for regulated or consequential decisions.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Identify the input before choosing the extractor

Different inputs require different stages. Clean digital prose can often go directly to semantic extraction. A scan cannot: it first needs optical character recognition (OCR), and a form or table may require layout analysis to preserve the relationship between labels, values, rows, and columns.

Input Likely first step What to watch
Digital prose, email, or plain-text notes Normalize the text, then extract fields Ambiguous references, inconsistent date formats, long inputs
Scanned page or image OCR, then semantic mapping Recognition errors, page orientation, low contrast, reading order
Forms and semi-structured documents OCR/layout or form analysis, then map to your schema Key-value pairing, repeated sections, missing fields
Tables Table/layout extraction, then map rows and columns Header association, merged cells, multi-page tables

For document-like inputs, AWS Textract describes AnalyzeDocument operations for detected text, forms, tables, query responses, and signatures; its response represents relationships such as form keys and values. That can supply a useful document-analysis stage, but mapping its output into a custom semantic record still needs to be designed and evaluated. See Textract analysis and Textract response objects.

Choose the extraction approach

There are three overlapping approaches, but they address different needs. Feature lists alone do not establish which will work best on your documents; compare candidates on representative examples from your own corpus.

Approach Best fit Evaluate
Schema-constrained LLM output Custom fields that depend on context or interpretation in prose Field accuracy, schema support, treatment of absent or ambiguous evidence, latency, cost, privacy, integration
Named-entity analysis Recognizing supported entity classes, such as people, organizations, or locations Supported entity types, language and domain fit, precision and recall, offsets and metadata, integration
Document-analysis/OCR service Scans, forms, tables, and inputs where layout matters OCR and layout accuracy on your actual files, representation of forms/tables, customization, throughput, cost, data handling

Schema-constrained output for custom fields

OpenAI’s Structured Outputs documentation says, “You can define structured fields to extract from unstructured input data, such as research papers.” It describes using a defined schema to constrain response shape. The documentation distinguishes this from function calling, which connects a model to application functions. For an extraction pipeline, the model can produce a record and your application can validate and store it; a function call is useful when the model must invoke an application action. Check current model support and the supported JSON Schema subset before implementation. A valid shape does not establish semantic truth. OpenAI Structured Outputs guide; OpenAI Function Calling.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Google’s Gemini documentation also describes JSON Schema-constrained output, including text extraction examples. This output-format feature is distinct from Google Cloud Natural Language’s entity-analysis API, which recognizes supported entities and returns associated information. Choose based on whether you need arbitrary custom fields or predefined entity analysis. Gemini structured outputs; Google Cloud Natural Language basics; analyzeEntities API reference.

Example: request a typed record with Python

The example shows the shape of a schema-constrained extraction request using OpenAI’s Responses API format. It uses the model identifier named in OpenAI’s August 2024 Structured Outputs announcement; check current availability and supported schema features for your account and model before relying on it. Install the SDK with pip install openai and set OPENAI_API_KEY in the environment. In production, test strict-mode compatibility with the current API and keep the schema narrow.

import json
from openai import OpenAI

client = OpenAI()
text = "On 12 March 2025, Acme Supply reported that order 4831 arrived late."
schema = {
    "type": "object",
    "properties": {
        "company": {"type": "string"},
        "order_id": {"type": "string"},
        "issue": {"type": "string"},
        "reported_date": {"type": "string"},
        "evidence": {"type": "string"}
    },
    "required": ["company", "order_id", "issue", "reported_date", "evidence"],
    "additionalProperties": False
}

response = client.responses.create(
    model="gpt-4o-2024-08-06",
    input=[
        {"role": "system", "content": (
            "Extract only information supported by the text. "
            "If a value is absent or unclear, return an empty string. "
            "Evidence must be a short verbatim supporting span."
        )},
        {"role": "user", "content": text}
    ],
    text={"format": {
        "type": "json_schema",
        "name": "issue_record",
        "strict": True,
        "schema": schema
    }}
)
record = json.loads(response.output_text)
print(json.dumps(record, indent=2))

This returns a parseable record only if the request and model support the selected schema format. The empty-string convention is a deliberate choice for this example, not a universal missing-value standard. In a real application, consider separate states for absent and ambiguous values and validate the evidence span against the original input.

Validate meaning, not just JSON

Treat extraction as a pipeline: parse the response, enforce field rules, verify support in the source, apply business rules, then save or route the record. A successful parse should never be the only acceptance check.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  1. Validate structure: confirm required keys, allowed properties, data types, array shapes, and permitted enum values.
  2. Validate source support: check that each important value appears in or is reasonably entailed by the input. For high-stakes fields, retain the source span or page location.
  3. Normalize carefully: standardize dates, currencies, and identifiers only under explicit rules. Preserve the original text when normalization could lose meaning.
  4. Apply cross-field rules: verify constraints such as a start date preceding an end date, totals reconciling, or an identifier matching an expected format.
  5. Route uncertainty: reject, request a second pass, or send for human review when required data is absent, contradictory, or ambiguous.

Do not infer missing facts from general knowledge unless the task explicitly authorizes enrichment and your system records that the value was inferred rather than extracted. Keep extraction and enrichment distinguishable.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Evaluate on examples from your own corpus

Before production, assemble representative documents and have people label the expected fields. Include ordinary cases as well as scans, unusual layouts, missing fields, conflicting statements, and ambiguous references if those occur in practice. Keep a held-out set so prompt or configuration changes can be checked against examples that were not used to tune the workflow.

  • Field-level precision: of the values the system returned, how many were correct?
  • Field-level recall: of the values that should have been extracted, how many did it find?
  • Schema validity: how often did outputs satisfy structural and allowed-value rules?
  • Error categories: separate OCR mistakes, omitted fields, unsupported guesses, normalization errors, and incorrect entity links.
  • Operational fit: measure latency, cost, throughput, privacy/data-handling fit, and integration effort for the actual workflow.

Report results by field and document type, not only as one overall score: an average can conceal that a critical field fails on a particular class of files. These sources do not establish a universal best service or independent comparative accuracy, so a corpus-specific evaluation is necessary.

OpenAI’s August 6, 2024 launch announcement reported 100% on its complex JSON-schema-following evaluation for gpt-4o-2024-08-06 and less than 40% for gpt-4-0613. Those are OpenAI-reported results for that schema-following evaluation, not an independent comparison and not a claim of perfect factual extraction from arbitrary text. OpenAI’s announcement.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Best Value

Handle production failures deliberately

  • Malformed or rejected schema request: inspect the API error and compare the schema with the provider’s current supported subset. Simplify unsupported constructs rather than silently dropping validation.
  • Valid JSON but wrong value: treat as a semantic extraction error. Tighten field definitions, require source evidence, add a validation rule, or send that case to review.
  • Empty or missing data: do not convert absence to a plausible guess. Preserve the agreed missing-value representation and check whether the source actually contains the field.
  • Scan yields garbled text: fix OCR and image quality/layout handling before tuning semantic extraction. The model cannot reliably recover content that OCR did not capture.
  • Long document is truncated or misses distant context: check input limits and how text is divided. If chunking is needed, preserve page/section context and reconcile duplicates or conflicting facts across chunks.
  • Records disagree across runs: retain the source and extraction configuration, compare outputs on a fixed evaluation set, and add deterministic business validation. Do not assume a retry will resolve an ambiguity.
  • Latency or cost is too high: measure by document type and stage. Consider simpler entity analysis for predefined entity needs, avoiding unnecessary OCR on digital text, or batching only where the chosen API supports it and the workflow can tolerate the delay.

Or skip the browser setup

ScreenshotNeo is a website screenshot API, not a text-extraction service; it is useful only if a URL screenshot is an input you need to capture before another system processes it. One GET request returns an image or PDF. See the ScreenshotNeo site and API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

ScreenshotNeo can remove cookie/consent banners, newsletter popups, and chat widgets before the shot. Bot checks, blank pages, and failed loads are never billed; an MCP server lets AI agents take screenshots; and 1,000 screenshots per month are free with no card, with paid plans starting at $5 for 3,000. Sign up for the free plan.

Make the workflow auditable

Store enough context to reproduce and investigate an extraction: the source document or a secure reference to it, the extracted record, validation outcomes, and the versioned schema and extraction configuration. Apply access controls and retention rules appropriate to the information being processed. Privacy, compliance suitability, and vendor data handling depend on your deployment and must be assessed for your use case; the documentation cited here does not settle them.

The durable principle is simple: specify the record first, use the right upstream handling for the input, and validate extracted values against evidence. Schema-constrained output makes structured data easier to consume, but reliability comes from evaluation and checks around it.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Frequently Asked Questions

What is the difference between structured output and function calling?

Structured output constrains the shape of a model response; function calling is for connecting a model to application functions or actions.

Can OCR alone extract custom business fields?

OCR turns image content into text and may preserve layout, but a separate mapping or semantic extraction step is often needed to produce custom fields.

Should I use an LLM or a named-entity API?

Use a corpus-specific evaluation: custom contextual fields favor testing schema-constrained extraction, while predefined supported entity classes may fit entity analysis better.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Read next

Recommended PC Tool
Recommended PC Tool
PC Slower Than It Used to Be?Free scan - under a minute
Outdated Drivers Are Slowing You DownFree scan - exact matches

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.