Driver FixRecommendedSound, Wi-Fi or graphics acting up? Check drivers firstFind missing or outdated drivers fast.Check DriversOctober DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsPC HealthRecommendedCrashes, freezes, slowdowns? Check your PC nowSpot repairable issues before they interrupt work.Check PC×
Skip to content
Computer Vision

How to Extract Text From Images With an LLM (Python Guide)

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Direct answer: send the image to a vision-capable language model with an explicit transcription instruction, then verify the response against the pixels. Ask it to preserve line breaks or columns when needed and to mark unreadable characters instead of guessing. Image quality, rotation, resolution and the model’s image-detail setting all affect the result; even capable systems can misread small or unusual text.

This guide shows a practical Python workflow, equivalent cURL and Node.js requests, validation techniques, and when dedicated OCR is a better fit.

What an LLM can and cannot do

Vision-capable models accept an image alongside text and can return a transcription or answer questions about visible words. OpenAI documents PNG, JPEG, WEBP and non-animated GIF inputs; Gemini documents PNG, JPEG, WEBP, HEIC and HEIF. Confirm the formats and limits for the exact model and endpoint you deploy because they change. See the OpenAI image and vision guide and Gemini image-understanding guide.

This is not guaranteed optical character recognition. OpenAI states that “Vision models can make mistakes.” Small type, rotation, handwriting, unusual fonts and some non-Latin scripts are particularly easy to misread. Treat names, dates, amounts, serial numbers and legal text as untrusted until checked character by character.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Prepare an image that is readable

Improve the source first

  • Use a sharp, in-focus image with even lighting and high contrast.
  • Rotate it so text is upright. Google specifically recommends checking rotation.
  • Crop away irrelevant borders and enlarge the region containing fine print. Cropping can also reduce token use.
  • For a multi-column page, keep enough surrounding context for reading order, or process each column separately if the model interleaves them.

Choose detail and resolution deliberately

OpenAI recommends original detail for fine visual tasks such as OCR when that control is available. Gemini notes that higher resolution can improve fine-text reading but increases token usage and latency. “Original” does not mean unlimited pixels: providers may still resize an image to model limits. Check the current documentation before relying on a setting.

A reliable transcription prompt

Ask for transcription, not a summary. This prompt is a useful starting point (it is guidance, not a guarantee):

Transcribe all visible text exactly.
Preserve line breaks and reading order where practical.
Do not infer unreadable characters; write [unclear] instead.
Return only the transcription, with no commentary.

Add task-specific rules when necessary: “Keep table rows separated,” “Label each column,” or “Preserve capitalization and punctuation.” If you need structured output, request a clearly defined JSON shape and still compare it with the image.

Python: send an image to a vision model

The following example uses the OpenAI Responses API shape documented in the official vision guide. Set OPENAI_API_KEY and, if required by your account, OPENAI_MODEL to a currently available vision-capable model. The image is embedded as a data URL, so no public image hosting is needed.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import base64
import mimetypes
import os
import requests

IMAGE_PATH = "receipt.jpg"
MODEL = os.environ.get("OPENAI_MODEL", "gpt-4.1-mini")

mime, _ = mimetypes.guess_type(IMAGE_PATH)
if mime not in {"image/png", "image/jpeg", "image/webp", "image/gif"}:
    raise ValueError("Use a supported PNG, JPEG, WEBP, or non-animated GIF image")

with open(IMAGE_PATH, "rb") as f:
    encoded = base64.b64encode(f.read()).decode("ascii")

payload = {
    "model": MODEL,
    "input": [{
        "role": "user",
        "content": [
            {"type": "input_text", "text": (
                "Transcribe all visible text exactly. Preserve line breaks where practical. "
                "Do not guess unreadable characters; mark them [unclear]. Return only the transcription."
            )},
            {"type": "input_image", "image_url": f"data:{mime};base64,{encoded}"}
        ]
    }]
}

response = requests.post(
    "https://api.openai.com/v1/responses",
    headers={
        "Authorization": f"Bearer {os.environ['OPENAI_API_KEY']}",
        "Content-Type": "application/json",
    },
    json=payload,
    timeout=90,
)
response.raise_for_status()
data = response.json()

# Responses API output is an array of content blocks; print text blocks.
for item in data.get("output", []):
    for block in item.get("content", []):
        if block.get("type") == "output_text":
            print(block.get("text", ""))

Install the only dependency with python -m pip install requests. If your selected model or API version uses different field names, follow that provider’s current image-input example rather than silently changing the payload.

Equivalent cURL request

For a publicly reachable image URL, the request can be shorter. Keep the image URL stable and accessible to the provider:

curl https://api.openai.com/v1/responses 
  -H "Authorization: Bearer $OPENAI_API_KEY" 
  -H "Content-Type: application/json" 
  -d '{
    "model":"'"${OPENAI_MODEL:-gpt-4.1-mini}"'",
    "input":[{"role":"user","content":[
      {"type":"input_text","text":"Transcribe all visible text exactly. Preserve line breaks where practical. Mark unreadable characters [unclear]. Return only the transcription."},
      {"type":"input_image","image_url":"https://example.com/page.jpg"}
    ]}]
  }'

Replace the example URL with an image you control and use the provider’s documented URL-access requirements.

Equivalent Node.js request

const imageUrl = "https://example.com/page.jpg";
const model = process.env.OPENAI_MODEL || "gpt-4.1-mini";
const response = await fetch("https://api.openai.com/v1/responses", {
  method: "POST",
  headers: {
    "Authorization": `Bearer ${process.env.OPENAI_API_KEY}`,
    "Content-Type": "application/json"
  },
  body: JSON.stringify({
    model,
    input: [{ role: "user", content: [
      { type: "input_text", text: "Transcribe all visible text exactly. Preserve line breaks where practical. Mark unreadable characters [unclear]. Return only the transcription." },
      { type: "input_image", image_url: imageUrl }
    ] }]
  })
});
if (!response.ok) throw new Error(`${response.status}: ${await response.text()}`);
const data = await response.json();
for (const item of data.output ?? [])
  for (const block of item.content ?? [])
    if (block.type === "output_text") process.stdout.write(block.text);

Validate before using the text

  1. Display the source beside the transcription at the same zoom level.
  2. Check every high-impact token character by character: account numbers, URLs, decimal points, minus signs, dates and capitalization.
  3. Look for dropped lines, duplicated lines and columns returned in the wrong order.
  4. Run a second pass with a focused prompt such as “Compare this proposed transcription with the image and list only discrepancies.” Treat that pass as another aid, not proof.
  5. For a critical record, have a person approve the final text and retain the original image.

Do not allow the model to replace uncertainty with plausible spelling. An explicit [unclear] marker is safer than a confident error.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

LLM or dedicated OCR?

Use an LLM when reading is part of a broader interpretation task—for example, answering a question about a label, explaining a diagram, or extracting a few fields from a mixed image. Use dedicated OCR or document processing when you repeatedly transcribe many pages, require bounding boxes, or need stable document structure.

Google Cloud Vision separates TEXT_DETECTION, which returns text and individual words with boxes, from DOCUMENT_TEXT_DETECTION, which is aimed at dense documents and exposes page, block, paragraph, word and break structure. Google directs scanned-document OCR, structured forms and entity extraction toward Document AI. See the Cloud Vision OCR guide.

Requirement Practical choice
One-off question about text in a photo Vision-capable LLM
High-volume, repeatable transcription Dedicated OCR pipeline
Word boxes or reading-order structure OCR/document service
Exact legal, financial or identifier data Either approach plus human or deterministic validation

Available documentation does not establish a universal accuracy winner, pricing comparison or privacy ranking. Evaluate representative images from your own scripts, layouts and error tolerance. Compare exact-character accuracy, small or rotated text, handwriting, non-Latin scripts, reading order, supported formats, latency, cost and data handling.

Troubleshooting common failures

The request is rejected

Check the model name, authentication header, JSON shape and supported media type. Provider APIs evolve; copy the current image-input schema from the linked guide.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The output is blank or incomplete

Verify that the image bytes or URL are valid, the URL is reachable by the provider, and the image is not an unsupported animated format. Crop and enlarge the text, then retry with higher detail or resolution where available.

Columns are scrambled

Crop to one column at a time or explicitly request column labels and reading order. Preserve a copy of the full page so the result can be reconstructed.

Characters are consistently wrong

Improve focus and contrast, rotate the image, and supply a closer crop. For dense scans or scripts the model handles poorly, move to a document OCR service and validate its output.

Requests are slow or expensive

Downscale irrelevant regions, crop aggressively and avoid sending the same image repeatedly. Higher resolution and original-detail modes consume more tokens and can increase latency. Cache results using a hash of the image and prompt when your data policy permits.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

If the image is on a web page and you first need a clean capture, ScreenshotNeo provides a screenshot API and MCP server. One GET request returns PNG, JPEG, WebP or PDF; it can accept cookies or consent banners like a visitor and remove more than 60 known consent platforms, newsletter popups and chat widgets before capture. Bot checks, blank pages, timeouts, failed loads and cache hits are not billed, and response headers identify the page verdict and billing status.

Use the documented API examples at ScreenshotNeo docs, then pass the resulting image to your LLM:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

ScreenshotNeo also offers an MCP server with take_screenshot, get_page_info and capture_pdf for Claude, Cursor and other MCP clients. Every feature is on every plan: 1,000 shots per month are free with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.

Operational checklist

  • Record the model, image format, detail setting and prompt with each transcription.
  • Keep the original file and the returned text together.
  • Redact sensitive images when your organization’s policy requires it.
  • Measure errors on your own sample before automating downstream decisions.
  • Route uncertain or high-impact records to human review.

Frequently Asked Questions

Can an LLM preserve the exact layout of a document?

It can be instructed to preserve line breaks, columns or table rows, but layout fidelity is not guaranteed. For bounding boxes and document structure, use a service that explicitly returns those fields.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Should I convert a PDF page to an image first?

Only when the provider or your workflow requires image input. If the source is already a scanned document, a document OCR service may preserve structure more reliably than converting pages and asking for free-form text.

Is a second LLM pass enough to prove the transcription is correct?

No. A second pass can find discrepancies, but exactness still requires comparison with the original image and, for consequential text, human approval.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Read next

Recommended PC Tool
Recommended PC Tool
Crashes, No Sound, or Screen Glitches?Free driver scan
PC Slower Than It Used to Be?Free scan - under a minute

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.