Hardware FixRecommendedDevice not working? Your driver may be the problemCheck updates for common hardware issues.Fix DriversOctober DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsSlow PC?RecommendedPC slow today? Run a repair scan before it gets worseResolve common Windows issues and optimize system performance.Scan Now×
Skip to content
Azure AI Vision

How to Use an Image API for OCR Text Extraction

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

To extract text from an image, call an OCR or vision API—not an image-search endpoint that finds visually similar pictures. Google Cloud Vision offers TEXT_DETECTION for text in ordinary images and DOCUMENT_TEXT_DETECTION for dense pages with document structure. Azure AI Vision Read is another managed option, especially if your application already uses Azure. Both return text from an image; neither is an image-similarity search API.

Image search and OCR solve different problems

An image-search API looks for images that resemble a query image or match visual criteria. OCR—optical character recognition—identifies characters visible inside an image and returns them as text. If you need to read a receipt, sign, screenshot, photograph, or scanned page, choose an OCR feature or vision service rather than a similarity-search operation.

For Google Cloud Vision, the two relevant feature names are TEXT_DETECTION and DOCUMENT_TEXT_DETECTION. Microsoft’s Azure AI Vision Read API provides a comparable managed OCR workflow. The best choice depends less on a universal accuracy ranking and more on your input source, output structure, existing cloud environment, regional requirements, and the results on your own representative images. The provider documentation cited here does not establish a directly comparable accuracy percentage.

Choose between Google Cloud Vision and Azure AI Vision Read

Consideration Google Cloud Vision Azure AI Vision Read
Best fit in this workflow Use TEXT_DETECTION for text in ordinary images; use DOCUMENT_TEXT_DETECTION for dense documents and hierarchical layout. A managed OCR workflow for images and PDFs; a natural fit when the application already uses Azure identity, networking, monitoring, or storage.
Input described by the provider guidance Cloud Storage URI or web URL. Google cautions that an external host may deny access or throttle requests. Image or PDF; the quickstart demonstrates submitting an image URL.
Processing model images:annotate accepts an annotation request. Asynchronous batch annotation is available for offline workloads. Read processing is asynchronous: submit a request, then query the returned operation result.
Output detail Text and word-level bounding polygons; document mode exposes page, block, paragraph, word, and break structure. Extracted text is converted to a character stream; page selection and page ranges are supported.
Batch size, regions, quotas, and price Google documents asynchronous batches of up to 2,000 image files and global, US, or EU regional OCR endpoints. Quotas and current price are not stated here; check Google Cloud documentation for your project and region. Batch limit, regional options, quota, and current price are not stated here; check Microsoft documentation for the API version and region you plan to use.
Accuracy comparison No directly comparable accuracy percentage is established in the provider guidance discussed here. Test both services, if relevant, on the same representative inputs.

Choose Google when you want its documented distinction between ordinary image text and dense document structure, or need its stated asynchronous batch and regional options. Choose Azure when Azure is already part of your application’s infrastructure or when its image/PDF and page-range workflow suits your job. These are practical selection criteria, not a claim that either service is universally more accurate.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
Sale
ScanSnap iX2500 Wireless or USB High-Speed Document Scanner, Black
  • OUR MOST ADVANCED SCANSNAP. Large touchscreen, fast 45ppm double-sided scanning, 100-sheet document feeder, Wi-Fi and USB connectivity, automatic optimizations, and support for cloud services. Upgraded replacement for the discontinued iX1600
  • CUSTOMIZABLE. SHARABLE. Select personalized profiles from the touchscreen. Send to PC, Mac, mobile devices, and clouds. QUICK MENU lets you quickly scan-drag-drop to your favorite computer apps
  • STABLE WIRELESS OR USB CONNECTION. Built-in Wi-Fi 6 for the fastest and most secure scanning. Connect to smart devices or cloud services without a computer. USB-C connection also available
  • PHOTO AND DOCUMENT ORGANIZATION MADE EFFORTLESS. Easily manage, edit, and use scanned data from documents, receipts, photos, and business cards. Automatically optimize, name, and sort files
  • AVOIDS PAPER JAMS AND DAMAGE. Features a brake roller system to feed paper smoothly, a multi-feed sensor that detects pages stuck together, and skew detection to prevent paper damage and data loss

Prepare access and choose the right input

Set up Google Cloud Vision

  1. Create or select a Google Cloud project.
  2. Enable the Vision API for that project.
  3. Configure billing and credentials for the project. The request example below uses an OAuth access token; keep it out of source code and logs.
  4. Choose either a Cloud Storage URI, such as gs://BUCKET/path/image.jpg, or an externally reachable web URL.

A URL is convenient for a quick test, but it makes the OCR request depend on another host continuing to serve the file and permit Google’s request. The host can deny access or throttle requests. For production, use a controlled Cloud Storage object when you need dependable access to the image. Ensure the identity making the request can access the chosen input.

Set up Azure when that ecosystem fits

Azure AI Vision Read uses an Ocp-Apim-Subscription-Key in the documented quickstart workflow. It accepts an image or PDF, begins an asynchronous operation, and returns an operation reference to query for the result. Exact endpoint construction and API-version details depend on the Azure resource and API version you configure; use the current Microsoft quickstart for those values rather than copying a guessed endpoint.

Send a Google OCR request

The request body contains an image source and a feature. Select TEXT_DETECTION for general text in an image. Substitute DOCUMENT_TEXT_DETECTION when the input is a dense page and you need page and paragraph hierarchy.

Rank #2
CZUR Shine Ultra Smart Portable Document Scanner, Thin Book Scanner
  • Design and Speed: Work with Windows XP/7/8/10/11 AND macOS 10.13 or later. Not compatible with Android and iOS. Designed for A3&A4(11.69*16.53 & 8.27*11.75 inch) document, any objects smaller than A3 size can be scanned with Ultra-fast scanning speed, about 1 second per page. Perfect device to scan FLAT papers
  • USB Document Camera & Scanner: Work as both a document camera for remote teaching&learning compatible with ZOOM; Goole Meet and a document scanner to scan papers and convert/OCR files. OCR supports 180+ languages for text recognition. Please note that Thai, Hebrew, and Arabic are currently not supported. If you need the complete OCR language support list, please feel free to contact us for more details
  • Patented Flattening Curved Book Page Technology: Shine Ultra applies CZUR’s patented technology to flatten the curved surface after pixel transformation to flattening of the book page (Only suitable for thinner books, ET series is recommended for thicker books)
  • High Resolution & AI Tech: CMOS 13MP (4160*3120, A4≈340 AND A3≈245 DPI) camera. Smart Paging and Auto Cropping; Combine Sides; Stamp Mode; and Multiple Color Modes
  • Height Adjustable & Portable: 2-level height adjustable neck. 90 degree foldable and lightweight 4 lbs with foot pedal for convenient operation
{
  "requests": [{
    "image": {"source": {"imageUri": "gs://BUCKET/path/image.jpg"}},
    "features": [{"type": "TEXT_DETECTION"}]
  }]
}

Send the JSON to https://vision.googleapis.com/v1/images:annotate using OAuth bearer authentication. Here is a complete cURL example for a Cloud Storage image. Set GOOGLE_OAUTH_ACCESS_TOKEN to an access token authorized for the project, and replace the bucket and object path.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
export GOOGLE_OAUTH_ACCESS_TOKEN="YOUR_OAUTH_ACCESS_TOKEN"

cat > request.json <<'JSON'
{
  "requests": [{
    "image": {"source": {"imageUri": "gs://BUCKET/path/image.jpg"}},
    "features": [{"type": "TEXT_DETECTION"}]
  }]
}
JSON

curl -sS -X POST 
  -H "Authorization: Bearer $GOOGLE_OAUTH_ACCESS_TOKEN" 
  -H "Content-Type: application/json" 
  --data-binary @request.json 
  "https://vision.googleapis.com/v1/images:annotate" 
  -o response.json

For a web-hosted image, replace the value of imageUri with its URL. That does not upload the file from your machine: Google must be able to retrieve it from the host.

Run the request and parse the result in Python

This example uses requests and reads a Cloud Storage URI from an environment variable. It prints the detected full text and, when available, each word’s bounding polygon. Set the token and URI before running it.

Rank #3
Sale
Brother DS-640 Compact Mobile Document Scanner, (Model: DS640)
  • FAST SPEEDS - Scans color and black and white documents a blazing speed up to 16ppm (1). Color scanning won’t slow you down as the color scan speed is the same as the black and white scan speed.
  • ULTRA COMPACT – At less than 1 foot in length and only about 1. 5lbs in weight you can fit this device virtually anywhere (a bag, a purse, even a pocket).
  • READY WHENEVER YOU ARE – The DS-640 mobile scanner is powered via an included micro USB 3. 0 cable allowing you to use it even where there is no outlet available. Plug it into you PC or laptop and you are ready to scan.
  • WORKS YOUR WAY – Use the Brother free iPrint&Scan desktop app for scanning to multiple “Scan-to” destinations like PC, Network, cloud services, Email and OCR. (2) Supports Windows, Mac and Linux and TWAIN/WIA for PC/ICA for Mac/SANE drivers. (3)
  • OPTIMIZE IMAGES AND TEXT – Automatic color detection/adjustment, image rotation (PC only), bleed through prevention/background removal, text enhancement, color drop to enhance scans. Software suite includes document management and OCR software. (4)
import os
import requests

access_token = os.environ["GOOGLE_OAUTH_ACCESS_TOKEN"]
image_uri = os.environ["OCR_IMAGE_URI"]  # e.g. gs://BUCKET/path/image.jpg
feature = os.environ.get("OCR_FEATURE", "TEXT_DETECTION")

if feature not in {"TEXT_DETECTION", "DOCUMENT_TEXT_DETECTION"}:
    raise ValueError("OCR_FEATURE must be TEXT_DETECTION or DOCUMENT_TEXT_DETECTION")

payload = {
    "requests": [{
        "image": {"source": {"imageUri": image_uri}},
        "features": [{"type": feature}]
    }]
}

response = requests.post(
    "https://vision.googleapis.com/v1/images:annotate",
    headers={
        "Authorization": f"Bearer {access_token}",
        "Content-Type": "application/json",
    },
    json=payload,
    timeout=60,
)
response.raise_for_status()
data = response.json()

result = data["responses"][0]
if "error" in result:
    raise RuntimeError(result["error"])

# The first text annotation contains the complete detected string, when present.
annotations = result.get("textAnnotations", [])
if annotations:
    print("Full detected text:")
    print(annotations[0].get("description", ""))

# Other text annotations can provide word-level text and polygon vertices.
for annotation in annotations[1:]:
    print({
        "text": annotation.get("description", ""),
        "vertices": annotation.get("boundingPoly", {}).get("vertices", []),
    })

# Document mode also returns structured full-text annotation data when present.
full_text = result.get("fullTextAnnotation")
if full_text:
    print("Structured full text:")
    print(full_text.get("text", ""))
    for page in full_text.get("pages", []):
        for block in page.get("blocks", []):
            for paragraph in block.get("paragraphs", []):
                words = [
                    "".join(symbol.get("text", "") for symbol in word.get("symbols", []))
                    for word in paragraph.get("words", [])
                ]
                print("Paragraph words:", words)

Install the dependency with python -m pip install requests. Run it with GOOGLE_OAUTH_ACCESS_TOKEN and OCR_IMAGE_URI set in the process environment. To request document structure, set OCR_FEATURE=DOCUMENT_TEXT_DETECTION. The code checks HTTP errors and also checks for an error in the first per-image response; successful HTTP status alone does not mean the OCR operation returned usable text.

Understand text, coordinates, and document structure

For a simple extraction, start with the complete detected string rather than joining every lower-level annotation yourself. The word-level annotations are useful when you need to associate text with image locations—for example, to highlight recognized words in a viewer or connect a label to a region. Their bounding polygons describe coordinates in the image; preserve the original image dimensions if you later need to map those coordinates onto a resized display.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For dense documents, request DOCUMENT_TEXT_DETECTION and traverse the returned hierarchy. Page, block, paragraph, word, and break information gives an application more structure than one flat string. That structure is useful when the next step depends on reading order or page grouping. It is not a guarantee that a complex layout will be interpreted exactly as a human expects; inspect results on the forms and scans your application actually receives.

Rank #4
Sale
ScanSnap iX1300 Wireless or USB Double-Sided Color Document Scanner, Black
  • FITS SMALL SPACES AND STAYS OUT OF THE WAY. Innovative space-saving design to free up desk space, even when it's being used
  • SCAN DOCUMENTS, PHOTOS, CARDS, AND MORE. Handles most document types, including thick items and plastic cards. Exclusive QUICK MENU lets you quickly scan-drag-drop to your favorite computer apps
  • GREAT IMAGES EVERY TIME, NO EXPERIENCE REQUIRED. A single touch starts fast, up to 30ppm duplex scanning with automatic de-skew, color optimization, and blank page removal for outstanding results without driver setup
  • SCAN WHERE YOU WANT, WHEN YOU WANT. Connect with USB or Wi-Fi. Send to Mac, PC, mobile devices, and cloud services. Scan to Chromebook using the mobile app. Can be used without a computer
  • PHOTO AND DOCUMENT ORGANIZATION MADE EFFORTLESS. ScanSnap Home all-in-one software brings together all your favorite functions. Easily manage, edit, and use scanned data from documents, receipts, business cards, photos, and more

Build your parser to tolerate missing or empty annotations. An image may contain no legible text, and different response fields serve different purposes. Keep the raw JSON for debugging where your data-handling policy permits it, but avoid logging access tokens or sensitive image contents.

Use Azure Read for an asynchronous image or PDF job

  1. Submit the image or PDF to Azure AI Vision Read using the resource’s configured endpoint, API version, and subscription key.
  2. Capture the operation reference returned by the submission request.
  3. Query the operation result until processing completes, following the status and retry guidance for the API version you use.
  4. Read the extracted text and use supported page selection or page ranges when you do not need to process every page.

Because this workflow is asynchronous, do not treat the initial submission as the extracted-text response. Your application needs to retain the operation reference and handle the later result. The exact polling interval, operation URL format, API version, quota, and pricing should come from the current Azure documentation for your resource; those details are not specified here.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Scale offline work and account for regions

Google documents asynchronous batch annotation for offline processing, with support for up to 2,000 image files and response JSON written to Cloud Storage. This is the relevant path when a collection can be processed as a batch rather than one interactive request at a time. Design the downstream job to read the result files from storage and associate each result with its input image.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Best Value
Sale
Epson Workforce ES-400 II High-Speed Color Duplex Desktop Document Scanner
  • FAST DOCUMENT SCANNING — Document scanner with feeder allows you to speed through stacks with a 50-sheet Auto Document Feeder (ADF); Efficient office scanner to help you scan more productively
  • INTUITIVE, HIGH-SPEED SOFTWARE — Quickly scan with this desktop document scanner; Epson ScanSmart Software lets you easily preview scans, email files, upload to the cloud, and more; Plus, automatic file naming saves even more time
  • SEAMLESS INTEGRATION — Easily incorporate your data into most document management software with the included TWAIN driver; Office document scanner integrates seamlessly with business workflows
  • EASY SHARING — Duplex scanner allows you to scan straight to email or popular cloud storage2 services like Dropbox, Evernote, Google Drive, and OneDrive for simple storage and sharing
  • SIMPLE FILE MANAGEMENT — Scanner allows the creation of searchable PDFs with Optical Character Recognition (OCR) and convert scans to editable Word or Excel files effortlessly; Designed for home and office document scanning

Google also documents global, US, and EU regional OCR endpoints. If where processing happens matters to your organization, select the applicable endpoint and verify that the rest of your storage and application path meets your requirements. Choosing a regional OCR endpoint alone does not establish the location of every copy of the image or result in your wider system.

Improve reliability and control cost

  • Prefer controlled inputs for production. A third-party web URL can stop working, block retrieval, or throttle requests. Put files in storage you control when that dependency is unacceptable.
  • Separate submission from completion. Google’s annotation call returns a response for the request; Azure Read uses an asynchronous operation-and-query workflow. Implement the lifecycle appropriate to the chosen provider instead of assuming the two are interchangeable.
  • Keep the original alongside extracted data. Retaining a reference to the source image and the OCR result allows a later review of suspicious text or coordinates, subject to your retention and privacy requirements.
  • Test with representative material. Include the image types, document density, languages, image quality, and layouts your own users provide. Provider feature names and broad capability statements are not a substitute for measuring whether the output is useful in your workflow.
  • Check current quotas and prices before launch. They can depend on provider, region, resource, and usage. The available provider information here does not support a like-for-like price or quota table.

Troubleshoot common OCR failures

  • Access denied or image cannot be fetched: Check the OAuth token and project setup, then verify that the Cloud Storage object is reachable by the identity making the request. For a web URL, confirm that the host permits Google’s retrieval; a denied or throttled external request is a documented risk.
  • HTTP request succeeds but no text appears: Inspect the per-image JSON response for an error and check whether text annotations are absent or empty. Confirm that the input really contains visible text and that the response parser is reading the right feature’s output rather than assuming every image yields annotations.
  • Dense-page output lacks the structure you need: Use DOCUMENT_TEXT_DETECTION rather than TEXT_DETECTION, then traverse the page/block/paragraph/word hierarchy instead of relying only on a flat string.
  • Coordinates do not line up in your interface: Compare the returned vertices with the original image’s coordinate space. If your display resizes or crops the image, apply the corresponding coordinate transformation before drawing boxes.
  • Azure returns no final text in the first response: That workflow is asynchronous. Save the returned operation reference and query its result rather than treating the submission response as completion.
  • Remote-URL tests work inconsistently: The source host may deny or throttle requests. Move the image to storage under your control for a more dependable production input path.

Or skip the browser setup

ScreenshotNeo is a website screenshot API, not an OCR service: it captures a webpage image, which you can then send to Google Cloud Vision, Azure Read, or another OCR provider. If the source is a webpage and your goal is to obtain a screenshot as the OCR input, one GET request can capture it. See the ScreenshotNeo API documentation for the request options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

ScreenshotNeo accepts cookie or consent banners like a visitor and removes 60+ known consent platforms, newsletter popups, and chat widgets before capture; each of those steps can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers report the page verdict and whether the request was billed. Its MCP server offers take_screenshot, get_page_info, and capture_pdf for AI agents and MCP clients. The Free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. These are screenshot features and do not perform OCR themselves.

Sign up for ScreenshotNeo to get 1,000 screenshots a month free with no card.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Further implementation choices

Use the full text or detailed geometry?

Choose the complete text string when the downstream task is search, indexing, or a rough transcription. Request and retain word-level polygons when your interface must point back to image regions. Choose document mode when page and paragraph grouping is material. These output choices affect how you consume the result; they do not eliminate the need to validate important extracted values.

One image at a time or an offline collection?

Use a request/response flow for an interactive image or a small, immediate task. For a large offline collection, Google’s documented batch mode can process up to 2,000 image files and write response JSON to Cloud Storage. Azure Read’s described workflow is asynchronous as well, but the available facts do not state a comparable batch size. Avoid assuming matching limits across services.

Remote URL or controlled storage?

A remote URL is useful for a quick test or content already published for retrieval. Controlled Cloud Storage avoids relying on an unrelated web host’s availability and access policy. Whichever source you select, check its accessibility from the service side, not just from your own browser.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Read next

Recommended PC Tool
Recommended PC Tool
Outdated Drivers Are Slowing You DownFree scan - exact matches
PC Slower Than It Used to Be?Free scan - under a minute

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.