What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

Use two separate stages: capture pixels with PyAutoGUI, then pass the resulting Pillow image to Tesseract through pytesseract. PyAutoGUI can save a full screen or a rectangular region and can find visual templates, but it does not read words. pytesseract is only the Python interface; the Tesseract OCR engine must also be installed and configured on your computer.

This separation makes the workflow easier to debug: first inspect the screenshot, then inspect the recognized text or structured word data. Neither library guarantees accurate recognition for every screen, so validate the output against representative images.

What the Python workflow does

A typical script follows this sequence:

  1. Install PyAutoGUI, Pillow and the pytesseract Python package.
  2. Install the separate Tesseract executable and make sure Python can find it.
  3. Capture the whole display or a region with pyautogui.screenshot().
  4. Send that image directly to pytesseract.image_to_string() for plain text, or to image_to_data() for word-level records and coordinates.
  5. Review the original image beside the OCR result and add application-specific validation.

The screenshot object is a Pillow image, so no intermediate file is required. Saving a file can still be useful when diagnosing a failed capture or preserving an audit trail.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Install and configure the dependencies

Python packages

In the virtual environment used by your project, install the libraries without assuming a particular release number:

#1 Best Overall
Sale
Epson Workforce ES-50 Compact & Lightweight Mobile Document Scanner
  • PORTABLE SCANNER FOR USE ON-THE-GO — The fastest and lightest mobile single-sheet-fed compact document scanner in its class¹
  • QUICK DOCUMENT SCANNING ― This Epson ultra-fast scanner scans a single page as quickly as 5.5 seconds²; Windows and Mac compatible
  • VERSATILE PAPER HANDLING ― Portable scanner scans documents up to 8.5 x 72 in; Also easily digitizes receipts and ID cards to make accounting, bookkeeping, and organizing simpler
  • INTUITIVE, HIGH-SPEED SOFTWARE — Epson ScanSmart Software³ is a smart tool allowing you to easily scan, review, and save; Stay organized easily with the help of this Epson scanner
  • EASY SETUP — USB-powered connect to your computer for quick and simple scanning; No batteries or external power supply required to operate portable document scanner; Standard Connectivity: USB 2.0
python -m pip install pyautogui pillow pytesseract

PyAutoGUI’s screenshot support depends on Pillow. The package documentation also names scrot as a Linux screenshot dependency. Check the current PyAutoGUI documentation for the requirement on your distribution rather than assuming that every desktop image has it.

The Tesseract engine

pytesseract does not contain the OCR engine. Install Tesseract using the package manager or installer appropriate for your operating system, then verify that the tesseract executable is on your PATH. If it is installed in a non-standard directory, set the command explicitly in Python:

import pytesseract

pytesseract.pytesseract.tesseract_cmd = r"C:PathTotesseract.exe"

Use the path syntax for your own operating system. A missing executable usually appears as a TesseractNotFoundError when the first OCR call runs.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Desktop and display prerequisites

PyAutoGUI captures an interactive desktop. A locked session, an unavailable display, a remote session with no desktop, or a headless server can prevent a useful screenshot even when Python itself is working. Test capture in the same user session and display configuration that will run the automation. The documentation also notes limitations around multiple monitors; confirm the current support before designing a multi-display workflow.

Capture a complete screen or a region

Full-screen capture

import pyautogui

image = pyautogui.screenshot()
image.save("screen.png")

screenshot() returns a Pillow image. Supplying a filename saves it as well, but retaining the returned object lets you send it straight to OCR.

Capture only the area containing text

import pyautogui

left, top, width, height = 100, 180, 900, 500
image = pyautogui.screenshot(region=(left, top, width, height))
image.save("table-region.png")

The region tuple is ordered as left coordinate, top coordinate, width and height. A smaller region reduces unrelated pixels and is often easier to inspect, but coordinates are tied to the current display layout. If a window moves or the display scaling changes, fixed coordinates may no longer select the intended content.

Rank #2
Sale
Brother DS-640 Compact Mobile Document Scanner, (Model: DS640)
  • FAST SPEEDS - Scans color and black and white documents a blazing speed up to 16ppm (1). Color scanning won’t slow you down as the color scan speed is the same as the black and white scan speed.
  • ULTRA COMPACT – At less than 1 foot in length and only about 1. 5lbs in weight you can fit this device virtually anywhere (a bag, a purse, even a pocket).
  • READY WHENEVER YOU ARE – The DS-640 mobile scanner is powered via an included micro USB 3. 0 cable allowing you to use it even where there is no outlet available. Plug it into you PC or laptop and you are ready to scan.
  • WORKS YOUR WAY – Use the Brother free iPrint&Scan desktop app for scanning to multiple “Scan-to” destinations like PC, Network, cloud services, Email and OCR. (2) Supports Windows, Mac and Linux and TWAIN/WIA for PC/ICA for Mac/SANE drivers. (3)
  • OPTIMIZE IMAGES AND TEXT – Automatic color detection/adjustment, image rotation (PC only), bleed through prevention/background removal, text enhancement, color drop to enhance scans. Software suite includes document management and OCR software. (4)

Capture after the interface reaches a known state

PyAutoGUI can click and type, but timing is application-specific. Put the application in the required state before capture, then use a deliberate wait or a visual check. Do not assume that a screenshot taken immediately after a click contains content that is still loading.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Extract plain text with pytesseract

import pyautogui
import pytesseract

image = pyautogui.screenshot(region=(100, 180, 900, 500))
text = pytesseract.image_to_string(image)
print(text)

The image is handed directly from PyAutoGUI to image_to_string(). The return value is a string and commonly includes line breaks. Treat it as recognized text, not as a reliable representation of the page’s underlying data: characters can be confused, columns can be reordered, and decorative elements can be interpreted as letters.

Specify a language when it is installed

Tesseract can use language data installed with the engine. Pass the appropriate language code through the lang argument, for example lang="eng", only when that language data is available in your installation:

text = pytesseract.image_to_string(image, lang="eng")

Language availability is an installation detail, not something the Python wrapper can supply by itself.

Get structured OCR data with image_to_data

When downstream code needs to locate words, filter low-confidence results, or associate text with coordinates, use image_to_data() instead of a plain string.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import pyautogui
import pytesseract
from pytesseract import Output

image = pyautogui.screenshot(region=(100, 180, 900, 500))
data = pytesseract.image_to_data(image, output_type=Output.DICT)

records = []
for i, word in enumerate(data["text"]):
    word = word.strip()
    if not word:
        continue
    records.append({
        "text": word,
        "confidence": data["conf"][i],
        "left": data["left"][i],
        "top": data["top"][i],
        "width": data["width"][i],
        "height": data["height"][i],
    })

for record in records:
    print(record)

The coordinates are relative to the captured image. Add the region’s left and top values if you need coordinates in the full-screen coordinate system. Confidence values are useful for triage, not proof: set a threshold only after checking representative screenshots from your application.

Rank #3
Sale
Epson Workforce ES-400 II High-Speed Color Duplex Desktop Document Scanner
  • FAST DOCUMENT SCANNING — Document scanner with feeder allows you to speed through stacks with a 50-sheet Auto Document Feeder (ADF); Efficient office scanner to help you scan more productively
  • INTUITIVE, HIGH-SPEED SOFTWARE — Quickly scan with this desktop document scanner; Epson ScanSmart Software lets you easily preview scans, email files, upload to the cloud, and more; Plus, automatic file naming saves even more time
  • SEAMLESS INTEGRATION — Easily incorporate your data into most document management software with the included TWAIN driver; Office document scanner integrates seamlessly with business workflows
  • EASY SHARING — Duplex scanner allows you to scan straight to email or popular cloud storage2 services like Dropbox, Evernote, Google Drive, and OneDrive for simple storage and sharing
  • SIMPLE FILE MANAGEMENT — Scanner allows the creation of searchable PDFs with Optical Character Recognition (OCR) and convert scans to editable Word or Excel files effortlessly; Designed for home and office document scanning

Turn word records into application data

For a stable layout, group words by nearby vertical coordinates, then sort each line by its horizontal coordinate. For changing layouts, prefer anchors and validation rules such as expected labels, numeric formats or date patterns. Keep the original image whenever an OCR result drives a consequential action.

PyAutoGUI image matching is not OCR

PyAutoGUI’s image-location helpers search for a visual template. They can answer questions such as “where does this button image appear?” They do not convert the letters inside that image into text. Its FAQ answers the question “Does PyAutoGUI do OCR?” with: “No, but this is a feature that’s on the roadmap.”

Use template matching when the target has a known appearance and OCR when the task is to read changing words. The confidence option for image matching requires OpenCV. Installing OpenCV for that option does not turn PyAutoGUI into an OCR engine; you still need Tesseract and pytesseract for text extraction.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Need Use Output
Capture the display pyautogui.screenshot() Pillow image
Find a known icon or button PyAutoGUI image-location helpers Screen coordinates or a match region
Read visible words pytesseract.image_to_string() Plain text string
Read words with positions pytesseract.image_to_data() Text, boxes and confidence fields

A complete reusable script

from pathlib import Path
import pyautogui
import pytesseract
from pytesseract import Output

REGION = (100, 180, 900, 500)
image = pyautogui.screenshot(region=REGION)
Path("capture.png").unlink(missing_ok=True)
image.save("capture.png")

plain_text = pytesseract.image_to_string(image, lang="eng")
data = pytesseract.image_to_data(image, lang="eng", output_type=Output.DICT)

words = []
for i, value in enumerate(data["text"]):
    value = value.strip()
    if value:
        words.append({
            "text": value,
            "confidence": data["conf"][i],
            "box": {
                "left": data["left"][i],
                "top": data["top"][i],
                "width": data["width"][i],
                "height": data["height"][i],
            },
        })

print(plain_text)
print(words)

This demonstrates the documented API handoff; it is not a guarantee that the engine is installed or that a particular screenshot will be recognized correctly. Remove the unlink line if an existing capture must be preserved.

Improve reliability without assuming accuracy

Capture consistently

  • Use a fixed window size and a named region where possible.
  • Wait for the target element to appear before capturing.
  • Save failed or suspicious captures so you can compare pixels with OCR output.
  • Keep display scaling, font size and theme stable for automated jobs.

Validate the result

Check required headings, expected field counts, numeric ranges and date formats. Flag low-confidence words for review rather than silently accepting them. Build a small representative image set from the screens your program actually encounters; no source here establishes a universal preprocessing recipe or accuracy percentage.

Handle documents correctly

Tesseract’s input guidance distinguishes ordinary images from documents. PDF OCR generally requires converting pages to images or using OCRmyPDF. A multi-image sequence is read only from its first image by Tesseract, so process pages individually or use a document-oriented workflow instead of passing a sequence and expecting every page to be recognized.

Rank #4
Canon Canoscan Lide 300 Scanner (PDF, AUTOSCAN, Copy, Send)
  • Scanner type: Document
  • Connectivity technology: USB
  • With Auto Scan Mode, the scanner automatically detects what you're scanning
  • Digitize documents and images
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshooting

“No module named pyautogui” or “No module named pytesseract”

Install the packages into the same Python environment that runs the script, then check the interpreter path with python -c "import sys; print(sys.executable)".

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

TesseractNotFoundError

The Python wrapper cannot find the external engine. Put the executable on PATH or assign pytesseract.pytesseract.tesseract_cmd to its full path.

Screenshot fails on Linux

Check that the process has access to the graphical display and that the screenshot dependency named by PyAutoGUI, scrot, is installed where required by your distribution.

The image is blank or from the wrong monitor

Confirm the session is unlocked, the intended display is active, and the coordinates match the current layout. Multi-monitor behavior should be checked against the current PyAutoGUI documentation before deployment.

Text is garbled

Inspect the saved image first. If the pixels are wrong, fix capture timing or the region. If the image is correct, test a tighter crop, a consistent scale and the correct installed language data. Use image_to_data() to identify which words are uncertain, then validate rather than claiming guaranteed recognition.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Template matching cannot use confidence

The confidence parameter requires OpenCV. Install and configure that optional dependency, or omit confidence and use the matching behavior supported by your installed PyAutoGUI version.

Best Value
Sale
ScanSnap iX2500 Wireless or USB High-Speed Document Scanner, Black
  • OUR MOST ADVANCED SCANSNAP. Large touchscreen, fast 45ppm double-sided scanning, 100-sheet document feeder, Wi-Fi and USB connectivity, automatic optimizations, and support for cloud services. Upgraded replacement for the discontinued iX1600
  • CUSTOMIZABLE. SHARABLE. Select personalized profiles from the touchscreen. Send to PC, Mac, mobile devices, and clouds. QUICK MENU lets you quickly scan-drag-drop to your favorite computer apps
  • STABLE WIRELESS OR USB CONNECTION. Built-in Wi-Fi 6 for the fastest and most secure scanning. Connect to smart devices or cloud services without a computer. USB-C connection also available
  • PHOTO AND DOCUMENT ORGANIZATION MADE EFFORTLESS. Easily manage, edit, and use scanned data from documents, receipts, photos, and business cards. Automatically optimize, name, and sort files
  • AVOIDS PAPER JAMS AND DAMAGE. Features a brake roller system to feed paper smoothly, a multi-feed sensor that detects pages stuck together, and skew detection to prevent paper damage and data loss

Performance, privacy and cost considerations

Capture only the area you need when full-screen images contain sensitive information or unnecessary pixels. OCR work scales with the image supplied to Tesseract, but the appropriate crop and timing depend on your interface; measure your own workload rather than relying on an old, environment-specific timing example. Avoid logging screenshots or extracted text when they contain credentials, personal information or payment details.

This local workflow has no per-image API charge, but it requires a usable desktop session and maintenance of Python, Pillow, PyAutoGUI, Tesseract and any optional OpenCV dependency. A browser-based service can be simpler when the source is a URL rather than a screen you control.

Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server. One GET request returns a PNG, JPEG, WebP or PDF. It accepts cookie and consent banners before capture and removes more than 60 known consent platforms, newsletter popups and chat widgets; each cleanup step can be disabled. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, and the response identifies the page verdict and billing status in headers.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For a URL, call the API directly (see the ScreenshotNeo documentation):

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const data = Buffer.from(await res.arrayBuffer());
require('fs').writeFileSync('shot.webp', data);

ScreenshotNeo also supports full-page captures with lazy images loaded, CSS-selector element capture, dark mode, 12 device presets or a custom viewport, retina scale, PDF paper and page controls, custom CSS and JavaScript, clicks, selector or network-idle waits, request and resource blocking, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, configurable caching, signed image links, asynchronous webhooks, bulk capture for up to 100 URLs per call, a usage API and an OpenAPI specification. Its parameter names are compatible with those used by other screenshot APIs, which can simplify migration.

The free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 screenshots; every feature is available on every plan. An MCP server exposes take_screenshot, get_page_info and capture_pdf to Claude, Cursor and other MCP clients, so an AI agent can perform the capture. Create a free ScreenshotNeo account to try it with 1,000 screenshots a month and no card.

Frequently Asked Questions

Can PyAutoGUI read text by itself?

No. It captures pixels and can locate visual templates; use pytesseract with the separate Tesseract engine for OCR.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Should I use image_to_string or image_to_data?

Use image_to_string for a plain text result. Choose image_to_data when you need words, bounding boxes or confidence fields.

Can Tesseract OCR a PDF and every image in a sequence automatically?

PDFs generally need page conversion or OCRmyPDF, and Tesseract’s input guidance says a multi-image sequence is read only from its first image. Process pages individually or use a document workflow.

Quick Recap

Bestseller No. 4
Canon Canoscan Lide 300 Scanner (PDF, AUTOSCAN, Copy, Send)
Canon Canoscan Lide 300 Scanner (PDF, AUTOSCAN, Copy, Send)
Scanner type: Document; Connectivity technology: USB; With Auto Scan Mode, the scanner automatically detects what you're scanning
$75.00

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.