Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

Use Amazon Bedrock Knowledge Bases when PDFs must become a searchable corpus. Choose the default parser for selectable-text PDFs, Bedrock Data Automation (BDA) or a foundation-model parser when figures, tables, charts, images, or layout matter, and an OCR pipeline for scans. For a single document, a direct model request can be simpler than creating a vector store. The right choice depends on the document, whether the answer is one-off or recurring, and how much visual meaning must be preserved.

Choose the extraction path first

“PDF extraction” in Bedrock is not one API or one parser. AWS provides several paths with different capabilities and billing models.

Path Best fit What it handles Billing and limits
Knowledge Bases default parser Selectable text in a repeat-use corpus Extracts text for chunking, embeddings, and retrieval; does not interpret visual content in charts, figures, tables, or images AWS says parsing has no usage charge
Bedrock Data Automation Managed multimodal ingestion Extracts and represents figures, charts, tables, images, and text without extra extraction prompting Priced by pages or images processed; applies to every PDF in that data source
Foundation-model parser Complex documents requiring an adjustable extraction instruction Model-based multimodal parsing with a customizable default prompt Priced by input and output tokens; applies to every PDF in that data source
Textract plus Bedrock Scanned pages and OCR-oriented workflows Textract can extract printed text, handwriting, layout elements, and data; Bedrock interprets the result Check current regional Textract and Bedrock prices and use the correct synchronous or asynchronous operation
Direct model request One document or a small application-controlled job Ask a supported model to interpret submitted content, subject to that model’s document-input limits Token and model charges; support for PDF bytes and formats differs by model

Advanced parsing is a data-source setting. If you select BDA or a foundation-model parser, Bedrock uses it for every PDF in that source, including text-only files. Separate text-only and visually rich collections when that makes cost and operations easier to control.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Text-only PDFs: build a Knowledge Base

A Knowledge Base is the appropriate default when people or applications will ask many questions over the same documents. Bedrock ingests the source, parses documents, chunks content, creates embeddings, and writes vectors to a configured vector store. Later, Retrieve returns relevant chunks; RetrieveAndGenerate retrieves chunks and asks a model to produce a grounded answer with source attribution.

#1 Best Overall
Canon Canoscan Lide 300 Scanner (PDF, AUTOSCAN, Copy, Send)
  • Scanner type: Document
  • Connectivity technology: USB
  • With Auto Scan Mode, the scanner automatically detects what you're scanning
  • Digitize documents and images

Set up the corpus

  1. Place the PDFs in a supported unstructured data source, such as an Amazon S3 location.
  2. Create an IAM role that allows Bedrock to read the source, invoke the selected embedding and foundation models, and write to the chosen vector store. Scope permissions to the specific bucket, prefixes, knowledge base, and model resources.
  3. Create the Knowledge Base and select an embedding model and vector store.
  4. Choose the default parser for selectable text. If any PDFs require visual interpretation, decide whether to move them to a separate data source and use BDA or a foundation-model parser there.
  5. Configure chunking, then start ingestion or synchronization.

Sync after additions, edits, or deletions so the vector index reflects the source. Treat extracted fields as candidates, not unquestionable truth; verify important values against the original page.

Retrieve chunks in Python

import boto3

bedrock_agent = boto3.client("bedrock-agent-runtime", region_name="us-east-1")

response = bedrock_agent.retrieve(
    knowledgeBaseId="YOUR_KNOWLEDGE_BASE_ID",
    retrievalQuery={"text": "What is the contract termination date?"},
    retrievalConfiguration={
        "vectorSearchConfiguration": {"numberOfResults": 5}
    },
)

for result in response.get("retrievalResults", []):
    text = result.get("content", {}).get("text", "")
    score = result.get("score")
    location = result.get("location", {})
    print(f"score={score}n{text}n{location}n")

Use Retrieve when your application must control ranking, filtering, validation, or its own response format. Use RetrieveAndGenerate when Bedrock should compose the answer from retrieved material and return citations. Keep the returned source metadata with the answer so users can inspect the supporting page or chunk.

Visually rich PDFs: BDA or a foundation-model parser

Text extraction alone can lose the meaning of a chart, table structure, figure, or image. BDA is the managed option when you want multimodal extraction without writing an extraction prompt. A foundation-model parser is useful when you need to customize instructions—for example, asking for a table as normalized records or requesting a particular field schema.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Control cost at the data-source boundary

Both advanced parsers process every PDF in the selected data source. A folder containing 10,000 text-only PDFs and 20 diagram-heavy PDFs can therefore incur advanced-parser processing for all 10,020 files. Splitting sources by parsing need can prevent that mismatch, but it adds synchronization and query-routing work. Confirm current regional pricing with your page count, parser, model, and output volume before committing.

Rank #2
Sale
Epson Workforce ES-50 Compact & Lightweight Mobile Document Scanner
  • PORTABLE SCANNER FOR USE ON-THE-GO — The fastest and lightest mobile single-sheet-fed compact document scanner in its class¹
  • QUICK DOCUMENT SCANNING ― This Epson ultra-fast scanner scans a single page as quickly as 5.5 seconds²; Windows and Mac compatible
  • VERSATILE PAPER HANDLING ― Portable scanner scans documents up to 8.5 x 72 in; Also easily digitizes receipts and ID cards to make accounting, bookkeeping, and organizing simpler
  • INTUITIVE, HIGH-SPEED SOFTWARE — Epson ScanSmart Software³ is a smart tool allowing you to easily scan, review, and save; Stay organized easily with the help of this Epson scanner
  • EASY SETUP — USB-powered connect to your computer for quick and simple scanning; No batteries or external power supply required to operate portable document scanner; Standard Connectivity: USB 2.0

Inspect layout-sensitive answers

Ask for page references or source attribution, then compare critical totals, labels, and units with the rendered page. Low-resolution scans, overlapping text, handwriting, and complex tables can produce plausible but incorrect values.

One-off extraction with a direct model request

For one document or a small, application-controlled workload, creating a Knowledge Base and vector store may be unnecessary. Bedrock’s Converse API supplies a common message interface for supported models. However, the Converse reference does not establish that every model accepts PDF bytes or the same document formats. Check the selected model’s input support, regional availability, size limits, and required permissions before sending a PDF directly. If direct document input is unavailable, extract text or render page images first, then submit those representations.

Minimal Converse pattern

import boto3

client = boto3.client("bedrock-runtime", region_name="us-east-1")

response = client.converse(
    modelId="YOUR_SUPPORTED_MODEL_ID",
    messages=[
        {
            "role": "user",
            "content": [
                {"text": "Extract the invoice number, date, supplier, and total. Return JSON and mark uncertain fields."}
                # Add a document or image content block only when this model's
                # current documentation confirms support for your input format.
            ],
        }
    ],
)
print(response["output"]["message"]["content"])

The runtime caller needs permission to invoke the model (including the appropriate bedrock:InvokeModel access). Validate the response against a schema, preserve the source page, and require human review for compliance, financial, or legal decisions.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Scanned PDFs and OCR

A scanned PDF is usually an image of a page, not a text document. OCR or visual interpretation must happen before reliable text retrieval. AWS’s Bedrock/Textract hands-on tutorial demonstrates DetectDocumentText with a single-page JPG or PNG. It explicitly excludes the different asynchronous Textract workflow required for multi-page PDFs, so that tutorial is not a complete multi-page implementation.

Rank #3
Sale
Brother DS-640 Compact Mobile Document Scanner, (Model: DS640)
  • FAST SPEEDS - Scans color and black and white documents a blazing speed up to 16ppm (1). Color scanning won’t slow you down as the color scan speed is the same as the black and white scan speed.
  • ULTRA COMPACT – At less than 1 foot in length and only about 1. 5lbs in weight you can fit this device virtually anywhere (a bag, a purse, even a pocket).
  • READY WHENEVER YOU ARE – The DS-640 mobile scanner is powered via an included micro USB 3. 0 cable allowing you to use it even where there is no outlet available. Plug it into you PC or laptop and you are ready to scan.
  • WORKS YOUR WAY – Use the Brother free iPrint&Scan desktop app for scanning to multiple “Scan-to” destinations like PC, Network, cloud services, Email and OCR. (2) Supports Windows, Mac and Linux and TWAIN/WIA for PC/ICA for Mac/SANE drivers. (3)
  • OPTIMIZE IMAGES AND TEXT – Automatic color detection/adjustment, image rotation (PC only), bleed through prevention/background removal, text enhancement, color drop to enhance scans. Software suite includes document management and OCR software. (4)

For production scans, choose the current asynchronous Textract document-processing path, verify its input constraints and output format in your region, collect all page results, and then send the extracted text or structured blocks to Bedrock. Preserve page numbers and confidence information. Review skewed pages, stamps, handwriting, columns, and tables manually when the value matters.

Return the original or parsed document

When an interface needs to display or download the document behind a retrieved result, call GetDocumentContent with the Knowledge Base, data-source, and document identifiers. The response provides a MIME type and a pre-signed URL that expires after five minutes. The caller needs both bedrock:Retrieve and bedrock:GetDocumentContent. If ACL-based access control is enabled, pass the user’s identity context so document permissions are enforced.

Operational checklist

  • Document classification: selectable text, scanned pages, or visual layout?
  • Workload shape: one answer, occasional batches, or a repeatedly queried corpus?
  • Parser scope: will every PDF in this source need multimodal parsing?
  • Model constraints: is the chosen model available in the target region and able to accept the submitted format?
  • Evidence: will responses retain chunk metadata, page numbers, and links to source content?
  • Security: are S3, model, vector-store, and document-content permissions least-privilege?
  • Validation: are low-confidence OCR and critical extracted fields routed for review?
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Common failures and fixes

Answers ignore a chart or table

Cause: the default parser extracted text only. Fix: use BDA or a foundation-model parser for that data source, or process the relevant pages as images.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Parsing costs are higher than expected

Cause: an advanced parser runs on every PDF in its source. Fix: separate text-only and multimodal collections, reduce unnecessary re-syncs, and check current page- or token-based prices.

Rank #4
Sale
Epson Workforce ES-400 II High-Speed Color Duplex Desktop Document Scanner
  • FAST DOCUMENT SCANNING — Document scanner with feeder allows you to speed through stacks with a 50-sheet Auto Document Feeder (ADF); Efficient office scanner to help you scan more productively
  • INTUITIVE, HIGH-SPEED SOFTWARE — Quickly scan with this desktop document scanner; Epson ScanSmart Software lets you easily preview scans, email files, upload to the cloud, and more; Plus, automatic file naming saves even more time
  • SEAMLESS INTEGRATION — Easily incorporate your data into most document management software with the included TWAIN driver; Office document scanner integrates seamlessly with business workflows
  • EASY SHARING — Duplex scanner allows you to scan straight to email or popular cloud storage2 services like Dropbox, Evernote, Google Drive, and OneDrive for simple storage and sharing
  • SIMPLE FILE MANAGEMENT — Scanner allows the creation of searchable PDFs with Optical Character Recognition (OCR) and convert scans to editable Word or Excel files effortlessly; Designed for home and office document scanning

A PDF cannot be sent to Converse

Cause: the selected model or request format does not support PDF input, or the file exceeds a limit. Fix: verify the model documentation and region; extract text or render pages before invoking the model.

Scanned pages return empty retrieval results

Cause: the source contains images without an OCR stage. Fix: run the appropriate asynchronous multi-page Textract workflow, preserve page output, and ingest the resulting text.

A source link no longer works

Cause: GetDocumentContent URLs expire after five minutes. Fix: request a new URL when the user opens the document rather than storing the signed URL permanently.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Retrieved text is stale

Cause: the source changed without a synchronization or direct ingestion/deletion operation. Fix: sync the data source and monitor ingestion completion before serving the new version.

Best Value
Sale
ScanSnap iX1300 Wireless or USB Double-Sided Color Document Scanner, Black
  • FITS SMALL SPACES AND STAYS OUT OF THE WAY. Innovative space-saving design to free up desk space, even when it's being used
  • SCAN DOCUMENTS, PHOTOS, CARDS, AND MORE. Handles most document types, including thick items and plastic cards. Exclusive QUICK MENU lets you quickly scan-drag-drop to your favorite computer apps
  • GREAT IMAGES EVERY TIME, NO EXPERIENCE REQUIRED. A single touch starts fast, up to 30ppm duplex scanning with automatic de-skew, color optimization, and blank page removal for outstanding results without driver setup
  • SCAN WHERE YOU WANT, WHEN YOU WANT. Connect with USB or Wi-Fi. Send to Mac, PC, mobile devices, and cloud services. Scan to Chromebook using the mobile app. Can be used without a computer
  • PHOTO AND DOCUMENT ORGANIZATION MADE EFFORTLESS. ScanSnap Home all-in-one software brings together all your favorite functions. Easily manage, edit, and use scanned data from documents, receipts, business cards, photos, and more

Or skip the browser setup

If your workflow also needs clean screenshots of source pages, ScreenshotNeo provides a website screenshot API and MCP server. It accepts cookie and consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be disabled. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status.

One request returns PNG, JPEG, WebP, or PDF:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

See the ScreenshotNeo documentation for authentication and options. Its MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients. Every plan includes features such as full-page and selector capture, device and retina settings, custom CSS and JavaScript, request blocking, cookies and headers, geolocation, caching, signed links, asynchronous webhooks, bulk capture, and a usage API. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.

Frequently Asked Questions

Can Bedrock extract handwriting from a PDF?

Use an OCR-oriented Textract workflow for scanned handwriting, then pass the extracted material to Bedrock for interpretation. Validate low-confidence results against the page.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Should I store every PDF in one Knowledge Base?

Not necessarily. Separate sources when parser requirements, access controls, update schedules, or cost profiles differ; query multiple sources only when your application can route and merge results safely.

How long does a GetDocumentContent download link last?

AWS documents a five-minute lifetime for the pre-signed URL, so generate it close to the time the user opens the document.

Quick Recap

Bestseller No. 1
Canon Canoscan Lide 300 Scanner (PDF, AUTOSCAN, Copy, Send)
Canon Canoscan Lide 300 Scanner (PDF, AUTOSCAN, Copy, Send)
Scanner type: Document; Connectivity technology: USB; With Auto Scan Mode, the scanner automatically detects what you're scanning
$75.00

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.