Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Scan for outdated or missing drivers - takes under a minute3Repair Windows errors before they cause bigger problemsOmniExtractBench is Datalab’s open benchmark for structured document extraction: it pairs PDFs with schemas and gold-standard JSON, then scores extraction results down to individual values. Datalab says the project is intended to make comparisons easier to inspect and more representative of varied document tasks. Its published vendor scores are company-reported results on this corpus—not an independently verified ranking or a forecast for your own documents.
What OmniExtractBench is—and what it is trying to change
Announced by Datalab on September 16, 2026, OmniExtractBench is both a document dataset and a software toolkit. The toolkit includes a scorer, prediction adapters and orchestration for running provider comparisons; the dataset contains PDFs, gold extraction JSON, schemas and a manifest. Datalab presents it as a way for customers to compare extraction vendors and for engineers to diagnose model failures.
Datalab’s rationale is that existing benchmarks can favor their creators, obscure how prediction harnesses work, provide scores that do not explain errors, or cover only a narrow range of documents. Those are Datalab’s claims about the benchmark landscape, not independently established findings. The project’s intended response is a shared corpus, inspectable scoring method and value-level explanations of results. Read Datalab’s announcement.
What documents the benchmark includes
Datalab reports 620 documents drawn from four sources. Its dataset card describes a 620-row train split; each manifest entry connects a document ID to a PDF, gold extraction JSON, inline schema and suite label. The suite labels make it possible to examine results by corpus component rather than only as one overall score.
PC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware match#1 Best Overall
- PORTABLE SCANNER FOR USE ON-THE-GO — The fastest and lightest mobile single-sheet-fed compact document scanner in its class¹
- QUICK DOCUMENT SCANNING ― This Epson ultra-fast scanner scans a single page as quickly as 5.5 seconds²; Windows and Mac compatible
- VERSATILE PAPER HANDLING ― Portable scanner scans documents up to 8.5 x 72 in; Also easily digitizes receipts and ID cards to make accounting, bookkeeping, and organizing simpler
- INTUITIVE, HIGH-SPEED SOFTWARE — Epson ScanSmart Software³ is a smart tool allowing you to easily scan, review, and save; Stay organized easily with the help of this Epson scanner
- EASY SETUP — USB-powered connect to your computer for quick and simple scanning; No batteries or external power supply required to operate portable document scanner; Standard Connectivity: USB 2.0
| Source suite | Documents | Coverage described by Datalab |
|---|---|---|
| ExtractBench | 329 | Forms, filings and decks |
| Internal documents | 202 | Dense scalar schemas and small documents |
| micro1 | 47 | Very large tables |
| LongArray | 42 | Large tables with repeated scalars |
| Total | 620 | Four source suites |
Datalab also calls out scans, dense tables, forms, research papers, credit agreements, resumes and filings as examples of challenging material. That is meaningful variety, but 620 documents do not establish that every industry, language, scan quality or document layout is represented. Whether the mix resembles a particular organization’s workload needs to be checked against its own document types and extraction schemas. The dataset card lists CC-BY-4.0 for the dataset. See the dataset card.
How the scorer evaluates an extraction
The repository describes a multi-step process: normalize documents, flatten predicted and gold JSON into addressed scalar values, normalize those values, then match ambiguous array entries using Hungarian matching. For nested arrays, matching is applied recursively. This approach is meant to avoid relying on array position when entries can be matched by content instead.
Rank #2
- FAST SPEEDS - Scans color and black and white documents a blazing speed up to 16ppm (1). Color scanning won’t slow you down as the color scan speed is the same as the black and white scan speed.
- ULTRA COMPACT – At less than 1 foot in length and only about 1. 5lbs in weight you can fit this device virtually anywhere (a bag, a purse, even a pocket).
- READY WHENEVER YOU ARE – The DS-640 mobile scanner is powered via an included micro USB 3. 0 cable allowing you to use it even where there is no outlet available. Plug it into you PC or laptop and you are ready to scan.
- WORKS YOUR WAY – Use the Brother free iPrint&Scan desktop app for scanning to multiple “Scan-to” destinations like PC, Network, cloud services, Email and OCR. (2) Supports Windows, Mac and Linux and TWAIN/WIA for PC/ICA for Mac/SANE drivers. (3)
- OPTIMIZE IMAGES AND TEXT – Automatic color detection/adjustment, image rotation (PC only), bleed through prevention/background removal, text enhancement, color drop to enhance scans. Software suite includes document management and OCR software. (4)
Datalab says every unique scalar address produces a Verdict: an atomic outcome intended to indicate whether a value matched, was misread, was missed, was invented or was fabricated. That makes it possible to investigate the individual extraction decisions behind a score rather than seeing only an aggregate. The project also says it handles null versus blank values consistently. For edge-case behavior, consult the repository’s metric specification; the headline description alone does not show how every task-specific business consequence is represented.
A deterministic scorer can make the same comparison reproducible under the same inputs and rules, but determinism is not the same as universal correctness. A scoring rule may treat each scalar consistently while a buyer still needs to decide which fields matter most, how costly a wrong value is, and whether the benchmark’s schemas resemble its production task. Explore the source code, scorer and metric documentation.
The Tool Desk
Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Rank #3
- FAST DOCUMENT SCANNING — Document scanner with feeder allows you to speed through stacks with a 50-sheet Auto Document Feeder (ADF); Efficient office scanner to help you scan more productively
- INTUITIVE, HIGH-SPEED SOFTWARE — Quickly scan with this desktop document scanner; Epson ScanSmart Software lets you easily preview scans, email files, upload to the cloud, and more; Plus, automatic file naming saves even more time
- SEAMLESS INTEGRATION — Easily incorporate your data into most document management software with the included TWAIN driver; Office document scanner integrates seamlessly with business workflows
- EASY SHARING — Duplex scanner allows you to scan straight to email or popular cloud storage2 services like Dropbox, Evernote, Google Drive, and OneDrive for simple storage and sharing
- SIMPLE FILE MANAGEMENT — Scanner allows the creation of searchable PDFs with Optical Character Recognition (OCR) and convert scans to editable Word or Excel files effortlessly; Designed for home and office document scanning
What Datalab’s vendor results say
The September 16, 2026 announcement reports accuracy, precision, recall and coverage for ten extraction configurations on the 620-document corpus. Datalab says each quality metric is averaged across documents the provider could process. Coverage is therefore essential context: a system that cannot process some documents is not being assessed on the same count as one with full coverage. Datalab attributes incomplete coverage to issues including output-limit exhaustion and rejection of schemas as too large.
| Configuration (as labeled by Datalab) | Accuracy | Precision | Recall | Coverage |
|---|---|---|---|---|
| Datalab, accurate | 93.85 | 95.32 | 95.11 | 620 |
| Datalab, balanced mode | 93.48 | 95.30 | 94.79 | 620 |
| Reducto, deep_extract v2 | 93.47 | 94.91 | 95.02 | 620 |
| Claude, opus 5 | 90.96 | 95.17 | 92.60 | 575 |
| Extend | 90.17 | 91.67 | 94.66 | 620 |
| Gemini, 3.7-flash | 86.79 | 94.48 | 88.81 | 526 |
| LlamaExtract | 84.93 | 86.57 | 93.13 | 616 |
| GPT, 5.6-sol | 83.85 | 95.11 | 84.99 | 615 |
| Mistral OCR | 76.78 | 85.93 | 79.27 | 574 |
| Azure CU with GPT-4.1-mini | 61.08 | 80.32 | 64.07 | 569 |
These are figures reported by Datalab for its 2026 benchmark settings, not an independent evaluation. The announcement does not establish independent replication, statistical uncertainty or performance on a buyer’s own documents. Model and configuration names above follow the announcement’s labels; they should not be read as timeless product versions. See the results and settings reported by Datalab.
Rank #4
- Scanner type: Document
- Connectivity technology: USB
- With Auto Scan Mode, the scanner automatically detects what you're scanning
- Digitize documents and images
How to use the results when evaluating a provider
The table is best treated as a shortlist of questions to investigate, not a purchasing decision. A higher aggregate score can be useful, but it does not tell you whether a system handles your layouts, required fields, failure tolerance or operating constraints. Compare systems on the same task settings and look beyond the overall mean.
- Check coverage: Ask which documents were not processed and why. A quality average over a smaller subset can conceal a practical inability to handle difficult inputs.
- Inspect error types: Use the address-level verdicts to distinguish omissions, misread values and invented values, then assess the consequences for your workflow.
- Review suite-level results: Compare performance on the corpus components closest to your documents, not only the combined score.
- Run a representative evaluation: Test with your own PDFs, schemas and acceptance criteria before choosing a provider; the public comparison does not predict your results.
Reuse, licensing and running the tools
The GitHub README describes installation options for the scorer alone or with optional harness and benchmark dependencies. It documents score and predict interfaces, plus orchestration for running providers on the dataset or a selected manifest. Its benchmark command requires provider API credentials and can incur costs.
Best Value
- OUR MOST ADVANCED SCANSNAP. Large touchscreen, fast 45ppm double-sided scanning, 100-sheet document feeder, Wi-Fi and USB connectivity, automatic optimizations, and support for cloud services. Upgraded replacement for the discontinued iX1600
- CUSTOMIZABLE. SHARABLE. Select personalized profiles from the touchscreen. Send to PC, Mac, mobile devices, and clouds. QUICK MENU lets you quickly scan-drag-drop to your favorite computer apps
- STABLE WIRELESS OR USB CONNECTION. Built-in Wi-Fi 6 for the fastest and most secure scanning. Connect to smart devices or cloud services without a computer. USB-C connection also available
- PHOTO AND DOCUMENT ORGANIZATION MADE EFFORTLESS. Easily manage, edit, and use scanned data from documents, receipts, photos, and business cards. Automatically optimize, name, and sort files
- AVOIDS PAPER JAMS AND DAMAGE. Features a brake roller system to feed paper smoothly, a multi-feed sensor that detects pages stuck together, and skew detection to prevent paper damage and data loss
The README lists Apache 2.0 for the code, while the dataset card lists CC-BY-4.0 for the dataset. These are distinct licenses: check the relevant current license files and dataset terms before reusing or redistributing either artifact. The dataset card describes joining user predictions to the manifest by doc_id.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




