AI agents · OpenClaw · self-hosting · automation

Quick Answer

Best Way to Parse Scanned PDFs With Low Latency (2026)

Published:

The short answer

The fastest way to parse scanned PDFs in 2026 is to avoid OCR where you can and parallelise it where you can’t. Check every page for a usable text layer and pull that text directly. Split the remaining scanned pages and send them concurrently to a synchronous OCR endpoint, never a batch job. Then pick the engine by what you need back: plain text (AWS Textract, Google Document AI, Azure Read at $1.50 per 1,000 pages) or tables and layout (Mistral OCR 4 or Datalab Convert at $4 per 1,000 pages, Reducto at $10 for the worst scans).

Latency in document parsing is mostly architecture, not model speed. A 40-page scanned contract sent as one asynchronous job waits in a queue and comes back as one blob; the same contract split into pages and fanned out comes back in roughly the time of the slowest page.

Step 1: Skip OCR on pages that don’t need it

Many “scanned” PDFs are hybrids: a scanned signature page inside a digitally generated document, or scans that were already OCR’d by the copier. Run a text-layer check with PyMuPDF or pdfplumber first. If a page returns a sensible amount of text with normal character distribution, use it. This costs milliseconds and nothing per page. Only pages with no text layer, or a garbage one, go to OCR.

Step 2: Split, then fan out to a synchronous endpoint

The synchronous limits decide your architecture:

ServiceSynchronous limitAsynchronous limitPlain OCR price (per 1,000 pages)
AWS Textract1 page per PDF/TIFF, 10 MB3,000 pages, 500 MB$1.50 first 1M pages, $0.60 after
Google Document AI (Enterprise OCR)15 pages per request (30 in imageless mode), 40 MB500 pages per document, 1 GB$1.50 up to 5M pages, $0.60 after; first 1,000 free
Azure Document Intelligence (Read)Async analyze/poll; 15 analyze requests per second by default on S0, 2,000 pages per documentSame API$1.50
Mistral OCR 4Synchronous APIBatch API at half price$4 standard, $2 batch
Datalab (Chandra)Hosted API, 25 requests/min on Free, 400 on Team ($400/month)—$4 Convert fast/balanced, $10 accurate
Reducto r-1 ParseSync for low-latency calls, async with webhooksBatch 20% off$10

Two practical consequences:

  • Textract: send each page as its own DetectDocumentText call and run them concurrently up to your account’s TPS quota. Never send a multi-page PDF to the async job if a user is waiting.
  • Google: chunk into 15-page requests (30 if you can use imageless mode) and process chunks in parallel.

Render scans at 200–300 DPI before upload; higher resolution mostly adds upload time, lower resolution costs accuracy on small print.

Step 3: Pick the engine by output, not by benchmark

Plain text, lowest latency and cost: AWS Textract or Google Document AI. Both are $1.50 per 1,000 pages and accept work synchronously. Choose the cloud you already run in; the IAM, storage triggers and private networking matter more than any accuracy difference on clean scans. Textract reads handwriting in English only.

Tables, multi-column layouts, forms: Mistral OCR 4 or Datalab. Mistral OCR 4 (released June 23, 2026) returns markdown with bounding boxes, block types and per-word confidence for $4 per 1,000 pages. Mistral quotes the fintech Rogo reaching equivalent accuracy to agentic document parsers “at roughly 8x lower cost and 17x lower latency,” and the IP firm Anaqua measuring it about 4x faster per page than its incumbent. Datalab’s hosted Chandra charges $4 per 1,000 pages for Convert in fast or balanced mode and $10 for accurate mode; Datalab says speculative decoding cut its API’s p50 latency by 25% and p99 by 3x. Don’t use the Batch API on either if a person is waiting: batch exists to save money, not time.

Hardest scans (faded, rotated, handwritten, merged-cell tables): Reducto. Its r-1 Parse offers a synchronous mode for low-latency calls at $10 per 1,000 pages. Route only the pages your cheaper engine returned with low confidence, and you keep both latency and cost down.

Step 4: Stream results back

Return each page as soon as it finishes instead of waiting for the whole document. For a chat or review UI, the user can start reading page 1 while page 30 is still processing. Keep page number and bounding boxes with every block so a later structured-extraction step can cite the source region.

The pick

  • Interactive app, plain text, under 15 pages: Google Document AI online processing in one request.
  • Interactive app, longer documents: page-level fan-out to Textract or Google, plus a text-layer shortcut.
  • You need tables and layout fast: Mistral OCR 4 synchronous, or Datalab Convert fast mode.
  • Accuracy on ugly scans beats everything: Reducto sync, but only on the pages that need it.
  • Data cannot leave your network: self-host Chandra 2 or Mistral OCR 4 (enterprise self-hosting) on GPUs; Chandra 2 does 2 pages per second on an H100 at 96 concurrent requests, which is throughput, not single-page latency.

Measure p50 and p95 on 200 of your own documents before committing. Vendor speed claims are customer quotes and marketing benchmarks, not comparable tests.

Related: the full OCR and document extraction API ranking, the most accurate document parsing API, the best OCR for handwriting and complex tables and how to build an OCR pipeline for millions of pages.

Last verified: October 6, 2026. Limits and prices are from vendor documentation and pricing pages; no vendor publishes comparable latency benchmarks.

Sources