Per-page OCR routing for a PDF gateway
Inspect each PDF page before choosing text extraction, OCR, or review. A page-level gateway keeps quality visible and limits external processing.
A PDF is often a mixture of page types. Inspect each page, choose the least expansive extraction path, and make uncertain text visible for review before it drives a workflow.
Syntalith Team
A PDF can contain selectable text on one page, a scanned image on the next, and a damaged or misleading text layer after that. A file-level decision such as "OCR this PDF" wastes work and can make a reliable page less reliable.
Use page-level routing when mixed PDFs create external-processing or quality concerns. The gateway should make extraction quality visible to the downstream process.
Route each page by evidence
Inspect the page before choosing a method. Useful signals include the presence of a text layer, character count, repeated replacement characters, reading order, image coverage, table structure, and a sample of extracted text.
Keep the decision explicit. A page can go to local extraction, OCR, or review. The gateway should retain the source file identifier and page number throughout the process.
Three paths through a PDF
| Page state | Route | Required record |
|---|---|---|
| Selectable text passes quality checks | Local text extraction | method, text sample or quality signal, and source version |
| Image or weak text layer needs recognition | Approved OCR service or local OCR | provider or model, processing purpose, output, and review state |
| Extraction is ambiguous | Human review or a specialist parser | reason, page reference, reviewer, and corrected output |
Some pages need a table or layout parser rather than plain OCR. A gateway can route by page type while keeping the downstream schema stable.
Keep OCR in a controlled boundary
Decide which pages may leave the organisation. Apply access rules before upload, remove unnecessary metadata, and define how temporary files and provider copies are deleted. Keep the original PDF in the system that owns it.
Log the extraction method, source version, provider, and result status. A reviewer should be able to open the source page and compare the extracted text. Do not treat a confidence number as a substitute for visual review when the field drives a consequential action.
The GDPR text is the primary reference for personal-data processing. The organisation still needs its own purpose, access, retention, processor, and transfer decisions.
Measure quality and cost together
Build a reviewed page set from the document types the workflow receives. Include selectable text, scans, tables, rotated pages, low-quality images, headers and footers, and pages with multiple columns.
Measure character and field correctness, reading order, table structure, routing accuracy, review rate, latency, external processing volume, and cost per accepted page. Separate extraction quality from downstream classification or summarisation quality.
Re-test after changing a parser, OCR provider, threshold, document source, or page normalisation step. Keep the source version with the result.
Build the gateway around exceptions
The gateway should return a clear status for an unreadable page, missing source, processing failure, provider timeout, and review request. The downstream system should stop or route the document when a required page has no accepted extraction.
Add a human review queue with the original page, extracted output, reason for escalation, and correction control. Preserve the corrected text as a derived version instead of overwriting the original extraction.
Decide whether page routing earns its scope
Page routing is useful when PDFs mix page types, external OCR has a data or cost boundary, or a downstream process needs traceable extraction. A single trusted document type may only need one parser with a quality check.
For a document workflow, bring sample PDFs, data classes, required fields, and the review threshold to a process scan. The AI apps service can cover the gateway and downstream integration.
FAQ
Should every PDF page go through OCR? No. Inspect the page and choose the least expansive method that meets the quality requirement.
Can the gateway guarantee perfect text? No. It can expose the extraction path, retain the source, and route uncertainty to review.
What is the most important field to preserve? Keep the source file version, page locator, extraction method, and review status linked to every downstream field.
Sources
Free process scan
Start with a free process scan.
- A 30-minute call with the engineer who would lead the work.
- A review of the processes that cost you the most time and money.
- A written summary of what to automate first and the likely cost range.
The scan chooses one process to assess, and within 2 business days you receive a recommendation, including when a simpler route is the better fit.
€0
30 minutes · written takeaway within 2 business days
Times are shown in your own time zone. We work with clients across time zones.
Describe the process in the form