Case Study
Universal OCR
In ProgressUpload almost any real-world document — image or PDF — and extract its text, with adaptive preprocessing and bounding-box-accurate results.
- Stack
- Python · FastAPI · RapidOCR · Tesseract · OpenCV · onnxruntime · Next.js · TypeScript · Tailwind CSS
- My Role
- TBD
- Live
- TBD
- Repo
- TBD

Overview
Universal OCR is the OCR engine MVP of a future commercial Document-AI SaaS: a FastAPI backend running a pluggable OCR engine abstraction (RapidOCR and Tesseract today, PaddleOCR or a cloud engine later with no application changes), paired with a Next.js frontend for upload, preview, and review.
Problem
Real-world documents — a photographed receipt, a skewed scan, a mixed native-text-and-scanned PDF, an ID card mixing two scripts on one page — don't arrive clean, and a fixed OCR pipeline applied blindly to every document produces inconsistent results across that range.
Solution
A pipeline that measures each page before deciding what to do with it: detect whether a PDF page already has machine-readable text and skip OCR for it; for image-only pages, adaptively preprocess (orientation, perspective, deskew, denoise, contrast, scale) based on measured image quality, not a fixed sequence; then run the page through whichever OCR engine actually covers its language, with every bounding box mapped back to the original document's own pixel space.
Key Features
- Accepts images (JPG, PNG, WEBP, TIFF, BMP, GIF, HEIC/HEIF) and PDFs (text, scanned, or mixed) — mixed PDFs are handled page-by-page, native text extracted directly and scanned pages routed through OCR
- Adaptive preprocessing chosen from measured image quality, not applied uniformly to every document
- RapidOCR by default (PP-OCR models via onnxruntime), with automatic fallback to Tesseract for languages RapidOCR's bundled models don't cover, and multi-language routing for documents that mix two scripts on one page — verified against a real Aadhaar card
- QR code and barcode detection as structured data alongside OCR, never mixed into the text stream — verified against a real, high-density Aadhaar QR code
- Structured output (document → pages → blocks → lines → words) with
bounding boxes and confidence scores in the original document's pixel
space, plus per-line
low_confidenceflags and document-level confidence stats - One page failing doesn't take down the rest of the document; bounded automatic retry with a different engine on low confidence; cooperative job cancellation
Architecture
A FastAPI backend exposes an async job API (POST /api/ocr, job status
polling, result fetch, cancel) backed by an OcrService/JobService
pair over an in-memory job store — designed so the eventual move to a
real queue (Redis/DB-backed) is a storage-layer swap with zero route or
frontend changes. The OcrEngine abstraction (RapidOcrEngine,
TesseractEngine, with a Paddle/cloud adapter to come) lets an
EngineRegistry pick the right engine per request from an explicit
choice or a language catalog. Every preprocessing step composes into a
single transform matrix, inverted once at the end, so all coordinates in
the final result are in the original page's pixel space regardless of
what rotation/deskew/scale correction the OCR model actually saw. The
Next.js frontend drives the whole upload → process → review flow, with
a box-overlay view reading directly off those stored coordinates.
Technology
Python, FastAPI, RapidOCR (onnxruntime), Tesseract via pytesseract,
OpenCV for preprocessing, pypdfium2 for PDF handling, zxing-cpp for
QR/barcode detection, Next.js (App Router), TypeScript, Tailwind CSS.
Development Approach
Built using an AI-assisted development workflow (AIDLC) — Claude Code, Kiro and Amazon Q accelerate implementation, while architecture, decisions and quality remain owned by the developer.
Challenges
Keeping bounding boxes accurate in the original document's coordinate
space, even after a page goes through rotation, perspective correction,
deskew, and scaling, meant composing every preprocessing transform into
one matrix and inverting it once at the end — rather than trying to
track coordinates through each step individually. Handwriting isn't
supported (/api/capabilities reports this honestly rather than
silently failing), and multi-column reading order isn't resolved yet —
bounding boxes are preserved specifically so a future layout-analysis
pass can fix that without re-running OCR.
Decisions & Trade-offs
Every request is treated as an async job, even a 200ms image, rather than a synchronous call — more upfront plumbing, in exchange for one code path that already satisfies status polling, gives the frontend's staged progress UI real data, and needs no route changes when a real queue replaces the in-memory job store. Authentication, billing, and multi-tenancy were deliberately left out of this phase so the OCR core stays decoupled from product concerns that don't exist yet.
Lessons Learned
TBD.