Case Study

Universal OCR

In Progress

Upload almost any real-world document — image or PDF — and extract its text, with adaptive preprocessing and bounding-box-accurate results.

Stack
Python · FastAPI · RapidOCR · Tesseract · OpenCV · onnxruntime · Next.js · TypeScript · Tailwind CSS
My Role
TBD
Live
TBD
Repo
TBD
Universal OCR's review screen: an uploaded invoice alongside its extracted text, with OCR bounding boxes overlaid on the source image

Overview

Universal OCR is the OCR engine MVP of a future commercial Document-AI SaaS: a FastAPI backend running a pluggable OCR engine abstraction (RapidOCR and Tesseract today, PaddleOCR or a cloud engine later with no application changes), paired with a Next.js frontend for upload, preview, and review.

Problem

Real-world documents — a photographed receipt, a skewed scan, a mixed native-text-and-scanned PDF, an ID card mixing two scripts on one page — don't arrive clean, and a fixed OCR pipeline applied blindly to every document produces inconsistent results across that range.

Solution

A pipeline that measures each page before deciding what to do with it: detect whether a PDF page already has machine-readable text and skip OCR for it; for image-only pages, adaptively preprocess (orientation, perspective, deskew, denoise, contrast, scale) based on measured image quality, not a fixed sequence; then run the page through whichever OCR engine actually covers its language, with every bounding box mapped back to the original document's own pixel space.

Key Features

  • Accepts images (JPG, PNG, WEBP, TIFF, BMP, GIF, HEIC/HEIF) and PDFs (text, scanned, or mixed) — mixed PDFs are handled page-by-page, native text extracted directly and scanned pages routed through OCR
  • Adaptive preprocessing chosen from measured image quality, not applied uniformly to every document
  • RapidOCR by default (PP-OCR models via onnxruntime), with automatic fallback to Tesseract for languages RapidOCR's bundled models don't cover, and multi-language routing for documents that mix two scripts on one page — verified against a real Aadhaar card
  • QR code and barcode detection as structured data alongside OCR, never mixed into the text stream — verified against a real, high-density Aadhaar QR code
  • Structured output (document → pages → blocks → lines → words) with bounding boxes and confidence scores in the original document's pixel space, plus per-line low_confidence flags and document-level confidence stats
  • One page failing doesn't take down the rest of the document; bounded automatic retry with a different engine on low confidence; cooperative job cancellation

Architecture

A FastAPI backend exposes an async job API (POST /api/ocr, job status polling, result fetch, cancel) backed by an OcrService/JobService pair over an in-memory job store — designed so the eventual move to a real queue (Redis/DB-backed) is a storage-layer swap with zero route or frontend changes. The OcrEngine abstraction (RapidOcrEngine, TesseractEngine, with a Paddle/cloud adapter to come) lets an EngineRegistry pick the right engine per request from an explicit choice or a language catalog. Every preprocessing step composes into a single transform matrix, inverted once at the end, so all coordinates in the final result are in the original page's pixel space regardless of what rotation/deskew/scale correction the OCR model actually saw. The Next.js frontend drives the whole upload → process → review flow, with a box-overlay view reading directly off those stored coordinates.

Technology

Python, FastAPI, RapidOCR (onnxruntime), Tesseract via pytesseract, OpenCV for preprocessing, pypdfium2 for PDF handling, zxing-cpp for QR/barcode detection, Next.js (App Router), TypeScript, Tailwind CSS.

Development Approach

Built using an AI-assisted development workflow (AIDLC) — Claude Code, Kiro and Amazon Q accelerate implementation, while architecture, decisions and quality remain owned by the developer.

Challenges

Keeping bounding boxes accurate in the original document's coordinate space, even after a page goes through rotation, perspective correction, deskew, and scaling, meant composing every preprocessing transform into one matrix and inverting it once at the end — rather than trying to track coordinates through each step individually. Handwriting isn't supported (/api/capabilities reports this honestly rather than silently failing), and multi-column reading order isn't resolved yet — bounding boxes are preserved specifically so a future layout-analysis pass can fix that without re-running OCR.

Decisions & Trade-offs

Every request is treated as an async job, even a 200ms image, rather than a synchronous call — more upfront plumbing, in exchange for one code path that already satisfies status polling, gives the frontend's staged progress UI real data, and needs no route changes when a real queue replaces the in-memory job store. Authentication, billing, and multi-tenancy were deliberately left out of this phase so the OCR core stays decoupled from product concerns that don't exist yet.

Lessons Learned

TBD.