just-in-time OCR is all the rage.
most pipelines parse every page before anyone asks a question. for an agent working through an ad-hoc data room, that's slow, expensive, and most of those pages never get read.
the better pattern is just-in-time OCR in two passes:
โ ๏ธ LiteParse (free, OSS, Rust, 50+ formats) does a fast layout-aware first pass: spatial text, bounding boxes, headings, tables, and a per-page complexity flag. a full data room in 32 seconds.
โ ๏ธ LlamaParse zooms in on only the pages that need it, by page number, and returns cell-level tables, bounding boxes, and confidence scores. the rest fills in the background.
pypdf and pdftotext can't do the first pass well. parsing everything up front can't do it cheaply. two passes gets you both.
full breakdown with numbers: