Digital PDFs or warped phone photos, TeleOCR parses them with one lightweight 1.2B vision-language model. 📜 Apache 2.0 License.
🤖
📃
🏆 Scores 96.87 overall on OmniDocBench v1.6, the highest among the listed specialized VLMs, and ranks #
1# in the ICDAR 2026 Sci-ImageMiner Challenge.
📷 Handles digital, photographed, curved, and degraded documents directly, without a separate dewarping model.
🧠 Combines geometry-aware synthesis, consensus-generated labels, image-based self-verification, and progressive training from vision-language alignment to reinforcement learning.
⚡ Supports structured parsing of text, tables, formulas, layouts, and reading order, with synchronous or asynchronous vLLM inference.