Digital PDFs or warped phone photos, TeleOCR parses them with one lightweight 1.2B vision-language model. ๐ Apache 2.0 License.
๐ค
๐
๐ Scores 96.87 overall on OmniDocBench v1.6, the highest among the listed specialized VLMs, and ranks #
1# in the ICDAR 2026 Sci-ImageMiner Challenge.
๐ท Handles digital, photographed, curved, and degraded documents directly, without a separate dewarping model.
๐ง Combines geometry-aware synthesis, consensus-generated labels, image-based self-verification, and progressive training from vision-language alignment to reinforcement learning.
โก Supports structured parsing of text, tables, formulas, layouts, and reading order, with synchronous or asynchronous vLLM inference.