๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

ModelScope
@ModelScope2022
Driving innovations with open communities. ๐Ÿ’ฌ Join our Discord:
๊ฐ€์ž… April 2024
183 ํŒ”๋กœ์ž‰ ์ค‘    16K ํŒฌ
Digital PDFs or warped phone photos, TeleOCR parses them with one lightweight 1.2B vision-language model. ๐Ÿ“œ Apache 2.0 License. ๐Ÿค– ๐Ÿ“ƒ ๐Ÿ† Scores 96.87 overall on OmniDocBench v1.6, the highest among the listed specialized VLMs, and ranks #1# in the ICDAR 2026 Sci-ImageMiner Challenge. ๐Ÿ“ท Handles digital, photographed, curved, and degraded documents directly, without a separate dewarping model. ๐Ÿง  Combines geometry-aware synthesis, consensus-generated labels, image-based self-verification, and progressive training from vision-language alignment to reinforcement learning. โšก Supports structured parsing of text, tables, formulas, layouts, and reading order, with synchronous or asynchronous vLLM inference.
๋” ๋ณด๊ธฐ