Unlimited-OCR is now on ModelScope! A 3.3B multilingual OCR model for one-shot parsing across single images, multi-page documents, and PDFs. License: MIT🚀
🤖
📄 Full-document parsing instead of cropped-region OCR
📏 32K output length for long OCR sequences
🧩 Base and gundam image modes for different document layouts
⚙️ Transformers inference + SGLang serving with OpenAI-compatible streaming requests
Built to push DeepSeek-OCR-style document parsing further.
Mistral Document AI (with OCR 4) and Mistral Medium 3.5 are now in Microsoft Foundry.
Extract structure from documents and power apps with reasoning, coding, and agent capabilities, all within a unified platform.
🎉 🎉 🎉 We're open-sourcing Chronicles-OCR, a visual perception benchmark evaluating VLLMs on ancient Chinese characters.
The dataset spans 3,000 years of evolution. It covers 7 historical scripts from Oracle Bone to Cursive, featuring 2,800 balanced images across highly diverse physical media.
We assess models on 4 core tasks:
• Character Spotting
• Fine-grained Recognition
• Ancient Text Parsing
• Script Classification
The evaluation reveals how visual distribution shifts affect model perception over time.
Explore the dataset and paper below. 👇
📄 Paper:
🔗 GitHub: