DeepSeek-V4 can see now. 👀 DeepSeek-V4-Flash-Vision-Exp is the family’s first experimental multimodal model. MIT License.
🤖
🏆 Beats Opus-4.8 on Agents’ Last Exam and ZeroBench, and trails it by just 0.7 on Chartography.
🖼️ Understands images, screenshots, and charts inside tool-using agent workflows.
⚡ Vision added, text strength retained: 83.9 on Terminal-Bench 2.1 and 59.3 on DeepSWE, above Opus-4.8’s 58.0.
🧠 Built on V4-Flash with new visual modules. Weights, prompt encoding, and minimal inference code are open.