DeepSeek-V4 can see now. ๐ DeepSeek-V4-Flash-Vision-Exp is the familyโs first experimental multimodal model. MIT License.
๐ค
๐ Beats Opus-4.8 on Agentsโ Last Exam and ZeroBench, and trails it by just 0.7 on Chartography.
๐ผ๏ธ Understands images, screenshots, and charts inside tool-using agent workflows.
โก Vision added, text strength retained: 83.9 on Terminal-Bench 2.1 and 59.3 on DeepSWE, above Opus-4.8โs 58.0.
๐ง Built on V4-Flash with new visual modules. Weights, prompt encoding, and minimal inference code are open.