🚀 ZDTaichu5.0-9B is now on ModelScope! 🤖
An on-device multimodal model from TaichuAI. At 9B parameters it runs on a single GPU and brings spatial reasoning, embodied AI and agentic tool use to edge deployment. Qwen3.5-9B backbone + C-RADIOv4-H vision encoder, 128K context, any-resolution image and video input.
🧭 Spatial reasoning: leads the compared 10B-scale open VLMs (Qwen3.5-9B, STEP3-VL-10B, gemma4-8B-E4B) and scores above Gemini 3 Pro, Grok 4 and GPT-5.2 on ViewSpatial, MMSI-Bench and MindCube-tiny
🛠️ Agent: highest among the compared open models on TAU2-Bench, Claw-Eval and IFEval
📄 First-tier results on documents, charts, OCR, visual math and video, with a ready-to-use vLLM branch and Docker image
🧠 Entropy-Gated Adaptive Recurrent Reasoning: extra latent refinement steps go only to the hard tokens
显示更多