Qwen-AgentWorld just dropped two releases on ModelScope! An open 35B total / 3B active MoE world model with 256K context, plus a 7-domain benchmark grounded in real environment observations. 🚀
🔗
Qwen-AgentWorld-35B-A3B
🌍 One model for 7 agent environments: MCP, Search, Terminal, SWE, Web, OS, and Android
🧪 47.73 → 56.39 on AgentWorldBench, surpassing Claude Sonnet 4.6 at 56.04
🧠 Three-stage training: CPT injects environment knowledge, SFT activates next-state prediction reasoning, and RL sharpens simulation fidelity
AgentWorldBench
🛠️ Covers 7 domains with 2,170 samples and 22.8 average turns
🔎 Scores predictions on format, factuality, consistency, realism, and quality
顯示更多