Qwen-AgentWorld just dropped two releases on ModelScope! An open 35B total / 3B active MoE world model with 256K context, plus a 7-domain benchmark grounded in real environment observations. ๐
๐
Qwen-AgentWorld-35B-A3B
๐ One model for 7 agent environments: MCP, Search, Terminal, SWE, Web, OS, and Android
๐งช 47.73 โ 56.39 on AgentWorldBench, surpassing Claude Sonnet 4.6 at 56.04
๐ง Three-stage training: CPT injects environment knowledge, SFT activates next-state prediction reasoning, and RL sharpens simulation fidelity
AgentWorldBench
๐ ๏ธ Covers 7 domains with 2,170 samples and 22.8 average turns
๐ Scores predictions on format, factuality, consistency, realism, and quality