Are we actually evaluating world models — or just collecting scores we can't explain? That gap is exactly what this research tackles.
HarnessEval-W: Agentifying the Evaluation of Visual Worlds
❓ What's broken about current world model evaluation?
💡 Existing benchmarks return scalar scores with no verifiable reasoning chains. There's no automated way to verify physics correctness, causality, or state persistence the way a human would — making it impossible to know why a model scored what it did.
❓ How does HarnessEval-W actually work?
💡 It uses a three-level agent hierarchy. Cases are routed to specialized skills, sub-agents answer specific sub-questions, and parent agents aggregate evidence into a final score — with every step recorded in an auditable evidence tree. Evaluation spans 8 settings across 3 categories (observation quality, transition correctness, world persistence), covering 330 cases and 18 models.
❓ How does it compare to existing benchmarks?
💡 Physical transition pairwise accuracy improves from WBench's 31.9% to 71.7%, and the draw rate collapses from 52.2% to just 1.8%. Human alignment is strong: Spearman ρ=0.93 for intentional transitions and ρ=0.87 for physical transitions, validated against 5,000 pairwise A/B human judgments.
❓ What does it reveal about today's models?
💡 Seedance 2.0 leads at 75.5, followed by Wan 2.7 (75.0) and Kling 3.0 (74.4). Fine-tuning analysis exposes sharp capability trade-offs: DreamX-World gains +4.8 on exploratory transitions but loses −11.9 on intentional ones. And render quality vs. physical observation correlates at r=−0.04 — nearly independent capabilities.
#
WorldModel# #
VideoGeneration#