Are we actually evaluating world models โ or just collecting scores we can't explain? That gap is exactly what this research tackles.
HarnessEval-W: Agentifying the Evaluation of Visual Worlds
โ What's broken about current world model evaluation?
๐ก Existing benchmarks return scalar scores with no verifiable reasoning chains. There's no automated way to verify physics correctness, causality, or state persistence the way a human would โ making it impossible to know why a model scored what it did.
โ How does HarnessEval-W actually work?
๐ก It uses a three-level agent hierarchy. Cases are routed to specialized skills, sub-agents answer specific sub-questions, and parent agents aggregate evidence into a final score โ with every step recorded in an auditable evidence tree. Evaluation spans 8 settings across 3 categories (observation quality, transition correctness, world persistence), covering 330 cases and 18 models.
โ How does it compare to existing benchmarks?
๐ก Physical transition pairwise accuracy improves from WBench's 31.9% to 71.7%, and the draw rate collapses from 52.2% to just 1.8%. Human alignment is strong: Spearman ฯ=0.93 for intentional transitions and ฯ=0.87 for physical transitions, validated against 5,000 pairwise A/B human judgments.
โ What does it reveal about today's models?
๐ก Seedance 2.0 leads at 75.5, followed by Wan 2.7 (75.0) and Kling 3.0 (74.4). Fine-tuning analysis exposes sharp capability trade-offs: DreamX-World gains +4.8 on exploratory transitions but loses โ11.9 on intentional ones. And render quality vs. physical observation correlates at r=โ0.04 โ nearly independent capabilities.
#
WorldModel# #
VideoGeneration#