登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

cv usk
@cv_usk
AI / Software Research Notes AI Agent, LLMOps, MLOps, Software Architecture 投稿は個人の意見です。
参加 May 2026
280 フォロー中    413 ファン
Are we actually evaluating world models — or just collecting scores we can't explain? That gap is exactly what this research tackles. HarnessEval-W: Agentifying the Evaluation of Visual Worlds ❓ What's broken about current world model evaluation? 💡 Existing benchmarks return scalar scores with no verifiable reasoning chains. There's no automated way to verify physics correctness, causality, or state persistence the way a human would — making it impossible to know why a model scored what it did. ❓ How does HarnessEval-W actually work? 💡 It uses a three-level agent hierarchy. Cases are routed to specialized skills, sub-agents answer specific sub-questions, and parent agents aggregate evidence into a final score — with every step recorded in an auditable evidence tree. Evaluation spans 8 settings across 3 categories (observation quality, transition correctness, world persistence), covering 330 cases and 18 models. ❓ How does it compare to existing benchmarks? 💡 Physical transition pairwise accuracy improves from WBench's 31.9% to 71.7%, and the draw rate collapses from 52.2% to just 1.8%. Human alignment is strong: Spearman ρ=0.93 for intentional transitions and ρ=0.87 for physical transitions, validated against 5,000 pairwise A/B human judgments. ❓ What does it reveal about today's models? 💡 Seedance 2.0 leads at 75.5, followed by Wan 2.7 (75.0) and Kling 3.0 (74.4). Fine-tuning analysis exposes sharp capability trade-offs: DreamX-World gains +4.8 on exploratory transitions but loses −11.9 on intentional ones. And render quality vs. physical observation correlates at r=−0.04 — nearly independent capabilities. #WorldModel# #VideoGeneration#
もっと見る