Can video generation models actually reason about the visual world — or do they just make things look plausible?
These sound similar but are fundamentally different. Placing objects correctly, respecting physical laws, and solving multi-step tasks with coherent intermediate states: this is what "Visual Grounded Intelligence" means. VGI-Bench quantifies how far current models are from achieving it.
Across 27 tasks and 810 instances, the strongest model — Seedance 2.0 — reaches only 51.0% overall. Three failure modes appear repeatedly: physical collapse (unrealistic deformations and object penetrations), rule violations (reaching plausible end states while ignoring constraints), and object inconsistency (losing temporal identity across frames). Open-source models lag far behind commercial ones, with the strongest open-source model at 19.1%.
What the denoising analysis reveals is even more telling. Self-correction during intermediate generation steps occurs in fewer than 1% of transitions. Wrong-to-wrong transitions at mid-denoising reach 23.1%, and later denoising steps primarily refine early hypotheses rather than fixing errors. Fine-tuning on 1M synthetic samples improved planning and spatial skills but degraded temporal and physical capabilities — a tradeoff that points to deeper architectural constraints. VGI-Bench makes clear that visual reasoning remains a genuine unsolved frontier for video generation.
Title: VGI-Bench: Probing Visual Intelligence in Video Generation Models
URL:
#
VideoGeneration# #
VisualReasoning#