State tracking is a core pillar of video understanding: it requires identifying entities and events, and mapping how their states evolve over time.
Frontier multimodal models are surprisingly bad at it, so we built a benchmark to measure it.
Meet VSTAT!