WAMs are popular because of their promise of better generalization. Is that true? We started playing with Video-Action-Model (VAMs) and realized a gap: video model backbones can compositionally generalize but VAMs often do not.
We coin this the Video-Action-Generalization (VAG) gap and present a study on how to explain and improve it. More details:
🧵 below