๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

Sihyun Yu
@sihyun_yu
๊ฐ€์ž… July 2020
798 ํŒ”๋กœ์ž‰ ์ค‘    1.5K ํŒฌ
Can MLLMs actually track what's happening in a video? Introducing VSTAT ๐ŸŽฏ, our new benchmark for visual state tracking. The tasks are simple: count cups, read typed words, count page flips. Humans solve them easily. MLLMs don't. ๐Ÿงต [1/11]
๋” ๋ณด๊ธฐ