Register and share your invite link to earn from video plays and referrals.

Tony Wu
@tonywu_71
Multimodal, RAG, Agents | ColPali co-first author | @centralesupelec ๐Ÿ‡ซ๐Ÿ‡ท x @Cambridge_Uni ๐Ÿ‡ฌ๐Ÿ‡ง | Core Researcher at @hcompany_ai ๐Ÿง‘๐Ÿปโ€๐Ÿ’ป
545 Following    1.5K Followers
๐Ÿ‘€ Meet NeoMME: a family of 260M and 800M Multimodal-Native Multilingual efficient Encoders One bidirectional Transformer processes text tokens and raw image patches, with no pretrained vision tower, text encoder, or decoder. (1/N ๐Ÿงต)
Show more
And STv6 supports ColPali-like models out-of-the-box, with beloved features like token pooling and similarity maps for interpretability!