註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Junlin Han
@han_junlin
Multimodal pre-training and post-training. final-year PhD student at Meta FAIR @AIatMeta and Oxford @OxfordTVG
加入 July 2020
1.1K 正在關注    1.3K 粉絲
We live in a multimodal world. We see, talk, act, and dream. Yet most LLMs still start with language pretraining. Why not train them natively with multimodal I/O from scratch? Because it’s SUPER HARD, adding modalities triggers training instability, design complexity, and often modal competition So what’s the path forward? Introducing: Towards Physics of Multimodal Pretraining ( We unpack the underlying mechanics of multimodal pretraining across 4 aspects: Knowledge Flow, Modality Synergy, Early Unification, and Recipe.
顯示更多
0
30
1K
171
轉發到社區