註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Xiaokang Chen
@PKUCXK
Researcher in Multimodal Team @deepseek_ai Previously Bachelor & Ph.D at Peking University @PKU1898 Opinions are my own.
加入 June 2022
63 正在關注    6.4K 粉絲
All about multimodal. 🧐
Demis Hassabis on the next 12 Months: - Full multimodal convergence: Models like Gemini will seamlessly take in and output text, images, audio, and video, with cross-pollination that boosts reasoning + creativity. - Breakthrough visual intelligence: Image models like Nano Banana Pro will produce highly accurate infographics and show near-human visual understanding. - Language + video fusion: Video models integrated with LLMs unlock richer analysis, storytelling, and step-by-step visual reasoning. - World models go mainstream like Genie 3 - Agents become reliable
顯示更多