๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

Xiaokang Chen
@PKUCXK
Researcher in Multimodal Team @deepseek_ai Previously Bachelor & Ph.D at Peking University @PKU1898 Opinions are my own.
๊ฐ€์ž… June 2022
63 ํŒ”๋กœ์ž‰ ์ค‘    6.4K ํŒฌ
All about multimodal. ๐Ÿง
Demis Hassabis on the next 12 Months: - Full multimodal convergence: Models like Gemini will seamlessly take in and output text, images, audio, and video, with cross-pollination that boosts reasoning + creativity. - Breakthrough visual intelligence: Image models like Nano Banana Pro will produce highly accurate infographics and show near-human visual understanding. - Language + video fusion: Video models integrated with LLMs unlock richer analysis, storytelling, and step-by-step visual reasoning. - World models go mainstream like Genie 3 - Agents become reliable
๋” ๋ณด๊ธฐ