Demis Hassabis on the next 12 Months:
- Full multimodal convergence: Models like Gemini will seamlessly take in and output text, images, audio, and video, with cross-pollination that boosts reasoning + creativity.
- Breakthrough visual intelligence: Image models like Nano Banana Pro will produce highly accurate infographics and show near-human visual understanding.
- Language + video fusion: Video models integrated with LLMs unlock richer analysis, storytelling, and step-by-step visual reasoning.
- World models go mainstream like Genie 3
- Agents become reliable