World models keep coming.
I think this is a real step toward AGI because the model is no longer generating static outputs. It’s continuously simulating and updating a world in real time from streaming audio, video, speech, and user interaction.
That’s much closer to intelligence interacting with reality instead of just predicting text.