Seeing InternVideo-Next explicitly characterize its architecture as "Encoder–Predictor–Decoder (EPD)" and defining the Predictor as a "latent world model" confirms a clear convergence in the field. 📉
From our early work on Context AutoEncoder (CAE, — where we decoupled the Encoder, Predictor ("Regressor" in the paper), and Decoder — to Meta's I-JEPA/V-JEPA shifting entirely to latent prediction, and now this.
It seems we are all validating the same core intuition:
Understanding ≠ Pixel Reconstruction. 💡
Decoupling representation learning from high-frequency detail generation to build World Models in latent space is evidently the path forward. Glad to see our early intuition resonating with the latest SOTA. 🫡
#
AI# #
ComputerVision# #
WorldModel# #
JEPA# #
CAE#