Generating an interactive world continuously for over an hour — on a single GPU. A new world model makes this possible.
Alaya-EVOKE: From Linear-Scaling Supervision to Endless World
Three key innovations that make EVOKE stand apart — a deep dive into each.
🗺 Highlight 1: External Geometric Memory for O(1) Context
Traditional world models blow up transformer context as generation grows longer, making sustained generation beyond a few minutes infeasible. EVOKE estimates monocular depth from generated frames, stores them as point clouds in a camera-pose-indexed "World State Bank," and retrieves geometry for the next chunk via pose-based lookup. No matter how long the session runs, each denoiser call receives exactly 1.5 seconds (9 frames) of input. "Extending a session increases only the number of recurrent calls, without increasing the context length of any individual call."
🎓 Highlight 2: Decoupling Supervision Horizon from Gradient Horizon
Long-term consistency requires supervision over long horizons — but backprop memory explodes. EVOKE solves this with a 14B Wan2.2 teacher that evaluates full trajectories jointly, while student gradients are computed per-chunk with detached history. A controlled experiment confirms the effect: the long-horizon student (30s supervision) stabilizes at 101% of opening brightness; the short-horizon student settles at 74% (Wilcoxon p=0.016). Long teacher supervision is statistically the key to temporal consistency.
⚡ Highlight 3: CFG-Free 3-Step Inference Tops VBench-2.0
No Classifier-Free Guidance, just 3 denoising steps over a coarse-to-fine latent pyramid — yet EVOKE scores 66.77 on VBench-2.0 (rank 1 of 10 systems), matching multi-step competitors. Generation speed: 2.11 seconds per 1.5-second chunk on a single H200.
Decoupling persistent state from denoiser context is the conceptual pivot that finally enables world models to run indefinitely without degrading.
#
WorldModel# #
VideoGenerationAI#