Register and share your invite link to earn from video plays and referrals.

alphaXiv
@askalphaxiv
High fidelity research
Joined November 2023
101 Following    56.5K Followers
“SenseNova-U1.5: Towards Native Unified Visual Intelligence” Most multimodal models still use separate representations for seeing and generating images. SenseNova-U1.5 instead uses one 8B encoder-free, VAE-free model to understand, reason about, generate, and edit images directly in pixel space, including native 4K generation. It then trains specialist RL experts for aesthetics, text, infographics, and editing, and distills them back into one unified model. The result is a stronger case that visual understanding and generation can share the same native representation rather than being separate systems.
Show more