“SenseNova-U1.5: Towards Native Unified Visual Intelligence”
Most multimodal models still use separate representations for seeing and generating images.
SenseNova-U1.5 instead uses one 8B encoder-free, VAE-free model to understand, reason about, generate, and edit images directly in pixel space, including native 4K generation.
It then trains specialist RL experts for aesthetics, text, infographics, and editing, and distills them back into one unified model.
The result is a stronger case that visual understanding and generation can share the same native representation rather than being separate systems.
Sensei, we've come a long way in the struggle against Decagrammaton.
Do you remember the first time you have taken the fight to Binah?
Now that you are up against all the prophets, which one made the most impression on you?!
#BlueArchive#