가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

alphaXiv
@askalphaxiv
High fidelity research
가입 November 2023
101 팔로잉 중    56.5K 팬
“SenseNova-U1.5: Towards Native Unified Visual Intelligence” Most multimodal models still use separate representations for seeing and generating images. SenseNova-U1.5 instead uses one 8B encoder-free, VAE-free model to understand, reason about, generate, and edit images directly in pixel space, including native 4K generation. It then trains specialist RL experts for aesthetics, text, infographics, and editing, and distills them back into one unified model. The result is a stronger case that visual understanding and generation can share the same native representation rather than being separate systems.
더 보기