登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

alphaXiv
@askalphaxiv
High fidelity research
参加 November 2023
101 フォロー中    56.5K ファン
“SenseNova-U1.5: Towards Native Unified Visual Intelligence” Most multimodal models still use separate representations for seeing and generating images. SenseNova-U1.5 instead uses one 8B encoder-free, VAE-free model to understand, reason about, generate, and edit images directly in pixel space, including native 4K generation. It then trains specialist RL experts for aesthetics, text, infographics, and editing, and distills them back into one unified model. The result is a stronger case that visual understanding and generation can share the same native representation rather than being separate systems.
もっと見る