注册并分享邀请链接,可获得视频播放与邀请奖励。

alphaXiv
@askalphaxiv
High fidelity research
加入 November 2023
101 正在关注    56.9K 粉丝
“SenseNova-U1.5: Towards Native Unified Visual Intelligence” Most multimodal models still use separate representations for seeing and generating images. SenseNova-U1.5 instead uses one 8B encoder-free, VAE-free model to understand, reason about, generate, and edit images directly in pixel space, including native 4K generation. It then trains specialist RL experts for aesthetics, text, infographics, and editing, and distills them back into one unified model. The result is a stronger case that visual understanding and generation can share the same native representation rather than being separate systems.
显示更多
0
2
113
17
转发到社区