๐ค MOSS-VL-Realtime is now open source on
@huggingface .
The 11B model family supports text, single and multiple images, single and multiple videos, and interleaved visual-text inputs in Chinese and English.
@MosiAI_Official
Highlights:
๐๏ธ Cross-Attention architecture separating visual encoding from language reasoning
๐งญ XRoPE for unified temporal-spatial positioning
๐งฉ Unified conversation templates for offline, streaming, and real-time interaction
๐ง 256K-token context window
๐ Apache-2.0 license
MOSS-VL-Realtime continues processing new frames while generating a response, allowing it to revise or interrupt that response as the scene evolvesโor remain silent when more evidence is needed.
Thank you
@sgl_project @lmsysorg for day-0 support! ๐
Huggingface:
Github:
Technical blog:
Join the community:
๐