🤗 MOSS-VL-Realtime is now open source on
@huggingface .
The 11B model family supports text, single and multiple images, single and multiple videos, and interleaved visual-text inputs in Chinese and English.
@MosiAI_Official
Highlights:
🏗️ Cross-Attention architecture separating visual encoding from language reasoning
🧭 XRoPE for unified temporal-spatial positioning
🧩 Unified conversation templates for offline, streaming, and real-time interaction
🧠 256K-token context window
📜 Apache-2.0 license
MOSS-VL-Realtime continues processing new frames while generating a response, allowing it to revise or interrupt that response as the scene evolves—or remain silent when more evidence is needed.
Thank you
@sgl_project @lmsysorg for day-0 support! 🚀
Huggingface:
Github:
Technical blog:
Join the community:
👇