๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

OpenMOSS
@Open_MOSS
An open research lab building artificial general intelligence. Join Discord โšก๏ธ
๊ฐ€์ž… January 2025
33 ํŒ”๋กœ์ž‰ ์ค‘    536 ํŒฌ
๐Ÿค— MOSS-VL-Realtime is now open source on @huggingface . The 11B model family supports text, single and multiple images, single and multiple videos, and interleaved visual-text inputs in Chinese and English.@MosiAI_Official Highlights: ๐Ÿ—๏ธ Cross-Attention architecture separating visual encoding from language reasoning ๐Ÿงญ XRoPE for unified temporal-spatial positioning ๐Ÿงฉ Unified conversation templates for offline, streaming, and real-time interaction ๐Ÿง  256K-token context window ๐Ÿ“œ Apache-2.0 license MOSS-VL-Realtime continues processing new frames while generating a response, allowing it to revise or interrupt that response as the scene evolvesโ€”or remain silent when more evidence is needed. Thank you @sgl_project @lmsysorg for day-0 support! ๐Ÿš€ Huggingface: Github: Technical blog: Join the community: ๐Ÿ‘‡
๋” ๋ณด๊ธฐ