๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

ModelScope
@ModelScope2022
Driving innovations with open communities. ๐Ÿ’ฌ Join our Discord:
๊ฐ€์ž… April 2024
183 ํŒ”๋กœ์ž‰ ์ค‘    16K ํŒฌ
NetEase-Youdao releases Confucius4-R2T2, a true-streaming ASR model for live captions, simultaneous translation, and voice agents. ๐Ÿค– ๐Ÿ† R2T2 achieves SOTA latency and recognition quality among the evaluated open-source models, while remaining competitive with leading closed-source systems. โšก Configurable 80msโ€“2s chunks deliver 200โ€“600ms average latency with near-offline recognition accuracy. ๐Ÿ“ Append-only decoding commits stable text without revising earlier words, avoiding transcript flicker and giving downstream agents reliable input. ๐ŸŒ Optimized for Chinese and English, with multilingual recognition, hotword prompts, and contextual prompts. ๐Ÿงฐ The GitHub repo includes inference code, a minimal example, and a vLLM backend for both offline and real-time streaming ASR. ๐Ÿ“œ Code: Apache 2.0. Weights: NetEase Model Use License Agreement.
๋” ๋ณด๊ธฐ