๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

cv usk
@cv_usk
AI / Software Research Notes AI Agent, LLMOps, MLOps, Software Architecture ๆŠ•็จฟใฏๅ€‹ไบบใฎๆ„่ฆ‹ใงใ™ใ€‚
๊ฐ€์ž… May 2026
279 ํŒ”๋กœ์ž‰ ์ค‘    408 ํŒฌ
๐Ÿฆข TL;DR: nails the timing of listening, speaking and interrupting, but task completion still lags. Tencent unveils an omni interaction model unifying speech, video and agentic execution. Title: Omni Interaction Agent Technical Report (Gander) URL: ๐Ÿ“Œ Highlights ๐Ÿง  Cerebellum-Brain design: real-time control layer + Claude/Codex-powered reasoning layer ๐ŸŽ™ Streams audio/video in 1-second chunks, no external VAD needed โœ… 100% turn-taking accuracy vs GPT-Realtime's 96% โš ๏ธ Task completion (Pass@1) only 0.400, below GPT-Realtime's 0.600 ๐Ÿ”€ Feeding transcribed text straight to the Brain lifts Pass@1 to 0.520 ๐ŸŒ Fusing audio+video boosts Daily-Omni by +19.13pt over single-modality ๐Ÿ“ฆ Weights, training code and eval harness planned for release The honest takeaway: nailing conversational timing and nailing task accuracy are still two separate problems. #MultimodalAI# #VoiceAgents#
๋” ๋ณด๊ธฐ