🦢 TL;DR: nails the timing of listening, speaking and interrupting, but task completion still lags. Tencent unveils an omni interaction model unifying speech, video and agentic execution.
Title: Omni Interaction Agent Technical Report (Gander)
URL:
📌 Highlights
🧠 Cerebellum-Brain design: real-time control layer + Claude/Codex-powered reasoning layer
🎙 Streams audio/video in 1-second chunks, no external VAD needed
✅ 100% turn-taking accuracy vs GPT-Realtime's 96%
⚠️ Task completion (Pass
@1) only 0.400, below GPT-Realtime's 0.600
🔀 Feeding transcribed text straight to the Brain lifts Pass
@1 to 0.520
🌐 Fusing audio+video boosts Daily-Omni by +19.13pt over single-modality
📦 Weights, training code and eval harness planned for release
The honest takeaway: nailing conversational timing and nailing task accuracy are still two separate problems.
#
MultimodalAI# #
VoiceAgents#