๐ฆข TL;DR: nails the timing of listening, speaking and interrupting, but task completion still lags. Tencent unveils an omni interaction model unifying speech, video and agentic execution.
Title: Omni Interaction Agent Technical Report (Gander)
URL:
๐ Highlights
๐ง Cerebellum-Brain design: real-time control layer + Claude/Codex-powered reasoning layer
๐ Streams audio/video in 1-second chunks, no external VAD needed
โ
100% turn-taking accuracy vs GPT-Realtime's 96%
โ ๏ธ Task completion (Pass
@1) only 0.400, below GPT-Realtime's 0.600
๐ Feeding transcribed text straight to the Brain lifts Pass
@1 to 0.520
๐ Fusing audio+video boosts Daily-Omni by +19.13pt over single-modality
๐ฆ Weights, training code and eval harness planned for release
The honest takeaway: nailing conversational timing and nailing task accuracy are still two separate problems.
#
MultimodalAI# #
VoiceAgents#