注册并分享邀请链接,可获得视频播放与邀请奖励。

Andi Marafioti
@andimarafioti
leading multimodal research @huggingface (prev @unity)
加入 April 2022
698 正在关注    8.5K 粉丝
Speech-to-speech no longer needs speech-to-text! Until now, our stack was VAD -> STT -> LLM -> TTS. Now it can send audio directly to multimodal LLMs: VAD → MLLM → TTS No STT. The model understands your voice. Now go build better voice agents!
显示更多
0
37
1.4K
149
转发到社区