So excited to introduce Muse Realtime Avatars which brings any character into a live conversation! This is a big milestone on a journey that started two years ago at Waveforms, and I’m so proud of what we’ve built. Since moving over to MSL, we've been rebuilding the realtime AI inference stack at Meta to enable both realtime voice and video conversations at scale.
From a technology POV, realtime multimodal inference is probably one of the most interesting and complex problems in AI Infra. It touches every part of the stack from model and hardware optimizations to LLM orchestration and the RTC stack. Every millisecond matters. Especially at scale. Realtime voice inference is cool, but compute-bound diffusion-based video models add a whole new dimension and bring an entirely new set of challenges. Our work makes it possible to serve these models at low latency and cost-effectively at Meta’s scale.
From an experience perspective, I find these interactions magical. Personal superintelligence can’t be truly personal unless you can express yourself in modalities beyond text. And so I’m super excited that Muse Realtime Avatars are coming to Muse! We'll have more to share soon, but for now check out this blogpost about our work :)