Register and share your invite link to earn from video plays and referrals.

Search results for AudioAI
AudioAI community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including AudioAI
The "answer only when spoken to" era of audio AI is over. Meet a model that keeps listening to sound, environment, and instructions — and acts on its own 🎧 Title: Audio Interaction Model URL: 🎧 Overview A unified, always-on streaming Large Audio Language Model (LALM). It runs a continuous perceive-decide-respond loop, listening to sound, environmental audio, and user instructions at once, then responding dynamically based on semantic understanding of the stream. ❓ Challenges Solved Most current LALMs run offline and handle isolated tasks (streaming ASR or voice chat separately). Real interaction needs an always-on model that listens in real time and reacts on the fly. 💡 Methodology & Proposed Approach The core is the SoundFlow framework operationalizing the perceive-decide-respond loop. ・Streaming-native data construction ・Comprehension-aware training ・Asynchronous low-latency inference for stable real-time interaction It trains on StreamAudio-2M (2.6M items) covering 7 fundamental abilities and 28 sub-tasks, plus Proactive-Sound-Bench to assess proactive intervention. 📊 Experimental Results / Use Cases ・Maintains competitive performance across 8 benchmarks ・Enables capabilities offline LALMs can't: real-time ASR, streaming audio instruction following, and proactive intervention Great for always-on voice assistants, real-time dialogue, and proactive audio assistance. #AudioAI# #LALM#
Show more
With cultural heritage as our oars and the spirit of our times as our vessel, we invite you to join us for a #DragonBoatFestival# audiovisual feast where ancient charm meets modern trends, and warmth blends with depth! Stay tuned on June 18th, #2026AdventuresonDragonBoatFestival#
Show more
With cultural heritage as our oars and the spirit of our times as our vessel, we invite you to join us for a #DragonBoatFestival# audiovisual feast where ancient charm meets modern trends, and warmth blends with depth! Stay tuned on June 18th, #2026AdventuresonDragonBoatFestival#
Show more
Today’s been an absolutely stacked day for AI news. In case you missed it: 1. GPT 5.6 Sol and Sol Ultra dropping tomorrow. Early reviews say it’s not as smart as Fable but very positive. 2. Grok 4.5 launches from Cursor / SpaceX that claims Opus 4.7 quality at 80tps and $2/M in $6/M out price. 3. Bytedance Seedream 5 Pro launches as an image ~#2# and nearly as good as GPT Image 2 at 4x cheaper cost: $0.045-$0.09/image, specializing in edits and infographics. Easily beats Meta’s Muse Image. 4. GPT-Live launches a full duplex non-turn based voice model which allows interactions while you’re talking seamlessly, a huge upgrade in audio AI for consumer Coming soon: 1. Seedance 2.5 Pro expected to launch soon (early July) to further extend Bytedance’s lead in SOTA video gen models. Will use Seedream 5 Pro as the frame generator. 2. Gemini 3.5 Pro launching soon (July 17). Google seems to have fallen quite behind frontier and people eagerly await to see what Google can put out here amidst a string of high profile departures. 3. GPT-6 rumored to launch soon (Polymarket spikes at August 14) with a new larger retrain, taking aim at Fable.
Show more
0
78
1.3K
95
Forward to community