๐ AuK is officially here. Nano banana๐ for audio
An open-source foundation model for unified speech generation and editing.
Natural-language instructions + reference audio. One interface.
Zero-shot TTS. Instruction-controlled generation. Content editing. Whisper-conversion. De-accent. Timbre/style/emotion edit. Speed/Pitch control. Enhancement, denoising, multi-speaker and music separation.
Also releasing AuK-Flash: 4-step inference. ~4.5ร faster under matched conditions.
Code, weights, and demo are live. Try it and share your feedback.
๐ค Paper & upvote:
โญ GitHub & star: