Under the hood:
Text → transformer prompt conditioning
Generation → text-conditioned diffusion transformer in compressed acoustic latent space
Output → semantic-acoustic autoencoder decodes a 44.1 kHz stereo waveform
Flow matching, teacher–student distillation, and post-training with human feedback reduce the generation path to a few high-quality steps.
顯示更多