Audex from Nvidia might be the first native audio model that isn't dumb as rocks.
Most audio models (audio understanding or speech / audio generation) just don't know anything about anything.
So it's cool to see a model that actually has smarts. I mean, when have you seen an audio model get a non-zero score on Terminal Bench?!
Again, it's as simple as an interleaved post training pipeline (nemotron omni really proved this approach imo).
Super cool model and release!