Introducing Jev-Omni, the first multimodal system one model. (Other OSS versions miss atleast a modality)
Supports all modalities: text, images, audio and video !
On Par with Jev on Typed-benchmarks.
Scaled -> 30k examples on 8xH200 (data mix matters a lot)
< 100ms on 1 H100
More work is coming, so follow along !