■ 補足
Jev-Omniは開発者が公開したマルチモーダル統合モデル。動画内ではセキュリティ映像の不審行動検知(200ms)、音声メッセージの緊急度判定(17ms)、サポート対応の解決可否判定など、複数モダリティを跨いだ判断を1モデルで完結させる例を紹介している。HuggingFaceには「akhilaaa3/Jev-Omni」名義で公開。⚠️Typed-benchmarksでの比較やレイテンシ数値は開発者自身の主張で、第三者検証は確認できていない。
■ 出典
Introducing Jev-Omni, the first multimodal system one model. (Other OSS versions miss atleast a modality)
Supports all modalities: text, images, audio and video !
On Par with Jev on Typed-benchmarks.
Scaled -> 30k examples on 8xH200 (data mix matters a lot)
< 100ms on 1 H100
More work is coming, so follow along !
顯示更多