Speech recognition gets much harder the moment people start talking like they actually do.
People switch languages mid-sentence, restart thoughts, use weak microphones and speak over traffic or room noise.
I’d test Hojo-ASR-Multi-V1 first on English and French code-switching, then add background noise and phone-recorded audio.
That should expose where the transcription starts to break.
Clean audio is easy. Real conversations aren’t.
People switch languages halfway through a sentence. They change their minds, speak over traffic noise, use bad microphones — and still expect the model to keep up.
So we want to test the messy stuff.
Send us a short, non-sensitive audio sample or tell us about a difficult voice scenario. We’ll test it with Hojo-ASR-Multi-V1 and share what works, what doesn’t, and where the model still needs improvement.
What should we test first?
Model: