Speaker diarization is one of the hardest problems in voice AI and the best-performing models can be too slow and expensive to run at scale.
We combined quality-aware quantization with index-based clustering to get
@pyannoteAI 's models running 9.6x faster at 3x the throughput.
Here's how we did it: