Here's our take on benchmarks for voice: they're hard to get right, but critical if you're building humanlike AI. Our live benchmarks page is up now: pass rate on real agent scenarios, latency including the worst-case (p99) moments, side-by-side conversation comparison between models over time. Feedback welcome!