Best models smallest to largest right now.
- Gemma-4-12B
- Qwen3.8-27B
- Laguna-S2.1
- Deepseek-V4-Flash
- Inkling-Small
- MiniMax-M3
- GLM-5.*
- Kimi-K2.7-Code
- Qwen3.8-Max
- Kimi-K3
Spoiled for choice, open weights community is much more exciting than the frontier.
You can clone your real voice and fine-tune run locally on CPU,
1.7B Qwen3-TTS-12Hz-1.7B-CustomVoice is insane.
- Voice Cloning & Voice Design models.
- 12Hz tokenizer, 97ms latency.
- high-quality preset voices across Chinese, English, Japanese, Korean & more
- Control emotion, tone, speed with plain English instructions
Not just another TTS this one gives you fine-grained control via natural language: “Speak like an excited teacher” or “angry disappointed tone”
Multilingual, streaming-ready, and surprisingly runnable on consumer hardware.