what happens when you build a voice model that was never actually built around speech?
most TTS systems map text to phonemes, then phonemes to sound.
Bark skips that step entirely, text goes straight to audio tokens, GPT-style, the same way a language model predicts the next word.
output runs about 13 seconds at a time by default.
install with:
pip install git+