FreeToken is fast. Comparing to Ollama, we have 3–4× faster decode, and 6–30× faster prefill
How? We introduce bandwidth-adaptive CPU–GPU execution + semantic-aware caching across agent turns.
More details in the technical report:
Your gaming PC can now serve frontier models at interactive speed using official checkpoints without extreme quantization!
Qwen3.6 35B → 8GB RTX 4060 laptop @ 39 tok/s
DeepSeek-V4-Flash 284B → RTX 5090 desktop @ 22-25 tok/s
GLM-5.2 753B → RTX PRO 6000 workstation @ 15 tok/s
Run your claude code or codex now with frontier model for $0
Meet FreeToken 🧵