Register and share your invite link to earn from video plays and referrals.

Shuo Yang
@Andy_ShuoYang
3rd year phd at Berkeley; Efficient ML System;
186 Following    5K Followers
FreeToken is fast. Comparing to Ollama, we have 3–4× faster decode, and 6–30× faster prefill How? We introduce bandwidth-adaptive CPU–GPU execution + semantic-aware caching across agent turns. More details in the technical report:
Show more
0
87
2.6K
224
Forward to community
Your gaming PC can now serve frontier models at interactive speed using official checkpoints without extreme quantization! Qwen3.6 35B → 8GB RTX 4060 laptop @ 39 tok/s DeepSeek-V4-Flash 284B → RTX 5090 desktop @ 22-25 tok/s GLM-5.2 753B → RTX PRO 6000 workstation @ 15 tok/s Run your claude code or codex now with frontier model for $0 Meet FreeToken 🧵
Show more
0
260
3.8K
509
Forward to community