註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Shuo Yang
@Andy_ShuoYang
3rd year phd at Berkeley; Efficient ML System;
加入 February 2023
186 正在關注    5K 粉絲
FreeToken is fast. Comparing to Ollama, we have 3–4× faster decode, and 6–30× faster prefill How? We introduce bandwidth-adaptive CPU–GPU execution + semantic-aware caching across agent turns. More details in the technical report:
顯示更多
0
87
2.6K
224
轉發到社區