注册并分享邀请链接,可获得视频播放与邀请奖励。

Shuo Yang
@Andy_ShuoYang
3rd year phd at Berkeley; Efficient ML System;
加入 February 2023
186 正在关注    5K 粉丝
FreeToken is fast. Comparing to Ollama, we have 3–4× faster decode, and 6–30× faster prefill How? We introduce bandwidth-adaptive CPU–GPU execution + semantic-aware caching across agent turns. More details in the technical report:
显示更多
0
87
2.6K
224
转发到社区