注册并分享邀请链接,可获得视频播放与邀请奖励。

Jon Durbin
@jon_durbin
Human. Backend dev
加入 December 2012
138 正在关注    7.2K 粉丝
Remember, the main point here is not even the decentralized training (even if it's what I'd consider a step change). The real magic is in the inference optimizations it enables. (sparse fp4 native compute for routed experts, fixed kv cache for most of the attention, sparse latent for the rest, tiny model weights, etc.) This is just from an unoptimized vllm branch on our model arch, not even really tuned yet. Maybe not an entirely fair comparison need a lot more models to compare against and some nuance and so on, but this is the whole point. Limitless tokens on cheap hardware.
显示更多
Dashboard is a quick work in progress, but for visibility into a run here ya go! This is around $11/b tokens, insane actually. MFU also insane. The whole thing, pretty legendary, and inference... Using a few nodes from @lium_io also!
显示更多
0
7
92
23
转发到社区