註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Jon Durbin
@jon_durbin
Human. Backend dev
加入 December 2012
138 正在關注    7.2K 粉絲
Remember, the main point here is not even the decentralized training (even if it's what I'd consider a step change). The real magic is in the inference optimizations it enables. (sparse fp4 native compute for routed experts, fixed kv cache for most of the attention, sparse latent for the rest, tiny model weights, etc.) This is just from an unoptimized vllm branch on our model arch, not even really tuned yet. Maybe not an entirely fair comparison need a lot more models to compare against and some nuance and so on, but this is the whole point. Limitless tokens on cheap hardware.
顯示更多
Dashboard is a quick work in progress, but for visibility into a run here ya go! This is around $11/b tokens, insane actually. MFU also insane. The whole thing, pretty legendary, and inference... Using a few nodes from @lium_io also!
顯示更多
0
7
92
23
轉發到社區