Register and share your invite link to earn from video plays and referrals.

Ning
@totheagi
Building agent empire. ex google brain (tensorflow, TPU), Fudan CS.
196 Following    4.3K Followers
one week of pushing Kimi K3 on the 80x 5090s. where are we now: peak single stream: 95 tok/s aggregate decode: 1386 tok/s prefill: 13180 tok/s we still see a lot of potential. that's about 10% of what these cards can actually stream.
Show more
we got the full Kimi K3, 2.8T params, running on 80x RTX 5090s. 20 tok/s single stream, day one, untuned. Last week we took GLM-5.2 from 30 to 110 tok/s on this same fleet. This number will climb. A first for open weights: frontier intelligence served with zero HBM, the scarcest silicon in AI. Just GDDR7 gaming cards, plain ethernet, and the official MXFP4 weights, nothing requantized. The most powerful open model on Earth, on the most abundant GPUs on Earth. Any lab, startup, or university can now own it, probe it, fine-tune it, run agents on it. @Kimi_Moonshot
Show more
0
315
7.1K
619
Forward to community
now it's live🔥 GLM-5.2 on RTX 5090s: 80-110 tok/s single stream, up from ~30. 3x faster overnight. come feel it: