Register and share your invite link to earn from video plays and referrals.

LightSeek Foundation
@lightseekorg
LightSeek is the creator of TokenSpeed, TorchSpec, and SMG — building next-generation AI infra systems
0 Following    2.8K Followers
TokenSpeed × Kimi K3, Part I: TP8 optimization on GB300. Now stable 3+ weeks in production across ~1,000 B300 and ~2,000 H200 GPUs. Thanks @NVIDIAAI for engineering support and GB300 NVL72 access, and community teams for their support and contributions. 👇
Show more
Releasing the @Kimi_Moonshot K3 Draft Collection — 3 draft models (EAGLE-3, DFlash2, DSpark) trained with TorchSpec and @vllm_project on @NVIDIAAI GB200. 🚀 We also shared the data recipes. More on the blog →
Show more
GLM 5.3 keeps the same architecture as 5.2 — no changes there, just a post-training weight update. But the jump in capability is huge. Truly impressive. TokenSpeed had day 0 support for GLM 5.3. Big thanks to the GLM team for open-sourcing such a strong model 👇
Show more
🎉 Congrats to @Zai_org on GLM-5.3-Flash — the first natively multimodal model in the GLM-5 series, and their first model combining sparse and linear attention. GLM-5.3-Flash introduces a new combination of linear attention for local dependencies and sparse attention for global context, with IndexPool keeping the indexing overhead low even at 1M-token context. This new architecture also brings a new set of inference challenges. TokenSpeed already has day-0 support for GLM-5.3-Flash, with the stack validated on both @NVIDIA and @AMD GPUs. 👇
Show more
TokenSpeed Day 0 Support for @Alibaba_Qwen 3.8 Flash Next. 🔹GDN + Qwen Sparse Attention hybrid architecture 🔹Gated residual connections 🔹N-gram embedding For N-gram embedding, TokenSpeed also supports FP8 precision. As a preview for Qwen 4, we will continue optimizing beyond Day 0.
Show more
We’re proud to be the Day 0 open-source inference engine partner for @Alibaba_Qwen 3.8. To serve this 2.4T-parameter model across multi-node @NVIDIAAI Blackwell inference, we optimized DP/EP scaling across nodes, delivering 30%+ faster performance than TP16, plus DSpark speculative decoding with single CUDA graph optimization👇
Show more
Kimi K3 is now supported in TokenSpeed. We worked closely with the @Kimi_Moonshot team to bring Day 0 support to both NVIDIA Blackwell ((G)B200/(G)B300) and AMD Instinct (MI350X/MI355X) in just one week. ✅ Prefix caching ✅ Speculative decoding ✅ Disaggregated serving ✅ CUDA Graph decode We also share how we optimized KDA, Gated MLA, Stable LatentMoE, AttnRes, and our unified Flat KV architecture. Blog ↓
Show more
Excited to see TokenSpeed Kernel featured by @PyTorch. Portable kernel APIs enable one runtime across multiple GPU architectures without performance regrets, while keeping runtime logic and backend kernels cleanly separated. Thanks to the PyTorch for highlighting our work.
Show more
Excited to see TOKENSPEED_MLA integrated into vLLM on Blackwell GPUs. Happy to see more DeepSeek-R1 / Kimi-K2.5 users benefit from the software optimizations and acceleration brought by TokenSpeed MLA. Looking forward to more optimizations and collaborations with the open-source community ahead.
Show more