Register and share your invite link to earn from video plays and referrals.

Max Lv
@m0d8ye
Chip Architect (2013-Present) at NVIDIA | Bitcoin Enthusiast since 2011. Open source developer: Views are my own.
Joined June 2009
351 Following    20K Followers
古早的 bitnet 邪求又被拿来压缩 LLM 了,说不定用在 DS4 和 Kimi K3 上效果更好。
Today, we're introducing Mach-1 Additive, a 35 billion parameter model that can inference without ever multiplying by a weight. At 1.7 bits per weight, Mach-1 recovers 95% of the performance of the original full precision model, Qwen 3.6 35b, across 12 agentic and reasoning benchmarks, while being 10x smaller. At 7GB, Mach-1 comfortably fits on consumer laptops with speeds of up to 120 tokens per second, making local inference not just feasible but useful. Unlike algorithms like BitNet, our approach requires minimal retraining, under 15 GPU hours, making it scalable to massive LLMs. Over the coming weeks, we will be announcing and serving models of up to 3 trillion parameters compressed using our algorithm. For now, you can visit our website to play with Mach-1 directly in your browser, or download our desktop app. We couldn't be more excited to launch Mach-1. We're looking forward to an energy efficient future for AI, powered by scaled intelligence density.
Show more