Register and share your invite link to earn from video plays and referrals.

Syzygy Research
@syzygyeng
We compress the largest open weights LLMs to run blazing fast with addition-only inference. Backed by @speedrun @pearvc
11 Following    1.1K Followers
Today, we're introducing Mach-1 Additive, a 35 billion parameter model that can inference without ever multiplying by a weight. At 1.7 bits per weight, Mach-1 recovers 95% of the performance of the original full precision model, Qwen 3.6 35b, across 12 agentic and reasoning benchmarks, while being 10x smaller. At 7GB, Mach-1 comfortably fits on consumer laptops with speeds of up to 120 tokens per second, making local inference not just feasible but useful. Unlike algorithms like BitNet, our approach requires minimal retraining, under 15 GPU hours, making it scalable to massive LLMs. Over the coming weeks, we will be announcing and serving models of up to 3 trillion parameters compressed using our algorithm. For now, you can visit our website to play with Mach-1 directly in your browser, or download our desktop app. We couldn't be more excited to launch Mach-1. We're looking forward to an energy efficient future for AI, powered by scaled intelligence density.
Show more