benchmarks of a 50% pruned Qwen3.6-35b-a3b and expert-specific quantization technique (made by me)
7.3gb model preforming => 51gb model, exiting to see where I can bring this technique to. I have some more things lined up too.
I need a DGX spark😭
AI Benchmarks with @Benchmark & @Vercel.
It's ① the most aptly named event in SF history and ② about one of the most important software categories of our generation.
The companies that benchmark models and guide the world towards skill, truth and efficiency.
Join us:
Beyond benchmarks, we’re seeing impressive real-world use cases. For example, in evaluations on real Chrome security bugs, 3.8 Flash Cyber produced 2.6x more correct patches for vulnerabilities than larger models.
Because of its advanced capabilities, we're giving Government agencies & cybersecurity partners access through the Fairwind program, read more:
Everyone benchmarks intelligence. Nobody benchmarks honesty.
Grok 4.6 has the lowest hallucination rate of any frontier model right now. GPT-5.6 Sol sits at 92%.
The smartest model in the room means nothing if it cancels your Stripe subscriptions during a migration.
LLM benchmarks are superfluous if your AI is neutered by safety-ism.
I can’t get over how different @grok + @bot feels to use.
Codex and Claude have been slowly boiling the frog. 🐸