benchmarks of a 50% pruned Qwen3.6-35b-a3b and expert-specific quantization technique (made by me)
7.3gb model preforming => 51gb model, exiting to see where I can bring this technique to. I have some more things lined up too.
I need a DGX spark😭
顯示更多