Founder, Aevonix | 36x DGX Spark Cluster | Building local AI agents with A.E.V.A. & Protagine | Sharing experiments, code & lessons from my Sparks & 6000 Pros
MiMo V2.6 Pro RL on 8 DGX Sparks: 17.8 → 68.3 tokens/s with DFlash.
3.8× end-to-end speedup across four long structured-output tests, versus the same runtime without speculation.
We’ve released the patches, setup and raw results: