Wait... The old Nvidia V100 cards have ~900 GB/s memory bandwidth. Those cards are almost a decade old at this point, but seems like they could still give pretty give great inference speeds, no?
The old card is faster on Qwen3.8-27B generation.
A 9-year-old $650 eBay V100 just outperforming a 2025 Strix Halo on Qwen3.8-27B decode.
- 56 tokens/sec.
Price-to-performance is wild right now.