Qwen 3.8 27B at 56tps; on 9 year old GPU btw
Nvidia V100 32GB ~$650 on EBay right now!
Using Dflash 2; disabling the ECC adds some more speed too! Thinking and prose is a bit slower, but 56-63 tps in code gen!
MTP runs faster for prose vs DFlash2 but slower sustained code generation speed. MTP also runs much faster power limited than DFlash does.
Working on a repo so you can get up and going quickly.
Fun fact, the Nvidia v100 was $11,500 per card when they first launched. Price you pay for future proofing I guess; they’re still great cards. Pcie 3.0 and the older software/architecture are the only drawbacks, but also those aren’t as much of an issue as you’d think. Especially when you consider the price today!