Here is the first speed test of
@Alibaba_Qwen Qwen3.8-Flash-Next, running on my 3090 😍
@UnslothAI UD-IQ3_XXS: 82GB
Total memory: 88GB
24 GB VRAM + 64 GB RAM
llama.cpp (pr #
27742#)
- Total usable context at f16 kv: 158K
- Total VRAM usage 23GB
Tested at depth up to 128k context.
- Avg decode speed: 21 t/s
- Peak decode speed: 33 t/s
- Peak prefill speed: 168 t/s @ 32k
All results were run without MTP or DFLASH