I spent 67 hours of model time to find out how much dumber 4 bit really makes Qwen3.8-27B.
FP8 vs NVFP4 vs AWQ INT4 vs GGUF Q4_K_M vs NInfer on my 4x RTX 3090. 4,800 tasks, 10,120 requests, 14.5M reasoning tokens, no token caps anywhere.
The results surprised me. Big thread, lets go ๐งต