Overnight benchmark run on GX10 for Poolside Laguna S 2.1 (Q4_K_M):
- 20.4 tok/s at 1K context
- 19.4 tok/s at 8K context
- 11.3 tok/s at 128K context
Local benchmark:
- 64% coding
- 60% debugging
- 70% repo edits
- 30% JSON
- 30% review
Optimized NVFP4 + vLLM/DFlash runs can produce higher throughput, especially with long generations and concurrent requests.