A 27B model with 96K context built an actual game locally on a 16GB RTX 5070 Ti at 75 tps.
This wasn't a benchmark but an actual working villager simulation.
Setup đ
đŽ RTX 5070 Ti 16GB
đ§ Qwen3.8-27B
đĻ UD-Q3_K_XL GGUF
đ Fully GPU-offloaded
đĻ beellama.cpp + Kvarn optimizations
đŽ MTP n=2
đ§Ž Kvarn3 KV cache
đ 96,256 context
⥠up to 75 tok/s generation
đĨ up to 1,700 tok/s prefill
đĒ Windows
The model incrementally built a browser-based village simulation with:
đ housing
đĻī¸ weather + seasons
đ day/night cycles
đ hunger
đĒĩ resources
đ deaths
đļ obstacle avoidance
đ¨âđž autonomous villagers
đ¯ The user deliberately chose a Q3_K_XL model + heavily compressed KV cache so the entire 27B model, MTP and ~96K context could stay inside 16GB VRAM instead of spilling weights to CPU.
Their conclusion was not to fear Q3 models or aggressive KV quantization when the alternative is CPU offload and a massive speed hit. đđĨ
â ī¸ This also uses beellama.cpp/Kvarn, not stock llama.cpp.
đ Reddit /r/LocalLLaMA/comments/1w821fg/ninfer_vs_llamacpp_vs_vllm_quality_speed/