A 27B model with 96K context built an actual game locally on a 16GB RTX 5070 Ti at 75 tps.
This wasn't a benchmark but an actual working villager simulation.
Setup ๐
๐ฎ RTX 5070 Ti 16GB
๐ง Qwen3.8-27B
๐ฆ UD-Q3_K_XL GGUF
๐ Fully GPU-offloaded
๐ฆ beellama.cpp + Kvarn optimizations
๐ฎ MTP n=2
๐งฎ Kvarn3 KV cache
๐ 96,256 context
โก up to 75 tok/s generation
๐ฅ up to 1,700 tok/s prefill
๐ช Windows
The model incrementally built a browser-based village simulation with:
๐ housing
๐ฆ๏ธ weather + seasons
๐ day/night cycles
๐ hunger
๐ชต resources
๐ deaths
๐ถ obstacle avoidance
๐จโ๐พ autonomous villagers
๐ฏ The user deliberately chose a Q3_K_XL model + heavily compressed KV cache so the entire 27B model, MTP and ~96K context could stay inside 16GB VRAM instead of spilling weights to CPU.
Their conclusion was not to fear Q3 models or aggressive KV quantization when the alternative is CPU offload and a massive speed hit. ๐๐ฅ
โ ๏ธ This also uses beellama.cpp/Kvarn, not stock llama.cpp.
๐ Reddit /r/LocalLLaMA/comments/1w821fg/ninfer_vs_llamacpp_vs_vllm_quality_speed/