๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

David Hendrickson
@TeksEdge
CEO & Founder | PhD | Startup Advisor | @Columbia | Author Generative Software Engineering | ๐Ÿ”” Follow for AI & Vibe Coding Tips ๐Ÿ‘‡
๊ฐ€์ž… July 2023
549 ํŒ”๋กœ์ž‰ ์ค‘    11.2K ํŒฌ
A 27B model with 96K context built an actual game locally on a 16GB RTX 5070 Ti at 75 tps. This wasn't a benchmark but an actual working villager simulation. Setup ๐Ÿ‘‡ ๐ŸŽฎ RTX 5070 Ti 16GB ๐Ÿง  Qwen3.8-27B ๐Ÿ“ฆ UD-Q3_K_XL GGUF ๐Ÿš€ Fully GPU-offloaded ๐Ÿฆ™ beellama.cpp + Kvarn optimizations ๐Ÿ”ฎ MTP n=2 ๐Ÿงฎ Kvarn3 KV cache ๐Ÿ“š 96,256 context โšก up to 75 tok/s generation ๐Ÿ“ฅ up to 1,700 tok/s prefill ๐ŸชŸ Windows The model incrementally built a browser-based village simulation with: ๐Ÿ  housing ๐ŸŒฆ๏ธ weather + seasons ๐ŸŒ™ day/night cycles ๐Ÿ– hunger ๐Ÿชต resources ๐Ÿ’€ deaths ๐Ÿšถ obstacle avoidance ๐Ÿ‘จโ€๐ŸŒพ autonomous villagers ๐ŸŽฏ The user deliberately chose a Q3_K_XL model + heavily compressed KV cache so the entire 27B model, MTP and ~96K context could stay inside 16GB VRAM instead of spilling weights to CPU. Their conclusion was not to fear Q3 models or aggressive KV quantization when the alternative is CPU offload and a massive speed hit. ๐Ÿ˜๐Ÿ”ฅ โš ๏ธ This also uses beellama.cpp/Kvarn, not stock llama.cpp. ๐Ÿ”— Reddit /r/LocalLLaMA/comments/1w821fg/ninfer_vs_llamacpp_vs_vllm_quality_speed/
๋” ๋ณด๊ธฐ