Register and share your invite link to earn from video plays and referrals.

David Hendrickson
@TeksEdge
CEO & Founder | PhD | Startup Advisor | @Columbia | Author Generative Software Engineering | 🔔 Follow for AI & Vibe Coding Tips 👇
Joined July 2023
549 Following    11.2K Followers
A 27B model with 96K context built an actual game locally on a 16GB RTX 5070 Ti at 75 tps. This wasn't a benchmark but an actual working villager simulation. Setup 👇 🎮 RTX 5070 Ti 16GB 🧠 Qwen3.8-27B đŸ“Ļ UD-Q3_K_XL GGUF 🚀 Fully GPU-offloaded đŸĻ™ beellama.cpp + Kvarn optimizations 🔮 MTP n=2 🧮 Kvarn3 KV cache 📚 96,256 context ⚡ up to 75 tok/s generation đŸ“Ĩ up to 1,700 tok/s prefill đŸĒŸ Windows The model incrementally built a browser-based village simulation with: 🏠 housing đŸŒĻī¸ weather + seasons 🌙 day/night cycles 🍖 hunger đŸĒĩ resources 💀 deaths đŸšļ obstacle avoidance 👨‍🌾 autonomous villagers đŸŽ¯ The user deliberately chose a Q3_K_XL model + heavily compressed KV cache so the entire 27B model, MTP and ~96K context could stay inside 16GB VRAM instead of spilling weights to CPU. Their conclusion was not to fear Q3 models or aggressive KV quantization when the alternative is CPU offload and a massive speed hit. 😁đŸ”Ĩ âš ī¸ This also uses beellama.cpp/Kvarn, not stock llama.cpp. 🔗 Reddit /r/LocalLLaMA/comments/1w821fg/ninfer_vs_llamacpp_vs_vllm_quality_speed/
Show more