Register and share your invite link to earn from video plays and referrals.

Search results for llamacpp
llamacpp community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including llamacpp
A 27B model with 96K context built an actual game locally on a 16GB RTX 5070 Ti at 75 tps. This wasn't a benchmark but an actual working villager simulation. Setup 👇 🎮 RTX 5070 Ti 16GB 🧠 Qwen3.8-27B 📦 UD-Q3_K_XL GGUF 🚀 Fully GPU-offloaded 🦙 beellama.cpp + Kvarn optimizations 🔮 MTP n=2 🧮 Kvarn3 KV cache 📚 96,256 context ⚡ up to 75 tok/s generation 📥 up to 1,700 tok/s prefill 🪟 Windows The model incrementally built a browser-based village simulation with: 🏠 housing 🌦️ weather + seasons 🌙 day/night cycles 🍖 hunger 🪵 resources 💀 deaths 🚶 obstacle avoidance 👨‍🌾 autonomous villagers 🎯 The user deliberately chose a Q3_K_XL model + heavily compressed KV cache so the entire 27B model, MTP and ~96K context could stay inside 16GB VRAM instead of spilling weights to CPU. Their conclusion was not to fear Q3 models or aggressive KV quantization when the alternative is CPU offload and a massive speed hit. 😁🔥 ⚠️ This also uses beellama.cpp/Kvarn, not stock llama.cpp. 🔗 Reddit /r/LocalLLaMA/comments/1w821fg/ninfer_vs_llamacpp_vs_vllm_quality_speed/
Show more
llama.cpp with MTP support makes local models fast enough to use as daily drivers 🚀 Qwen3.6-27B dense generation below on A10G: From 25 tok/st to 45 tok/s (+78%)!
llama.cpp adds MTP for the Qwen3.6 family This is a significant milestone for the local AI ecosystem. The performance jump with these changes is massive and elevates local inference on commodity hardware further. Special thanks to Aman Gupta for leading this development!
Show more
llama.cpp now supports Qwen3-ASR, Qwen3-Omni and Gemma 4 audio/vision input 🔥 Mixed modalities is the future 😼😼
llama.cpp at 100k stars now that 90% of the code worldwide is being written by AI agents, I predict that within 3-6 months, 90% of all AI agents will be running locally with llama.cpp 😄 Jokes aside, I am going to use this small milestone as an opportunity to reflect a bit on the project and the state of AI from the perspective of local applications. There is a lot to say and discuss and yet it feels less and less important to try to make a point. Opinions about viability of local LLMs are strongly polarized, details are overlooked, the scientific approach is lacking. Arguments are predominantly based on vibes and hype waves. One thing is clear though - local LLMs are used more and more. I expect this trend to continue and likely 2026 will end up being one of the most important years for the local AI movement. I admit that I didn't expect the agentic era to come so quickly to the local LLM space. One year ago, the available models were too computationally expensive for doing long-context tasks. There wasn't an obvious path towards meaningful agentic applications. The memory and compute requirements were huge. Last summer, with the release of gpt-oss, things started to change. It was the first time we saw a glimpse of tool calling that actually works well within the resource constraints of our daily devices. Later in the year, even better models were released and by now, useful local agentic workflows are a reality. Comparing local vs hosted capabilities at a given moment of time is pointless. To try put things into perspective: - We don't need frontier intelligence to automate searches and sending emails - We don't need trillion parameter models to be able to summarize articles or technical documents - We don't need massive GPU data centers to control our home appliances or turn the lights off in the garage I believe that there is a certain level of intelligence we as humans can comprehend and meaningfully utilize to improve our working process. Beyond that level, access to more intelligence becomes unnecessary at best and counterproductive at worst. I also believe that that level of useful artificial intelligence is completely within reach locally and it has always been just a matter of implementing the right software stack to bring it to the end user. With llama.cpp, I am confident that we continue to be on the right track of building that software stack! The llama.cpp project is going stronger than ever. With more than 1500 contributors, the project keeps growing steadily. From technical point of view, I think that llama.cpp + ggml is the only solution that actually makes sense. That is, the software stack must run efficiently on every possible device, hardware and operating system. The technology is too important to be vendor-locked. It has to be developed in the open, by the community, together with the independent hardware vendors. This is the only right way to build something that will truly make a difference in the long run. I won't try to convince you about what is currently and will be possible with local AI. We will just continue to build as usual. I am confident that after the smoke clears and we look objectively at what we have built together, the benefits will be obvious to everyone. Big shoutout to all llama.cpp maintainers. I feel extremely lucky to be able to work together with so many talented contributors. Every day I learn something new and I feel there is so much more cool stuff that we are going to build. Also, I am really thankful that the project continues to have reliable partners to support it! Cheers!
Show more
0
277
3.2K
392
Forward to community
The Llama app for Mac now comes with a simple request builder for llama.cpp's REST API
🤯 llama.cpp made DeepSeek V4 prefill up to 72% faster on a 4× RTX 3090 rig! 👀 New committed PR. Basically one new way of splitting the model. And it merged into llama.cpp mainline TODAY. 🔥 PR #26490# adds tensor splitting for DeepSeek 4. Independently tested ... 🧠 DeepSeek-V4-Flash-0731 📦 UD-IQ2_M — 84.7GB 🔥 4× RTX 3090 24GB 🖥️ Old Threadripper 1950X 🔌 PCIe 3.0 ❌ No NVLink 15K prompt processing: Layer split → 369 tok/s Tensor split → 636 tok/s 🚀 +72.3% PREFILL And VRAM became almost perfectly balanced across all 4 GPUs. 🎯The PR author also reports about +50% prefill on 4× RTX 4090s. Generation on the 3090 rig actually went: 40.3 → 38.5 tok/s So this won't make the words come out 72% faster. But if you're feeding DeepSeek huge prompts, documents, RAG context or codebases, it will process them faster. 🔥 llama.cpp PR #26490#
Show more
K2 Horizon got another llama.cpp-family path, this time through TurboQuant. Spark-X2.5 already landed in upstream llama.cpp. I’ve actually got it running myself. A brand-new K2 Horizon TurboQuant PR adds … 🧠 K2 Horizon model support 📦 Hugging Face → GGUF conversion 💬 tokenizer + chat templates ⚙️ llama.cpp-style inference And they’ve already tested 🧠 K2-Horizon-3.7B Q4_K_M 🍎 M3 Pro / 18GB ⚡ ~35 tok/s 📚 64K configured context The PR also adds Spark-X2.5 to the TurboQuant fork: ⚡ ~36 tps 📚 256K configured context Caveats ⚠️ This is still an OPEN PR ⚠️ Only Metal has been tested ⚠️ Other K2 sizes + MoVA/MoE variants remain untested K2 Horizon already has an IFM llama.cpp fork, but it still isn’t supported in normal upstream llama.cpp. So what I’m watching here is whether TurboQuant becomes another path for K2.
Show more
My talk about llama.cpp speculative decoding (MTP, dflash, dspark) at dotAI 🦙🦙 replay available soon!
Designed for local llama.cpp use. Please validate thoroughly for your workload. -