🤯 llama.cpp made DeepSeek V4 prefill up to 72% faster on a 4× RTX 3090 rig! 👀 New committed PR.
Basically one new way of splitting the model.
And it merged into llama.cpp mainline TODAY. 🔥
PR #
26490# adds tensor splitting for DeepSeek 4.
Independently tested ...
🧠 DeepSeek-V4-Flash-0731
📦 UD-IQ2_M — 84.7GB
🔥 4× RTX 3090 24GB
🖥️ Old Threadripper 1950X
🔌 PCIe 3.0
❌ No NVLink
15K prompt processing:
Layer split → 369 tok/s
Tensor split → 636 tok/s
🚀 +72.3% PREFILL
And VRAM became almost perfectly balanced across all 4 GPUs.
🎯The PR author also reports about +50% prefill on 4× RTX 4090s.
Generation on the 3090 rig actually went:
40.3 → 38.5 tok/s
So this won't make the words come out 72% faster.
But if you're feeding DeepSeek huge prompts, documents, RAG context or codebases, it will process them faster.
🔥 llama.cpp PR #
26490#