๐คฏ llama.cpp made DeepSeek V4 prefill up to 72% faster on a 4ร RTX 3090 rig! ๐ New committed PR.
Basically one new way of splitting the model.
And it merged into llama.cpp mainline TODAY. ๐ฅ
PR #
26490# adds tensor splitting for DeepSeek 4.
Independently tested ...
๐ง DeepSeek-V4-Flash-0731
๐ฆ UD-IQ2_M โ 84.7GB
๐ฅ 4ร RTX 3090 24GB
๐ฅ๏ธ Old Threadripper 1950X
๐ PCIe 3.0
โ No NVLink
15K prompt processing:
Layer split โ 369 tok/s
Tensor split โ 636 tok/s
๐ +72.3% PREFILL
And VRAM became almost perfectly balanced across all 4 GPUs.
๐ฏThe PR author also reports about +50% prefill on 4ร RTX 4090s.
Generation on the 3090 rig actually went:
40.3 โ 38.5 tok/s
So this won't make the words come out 72% faster.
But if you're feeding DeepSeek huge prompts, documents, RAG context or codebases, it will process them faster.
๐ฅ llama.cpp PR #
26490#