๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

David Hendrickson
@TeksEdge
CEO & Founder | PhD | Startup Advisor | @Columbia | Author Generative Software Engineering | ๐Ÿ”” Follow for AI & Vibe Coding Tips ๐Ÿ‘‡
๊ฐ€์ž… July 2023
549 ํŒ”๋กœ์ž‰ ์ค‘    10.6K ํŒฌ
๐Ÿคฏ llama.cpp made DeepSeek V4 prefill up to 72% faster on a 4ร— RTX 3090 rig! ๐Ÿ‘€ New committed PR. Basically one new way of splitting the model. And it merged into llama.cpp mainline TODAY. ๐Ÿ”ฅ PR #26490# adds tensor splitting for DeepSeek 4. Independently tested ... ๐Ÿง  DeepSeek-V4-Flash-0731 ๐Ÿ“ฆ UD-IQ2_M โ€” 84.7GB ๐Ÿ”ฅ 4ร— RTX 3090 24GB ๐Ÿ–ฅ๏ธ Old Threadripper 1950X ๐Ÿ”Œ PCIe 3.0 โŒ No NVLink 15K prompt processing: Layer split โ†’ 369 tok/s Tensor split โ†’ 636 tok/s ๐Ÿš€ +72.3% PREFILL And VRAM became almost perfectly balanced across all 4 GPUs. ๐ŸŽฏThe PR author also reports about +50% prefill on 4ร— RTX 4090s. Generation on the 3090 rig actually went: 40.3 โ†’ 38.5 tok/s So this won't make the words come out 72% faster. But if you're feeding DeepSeek huge prompts, documents, RAG context or codebases, it will process them faster. ๐Ÿ”ฅ llama.cpp PR #26490#
๋” ๋ณด๊ธฐ