๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

Inty News
@__Inty__
โค๏ธ๐Ÿ‡บ๐Ÿ‡ธ ๐Ÿ‡บ๐Ÿ‡ธโค๏ธ ๆˆ‘ไธบไฝ ๅˆ†ไบซไธ–็•Œ็ƒญ็‚นๆ–ฐ้—ป | ๆ—ถๆ”ฟ่Šๅคฉ็พค ๐Ÿ‘‡
๊ฐ€์ž… August 2018
43 ํŒ”๋กœ์ž‰ ์ค‘    752.3K ํŒฌ
1 token/็ง’ ๐Ÿ˜‚
You can Run DeepSeek V4 Flash (33B MoE) running on 8 GB RAM, CPU-only, zero GPU. - Cold first-token: 5.33s - Max RSS: 5.9-6.2 GiB - Process swaps: 0 The entire model stayed memory-mapped on NVMe. Linux demand-paged only the needed weights into RAM. This is not fast, Itโ€™s not production. But it proves something important: Model size vs RAM is now a performance problem, not a hard barrier. VRAM, RAM, NVMe is becoming a real memory hierarchy for local AI. -
๋” ๋ณด๊ธฐ