๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

Md Ismail ล ojal๎จ€ ๐Ÿ•ท๏ธ
@0x0SojalSec
Cyber_Security_Re-searcher || Ai Re-searcher || AI-Sec|| Malware Analysis II iOS || Pwn || 0SINT || Project AI-StrikeSec || 0ldAccounts Suspended @0xSojalSec ||
๊ฐ€์ž… October 2021
6K ํŒ”๋กœ์ž‰ ์ค‘    55.7K ํŒฌ
You can Run DeepSeek V4 Flash (33B MoE) running on 8 GB RAM, CPU-only, zero GPU. - Cold first-token: 5.33s - Max RSS: 5.9-6.2 GiB - Process swaps: 0 The entire model stayed memory-mapped on NVMe. Linux demand-paged only the needed weights into RAM. This is not fast, Itโ€™s not production. But it proves something important: Model size vs RAM is now a performance problem, not a hard barrier. VRAM, RAM, NVMe is becoming a real memory hierarchy for local AI. -
๋” ๋ณด๊ธฐ