๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

Md Ismail ล ojal๎จ€ ๐Ÿ•ท๏ธ
@0x0SojalSec
Cyber_Security_Re-searcher || Ai Re-searcher || AI-Sec|| Malware Analysis II iOS || Pwn || 0SINT || Project AI-StrikeSec || 0ldAccounts Suspended @0xSojalSec ||
๊ฐ€์ž… October 2021
6K ํŒ”๋กœ์ž‰ ์ค‘    55.7K ํŒฌ
You donโ€™t need 128GB of unified memory to run Qwen 3.8 Next Flash & DeepSeek V4 Flash. Run Qwen 3.8 Next Flash on Mac 20GB RAM it Wallm streams routed experts off the SSD instead of loading the full checkpoint into unified memory. - feeds on Codex, 11K input: 60 tok/s prefill, 8 tok/s decode. - TTFT is the pain: 17โ€“76 s depending on prompt length - Peak RAM: 23 to 36 GiB as context grows This is not MLX loading a 2-bit quant into 96GB+. Custom Swift/Metal runtime, Experts live on NAND until the router asks for them.
๋” ๋ณด๊ธฐ