Register and share your invite link to earn from video plays and referrals.

Inty News
@__Inty__
❤️🇺🇸 🇺🇸❤️ 我为你分享世界热点新闻 | 时政聊天群 👇
Joined August 2018
43 Following    752.3K Followers
1 token/秒 😂
You can Run DeepSeek V4 Flash (33B MoE) running on 8 GB RAM, CPU-only, zero GPU. - Cold first-token: 5.33s - Max RSS: 5.9-6.2 GiB - Process swaps: 0 The entire model stayed memory-mapped on NVMe. Linux demand-paged only the needed weights into RAM. This is not fast, It’s not production. But it proves something important: Model size vs RAM is now a performance problem, not a hard barrier. VRAM, RAM, NVMe is becoming a real memory hierarchy for local AI. -
Show more