Register and share your invite link to earn from video plays and referrals.

Teortaxes▶️ (DeepSeek 推特🐋铁粉 2023 – ∞)
@teortaxesTex
We're in a race. It's not USA vs China but humans and AGIs vs ape power centralization. @deepseek_ai stan #1#, 2023–Deep Time «C’est la guerre.» ®1
Joined September 2010
3.3K Following    76.9K Followers
What I find annoying is that everyone celebrates how cheap Engram is but nobody wants to ask "is Engram any good tho?" Like, you're not getting a free offload of FFNs to DRAM. How much does it add to modeling? I trust DeepSeek, but that's me. Does… everyone just trust them?
Show more
🫡梁文峰在记忆层动了手脚!为什么bDeepSeek V4.1 Flash 又便宜又强?!不是堆了更多显卡,是把“背课文”从 GPU 里拆走了! V4.1 Flash 的 Engram 把常见词组变成查表,不该算的不再占显存。 SemiAnalysis 拆完记忆层,卡上的数字更刺眼: 🔹输入只激活 80 亿、输出 160 亿,剩下约 1960 亿是查找表 🔹浅层记物体,深层记关系,熟词组直接查,推理留给模型 🔹表只看 token ID,能扔进内存;B300 上从 TP4 降到 TP2,性价比最高抬约 1.6 倍 🔹4 张 GB300 把表卸到主机后,KV 容量还能再涨约 36% 🔹B200 上内存卸载打赢 SSD:同等体验附近,每美元总 token 约 1.21 亿对 5200 万 🔹双机 GB300 上,预填充能到约 5.6 万 token/秒,单用户解码约 253 token/秒 便宜来自少算、少占 HBM; 强来自该查就查、该算就算。 显存带宽比显存容量更值钱。 #DeepSeek# #V41Flash# #Engram# #HBM# #B300# #GB300# #SemiAnalysis# #梁文峰# #梁圣#
Show more