Register and share your invite link to earn from video plays and referrals.

Search results for Engram
Engram community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including Engram
Engram tables in DSV4.1 are super cool! Everything is just compression, compression, compression. After that, it's all caching, caching, caching.
Engrams Embedding Entendre: Codesign for Efficient DRAM/SSD Offloading New Model Architecture Implications for TAM of DRAM/NVMe, DeepSeek V4.1 Flash, AgentX, InferenceX, NVMe experiments
Show more
Experiments have shown that offloading Engram to DRAM results in up to 50% better performance on H200, B200, B300, and even GB300 NVL72. Our technical team has collaborated with 10x engineers from @EmadBarsoumPi & @AnushElangovan to upstream support for @vllm_project ROCm Engram DRAM offloading too, and we are already seeing up to 50% better perf too in PR 57491. As such, in PRs 985, 1002, 1005, 1006 @NVIDIA, @AIatAMD and @SemiAnalysis_ are updating the default recommended recipes to use Engram offloading.
Show more
Ngram module (figure from DeepSeek's Engram) They first ablate at which layer to put it but couldnt find a clear winner. As such they just put one at layer 2 which allows it to be prefetched from CPU together with layer 1 computation From the figures in this section, this seems to be the biggest area where loss is not a good signal on its own. Downstream loss goes down as the vocab size increases, but same isnt true for evals
Show more
Thank you to @Nasdaq for supporting Engram on our launch day yesterday! Some have commented that this photo looks AI-generated. It's not. This really happened. Feel free to send this picture to your moms. We're certainly going to.
Show more
Recent research has highlighted the promise of scaling memory embeddings in LLM training. While Engram and STEM index memory by token identity or local n-grams, can we design more flexible memory routing that captures how each token’s meaning changes with context? This paper introduces Mixture-of-Memory Embeddings (MoME), which uses a learned router to sparsely select among multiple memory slots for each token. The architecture outperforms strong memory baselines across three model families. Interestingly, both qualitative and quantitative analyses show that the learned routers’ activations correlate with the context-dependent senses of polysemous words. The authors also train a sub-billion-parameter model that achieves competitive CORE-22 performance against similarly sized base models, including Qwen3-0.6B and Llama 3.2-1B. Code and pretrained models are publicly available!
Show more
MONEY PRINTER ALERT🚨 NVIDIA vLLM B200 CAN GENERATE UP TO💰️$15 BILLION💰️OF ANNUAL PROFITS PER GIGAWATT serving the open DeepSeekv4.1 Flash model at the official interactivity & official selling prices. Using Engram DRAM offloading on NVIDIA results in a 50% increase in revenue per GigaWatt.
Show more
Nvidia dropped an official DeepSeek-V4.1-Flash NVFP4 build. And this thing is BIG. DeepSeek V4.1 Flash stats 🧠 552B backbone 📚 +196B Engram conditional memory ⚡ only 8B active during prefill 🚀 16B active during decode 👁️ native vision 📖 1 MILLION token context 🧩 384 routed experts across 40 layers 📜 MIT license Nvidia has converted its routed MoE experts to NVFP4 W4A4 specifically for Blackwell GPUs. This is not a compressed giant model down to 4-bit. DeepSeek's experts were already stored in MXFP4, Nvidia instead converts them to its Blackwell-friendly NVFP4 format. And because NVFP4 uses finer scaling, the checkpoint actually gets slightly larger. 💾 Source: ~476 GiB 💾 NVIDIA NVFP4: ~492 GiB 48 safetensor shards. 😳 So why bother? Because NVIDIA is optimizing how those 4-bit experts execute on Blackwell. And impressively, NVIDIA's evaluations show basically no obvious quality collapse from the conversion. For example: 🧠 GPQA Diamond 91.04 → 91.29 💻 SciCode 54.40 → 55.84 🛠️ Terminal-Bench 2.1 81.60 → 82.16 👁️ MMMU-Pro 74.05 → 73.70 Some slightly up. Some slightly down. Essentially benchmark parity. And it already has: ✅ vLLM support ✅ SGLang support ✅ reasoning parser ✅ tool calling ✅ image input ✅ 1M context ⚠️ Nvidia validated it on 4× GB300 GPUs so not a local model (yet). The checkpoint is still ~492 GiB. 🔗 HF: /nvidia/DeepSeek-V4.1-Flash-NVFP4
Show more
Congrats to @deepseek_ai on releasing DeepSeek-V4.1-Flash! > 552B backbone, with a causal encoder-decoder activating just 8B params at prefill, 16B at decode > 196B Engram memory accessed through sparse lookups > ~8x less persistent KV than V4-Flash through bounded replay
Show more
Deepseek V4.1 Flash 552B total, 8/16B active with a new arch trained on 45T tokens, there are different active parameters for input/output tokens with the encoder/decoder arch, engram, new sparse attention, new mHC, native vision very high benchmarks (beating K3), insane efficiency, and as always amazing tech report this is probably the most novel arch i've seen in a while, pretty insane
Show more
0
34
1.3K
114
Forward to community