Engram: Conditional Memory via Scalable Lookup
Adds an O(1) hashed N-gram lookup memory as a sparsity axis beside MoE; a 27B model gains +3.4 MMLU and 97.0 on multi-query needle retrieval.
Offloads static knowledge lookup to hashed N-gram embeddings so transformer depth is spent on reasoning. Finds a U-shaped trade-off between neural compute and memory, and beats iso-parameter, iso-FLOP MoE baselines. Not used in DeepSeek-V4 per reporting.
- Date
- Monday, 12 January 2026
- Lab
- DeepSeek
- Kind
- paper
- Access
- open weights
Figures
| Measure | Value | Measured by |
|---|---|---|
| MMLU / BBH / HumanEval gain | +3.4 / +5.0 / +3.0 27B Engram vs MoE baseline at same params and FLOPs | authors |
| Multi-Query NIAH | 84.2 -> 97.0 | authors |
arXiv v1 2026-01-12. 'Engram absent from V4' is a search-result claim from a third-party blog and was not confirmed in DeepSeek's V4 paper summary (which lists CSA/HCA attention, mHC and Muon).
Sources
This record was checked against its sources on 6 October 2026. How we check