Engram: conditional memory via scalable lookup
Adds a second sparsity axis to LLMs, hashed N-gram embedding lookup for static knowledge alongside MoE compute. Engram-27B beats an iso-FLOPs MoE-27B.
Retrieves frequent N-gram patterns by O(1) hashed lookup so the backbone spends capacity on reasoning, with tables offloadable to host memory. Reports a U-shaped scaling law for splitting parameters between experts and memory; gains on MMLU, BBH, HumanEval and long context.
- Date
- Monday, 12 January 2026
- Lab
- DeepSeek
- Kind
- paper
- Access
- research preview
Figures
| Measure | Value | Measured by |
|---|---|---|
| Multi-Query NIAH | 84.2 to 97.0 MMLU +3.4, BBH +5.0, HumanEval +3.0 vs MoE baseline at iso-parameter, iso-FLOPs | company |
Code released under Apache 2.0, paper CC BY 4.0. Lead author Xin Cheng (Peking University) with DeepSeek co-authors. Not part of the V4 architecture list in the V4 report abstract (which names hybrid attention, mHC and Muon); TileKernels does include Engram gating kernels.
Sources
This record was checked against its sources on 6 October 2026. How we check