Titans: Learning to Memorize at Test Time
Adds a neural long-term memory module that updates its weights at test time, paired with attention for short-term context; scales past 2M tokens.
Memory is a small network trained online by a 'surprise' signal (gradient), unlike a fixed-size hidden state. Positions attention as short-term memory and the neural module as long-term. Basis for Google's later Nested Learning/Hope work.
- Date
- Tuesday, 31 December 2024
- Lab
- Google Research
- Kind
- paper
- Access
- paper only
Figures
| Measure | Value | Measured by |
|---|---|---|
| Context length | >2M tokens needle-in-haystack and language modelling vs Transformers and linear RNNs | authors |
Authors Ali Behrouz, Peilin Zhong, Vahab Mirrokni; arXiv v1 2024-12-31. Results at small scale; independent replications reported mixed and Google has not publicly said Gemini uses Titans (not found).
Sources
This record was checked against its sources on 6 October 2026. How we check