End-to-End Test-Time Training for Long Context (TTT-E2E)
A sliding-window Transformer keeps learning on its context at test time via next-token prediction, matching full-attention scaling with constant latency (2.7x faster at 128K).
Treats long-context as continual learning, so the model compresses what it reads into its weights, with the initialisation meta-learned in training. Needs no new architecture; at 128K it is 2.7x faster than full attention in the authors' tests.
- Date
- Monday, 29 December 2025
- Lab
- Stanford / NVIDIA / Astera / UC Berkeley / UCSD
- Kind
- paper
- Access
- paper only
Figures
| Measure | Value | Measured by |
|---|---|---|
| Speedup vs full attention at 128K | 2.7x 3B models on 164B tokens; constant per-token latency | authors |
arXiv v1 2025-12-29; authors incl. Yu Sun, Yejin Choi, Carlos Guestrin, Sanmi Koyejo. Code on GitHub. Training cost of the meta-learning stage is higher than standard training; small-scale (3B).
Sources
This record was checked against its sources on 6 October 2026. How we check