AI Research Atlas

End-to-End Test-Time Training for Long Context (TTT-E2E)

Stanford / NVIDIA / Astera / UC Berkeley / UCSD · 29 December 2025

A sliding-window Transformer keeps learning on its context at test time via next-token prediction, matching full-attention scaling with constant latency (2.7x faster at 128K).

Treats long-context as continual learning, so the model compresses what it reads into its weights, with the initialisation meta-learned in training. Needs no new architecture; at 128K it is 2.7x faster than full attention in the authors' tests.

Date
Monday, 29 December 2025
Lab
Stanford / NVIDIA / Astera / UC Berkeley / UCSD
Kind
paper
Access
paper only

Figures

MeasureValueMeasured by
Speedup vs full attention at 128K2.7x
3B models on 164B tokens; constant per-token latency
authors

arXiv v1 2025-12-29; authors incl. Yu Sun, Yejin Choi, Carlos Guestrin, Sanmi Koyejo. Code on GitHub. Training cost of the meta-learning stage is higher than standard training; small-scale (3B).

Sources

  1. arxiv.org/abs/2512.23675

This record was checked against its sources on 6 October 2026. How we check

Related