AI Research Atlas

Quiet-STaR

Stanford · 14 March 2024

Trains a model to generate a hidden rationale at every token during pretraining-style learning; zero-shot GSM8K 5.9% to 10.9%.

Generalises STaR from curated Q&A to arbitrary web text. The model samples 'thoughts' in parallel at each position, and is rewarded when the thought improves prediction of future tokens. An early precursor of RL on thinking tokens, without task-specific fine-tuning.

Date
Thursday, 14 March 2024
Lab
Stanford
Kind
paper
Access
paper only

Figures

MeasureValueMeasured by
GSM8K (zero-shot)5.9% -> 10.9%
after continued pretraining on web text
authors
CommonsenseQA (zero-shot)36.3% -> 47.2%authors

Authors are Eric Zelikman, Georges Harik, Yijia Shao, Varuna Jayasiri, Nick Haber and Noah Goodman. Builds on STaR (2022) by the same first author. arXiv v1 2024-03-14.

Sources

  1. arxiv.org/abs/2403.09629

This record was checked against its sources on 6 October 2026. How we check