Quiet-STaR
Trains a model to generate a hidden rationale at every token during pretraining-style learning; zero-shot GSM8K 5.9% to 10.9%.
Generalises STaR from curated Q&A to arbitrary web text. The model samples 'thoughts' in parallel at each position, and is rewarded when the thought improves prediction of future tokens. An early precursor of RL on thinking tokens, without task-specific fine-tuning.
- Date
- Thursday, 14 March 2024
- Lab
- Stanford
- Kind
- paper
- Access
- paper only
Figures
| Measure | Value | Measured by |
|---|---|---|
| GSM8K (zero-shot) | 5.9% -> 10.9% after continued pretraining on web text | authors |
| CommonsenseQA (zero-shot) | 36.3% -> 47.2% | authors |
Authors are Eric Zelikman, Georges Harik, Yijia Shao, Varuna Jayasiri, Nick Haber and Noah Goodman. Builds on STaR (2022) by the same first author. arXiv v1 2024-03-14.
Sources
This record was checked against its sources on 6 October 2026. How we check