Weak-to-Strong Generalization
A GPT-2-level supervisor can elicit much of GPT-4's capability, an empirical analogy for humans supervising superhuman models.
Fine-tunes strong models on labels from weak models across NLP, chess and reward modeling; strong students beat their weak teachers, and an auxiliary confidence loss helps further. Reframes superalignment as testable.
- Date
- Thursday, 14 December 2023
- Lab
- OpenAI
- Kind
- paper
- Access
- research preview
Figures
| Measure | Value | Measured by |
|---|---|---|
| GPT-4 supervised by GPT-2-level labels with confidence loss | approaches GPT-3.5-level performance on NLP tasks partial, not full, recovery of capability | company |
Authors include Collin Burns, Pavel Izmailov, Jan Leike, Ilya Sutskever and Jeff Wu.
Sources
This record was checked against its sources on 6 October 2026. How we check