AI Research Atlas

Weak-to-Strong Generalization

OpenAI · 14 December 2023

A GPT-2-level supervisor can elicit much of GPT-4's capability, an empirical analogy for humans supervising superhuman models.

Fine-tunes strong models on labels from weak models across NLP, chess and reward modeling; strong students beat their weak teachers, and an auxiliary confidence loss helps further. Reframes superalignment as testable.

Date
Thursday, 14 December 2023
Lab
OpenAI
Kind
paper
Access
research preview

Figures

MeasureValueMeasured by
GPT-4 supervised by GPT-2-level labels with confidence lossapproaches GPT-3.5-level performance on NLP tasks
partial, not full, recovery of capability
company

Authors include Collin Burns, Pavel Izmailov, Jan Leike, Ilya Sutskever and Jeff Wu.

Sources

  1. arxiv.org/abs/2312.09390

This record was checked against its sources on 6 October 2026. How we check