GPT-Red: self-improvement for robustness
GPT-Red, an automated red-teaming model trained by self-play, attacks OpenAI models and is used to adversarially train GPT-5.6 against prompt injection.
Human red-teaming does not scale and standard robustness evals are saturated, so OpenAI trained an internal-only attacker that sends prompts, observes responses and iterates. Earlier models were highly vulnerable to its prompt-injection attacks; GPT-5.6 was adversarially trained on its outputs and became much less vulnerable.
- Date
- Wednesday, 15 July 2026
- Lab
- OpenAI
- Kind
- paper
- Access
- research preview
OpenAI frames automated red-teaming as safety "self-improvement", meaning it uses today's models to make future models safer. Robustness numbers are company-measured and not independently verified.
Sources
This record was checked against its sources on 6 October 2026. How we check