AI Research Atlas

GPT-Red: self-improvement for robustness

OpenAI · 15 July 2026

GPT-Red, an automated red-teaming model trained by self-play, attacks OpenAI models and is used to adversarially train GPT-5.6 against prompt injection.

Human red-teaming does not scale and standard robustness evals are saturated, so OpenAI trained an internal-only attacker that sends prompts, observes responses and iterates. Earlier models were highly vulnerable to its prompt-injection attacks; GPT-5.6 was adversarially trained on its outputs and became much less vulnerable.

Date
Wednesday, 15 July 2026
Lab
OpenAI
Kind
paper
Access
research preview

OpenAI frames automated red-teaming as safety "self-improvement", meaning it uses today's models to make future models safer. Robustness numbers are company-measured and not independently verified.

Sources

  1. openai.com/index/unlocking-self-improvement-gpt-red/

This record was checked against its sources on 6 October 2026. How we check

Related