AI Research Atlas

gpt-oss-safeguard-120b and 20b

OpenAI · 29 October 2025

gpt-oss-safeguard, open-weight reasoning classifiers fine-tuned from gpt-oss, interpret a developer-written safety policy at inference time.

Rather than training a classifier on labeled data, the model reasons over a policy supplied at inference, so policies can change without retraining. Research preview under Apache 2.0, built with ROOST (a nonprofit trust-and-safety tooling group).

Date
Wednesday, 29 October 2025
Lab
OpenAI
Kind
open-weights
Access
open weights
Price
free; Apache 2.0

Sources

  1. openai.com/index/introducing-gpt-oss-safeguard/
  2. developers.openai.com/api/docs/changelog

This record was checked against its sources on 6 October 2026. How we check

Related