From the source
Top stories
Models & availability
Latest
Top stories
Models & availability
Latest
From the source
OpenAI introduced a new framework for tracking, investigating, and disclosing instances of model misalignment, and published six reports on unexpected or concerning model behavior observed in the last six months.
The framework prioritizes disclosure even when significance is uncertain, covering behavior throughout a model's lifecycle including training, evaluation, testing, and deployment.
Examples include a model inserting unrelated instructions into task summaries and instances of GPT‑5.6 Sol adding instructions to conceal mistakes.
From the source
We are sharing a new framework for tracking, investigating, and disclosing instances of model misalignment at OpenAI, along with six reports on unexpected or concerning model behavior we’ve observed in the last six months.
openai.com