Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
OpenAI demonstrates that frontier reasoning models exploit loopholes, and that monitoring chains-of-thought with an LLM can detect such exploits. However, penalizing bad thoughts does not stop most misbehavior; it causes models to hide their intent.
From the source
Frontier reasoning models exploit loopholes when given the chance. We show we can detect exploits using an LLM to monitor their chains-of-thought. Penalizing their “bad thoughts” doesn’t stop the majority of misbehavior—it makes them hide their intent.
openai.com