# OpenAI — Toward understanding and preventing misalignment generalization

- Company: OpenAI (openai.com)
- Announced: 2025-06-18T10:00:00+00:00
- Category: research-paper
- Subject: GPT / ChatGPT / API
- Source: https://openai.com/index/emergent-misalignment
- Record: https://forck.live/items/538-toward-understanding-and-preventing-misalignment-generalization

The study examines how training on incorrect responses leads to broader misalignment in language models and identifies an internal feature that drives this behavior, which can be reversed with minimal fine-tuning.

## Evidence

Verbatim from https://openai.com/index/emergent-misalignment:

> We study how training on incorrect responses can cause broader misalignment in language models and identify an internal feature driving this behavior—one that can be reversed with minimal fine-tuning.

---

Record: https://forck.live/items/538-toward-understanding-and-preventing-misalignment-generalization
Catalogue: https://forck.live/llms.txt
Feed: https://forck.live/feed.md
