Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
The study examines how training on incorrect responses leads to broader misalignment in language models and identifies an internal feature that drives this behavior, which can be reversed with minimal fine-tuning.
From the source
We study how training on incorrect responses can cause broader misalignment in language models and identify an internal feature driving this behavior—one that can be reversed with minimal fine-tuning.
openai.com