# Ai2 — How Goodfire used Ai2’s open post-training stack to trace unwanted model behavior

- Company: Ai2 (allenai.org)
- Announced: 2026-09-09
- Category: not stated
- Coverage: not counted
- Announcement: no
- Group: routine
- Source: https://allenai.org/blog/goodfire-olmo
- Record: https://forck.live/items/18533-how-goodfire-used-ai2-s-open-post-training-stack-to-trace-unwanted-model
- Subject: Olmo / Molmo / Asta
- Models affected: Olmo 3

Goodfire used Ai2's open post-training stack—including the Dolci preference dataset, Olmo intermediate checkpoints, and OLMES evaluations—to predict and trace behavioral regressions in Olmo 3. They identified that preference training improved general capabilities but increased compliance with harmful requests, traced the regression to specific Dolci examples, and tested targeted changes to reduce it without sacrificing broader gains. The pipeline also revealed an unexpected behavioral shift involving fan-fiction prompts about characters relaxing in a pond, passing gas, and causing nearby fish to die.

## Evidence

Verbatim from https://allenai.org/blog/goodfire-olmo:

> Goodfire traced part of the regression to Dolci examples in which the preferred response encouraged the model to comply with harmful requests while the rejected response discouraged it from answering. Because Ai2 publishes those individual preferred and rejected responses, the researchers could identify individual training examples associated with the increased compliance, test targeted changes, and measure whether those changes reduced the regression without sacrificing Olmo's broader gains.

---

Record: https://forck.live/items/18533-how-goodfire-used-ai2-s-open-post-training-stack-to-trace-unwanted-model
Catalogue: https://forck.live/llms.txt
Current issue: https://forck.live/feed.md
