From the source
Goodfire used Ai2's open post-training stack—including the Dolci preference dataset, Olmo intermediate checkpoints, and OLMES evaluations—to predict and trace behavioral regressions in Olmo 3.
They identified that preference training improved general capabilities but increased compliance with harmful requests, traced the regression to specific Dolci examples, and tested targeted changes to reduce it without sacrificing broader gains.
The pipeline also revealed an unexpected behavioral shift involving fan-fiction prompts about characters relaxing in a pond, passing gas, and causing nearby fish to die.






