# Together AI — Large Reasoning Models Fail to Follow Instructions During Reasoning: A Benchmark Study

- Company: Together AI (together.ai)
- Announced: 2025-10-22T00:00:00+00:00
- Category: research-paper
- Subject: Inference platform
- Models affected: GPT-OSS-120B, Qwen3-235B, DeepSeek-R1
- Source: https://www.together.ai/blog/large-reasoning-models-fail-to-follow-instructions-during-reasoning-a-benchmark-study
- Record: https://forck.live/items/2404-large-reasoning-models-fail-to-follow-instructions-during-reasoning-a

This blog post introduces ReasonIF, a benchmark for evaluating instruction-following in reasoning traces of large reasoning models (LRMs). It finds that frontier LRMs such as GPT-OSS-120B, Qwen3-235B, and DeepSeek-R1 fail to follow reasoning instructions more than 75% of the time, and performance degrades with task difficulty.

## Evidence

Verbatim from https://www.together.ai/blog/large-reasoning-models-fail-to-follow-instructions-during-reasoning-a-benchmark-study:

> We find frontier LRMs, including GPT-OSS-120B, Qwen3-235B, and DeepSeek-R1 fail to follow reasoning instructions more than 75% of time.

---

Record: https://forck.live/items/2404-large-reasoning-models-fail-to-follow-instructions-during-reasoning-a
Catalogue: https://forck.live/llms.txt
Feed: https://forck.live/feed.md
