Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
This blog post introduces ReasonIF, a benchmark for evaluating instruction-following in reasoning traces of large reasoning models (LRMs). It finds that frontier LRMs such as GPT-OSS-120B, Qwen3-235B, and DeepSeek-R1 fail to follow reasoning instructions more than 75% of the time, and performance degrades with task difficulty.
From the source
We find frontier LRMs, including GPT-OSS-120B, Qwen3-235B, and DeepSeek-R1 fail to follow reasoning instructions more than 75% of time.
together.ai