From the source
TII released Falcon H1R 7B, a decoder-only large language model built on the Falcon-H1 Base.
The 7B-parameter model matches or outperforms reasoning models 2–7× larger on math, code, and general benchmarks.
Its training uses a two-stage pipeline of supervised fine-tuning and GRPO reinforcement learning, and it employs a Deep Think with Confidence (DeepConf) method for test-time scaling.




