Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
Intel AI Software Group introduces DeepMath, a math reasoning agent based on Qwen3-4B Thinking fine-tuned with GRPO. The model emits Python snippets for intermediate steps, runs them in a sandbox, and uses the results in its reasoning. It reduces output lengths by up to 66% and improves accuracy on math datasets.
From the source
DeepMath is an aligned math reasoning agent built on Qwen3-4B Thinking and fine-tuned with GRPO (Group Relative Policy Optimization).
huggingface.co