Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
ServiceNow AI converted their 15B reasoning model to a Mamba hybrid (Apriel-H1-15b-Thinker-SFT) achieving 2.1x throughput with minimal quality loss, using reverse KL divergence and staged distillation on high-quality reasoning traces.
From the source
We converted our 15B reasoning model to a Mamba hybrid achieving 2.1x throughput with minimal quality loss.
huggingface.co