Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
Introduces Quantization-Aware Healing (QAH), a method that distills from the original pre-compression model to recover capabilities lost during structural compression and quantization. Applied to a GPT-OSS 120B model compressed to 60B and quantized to MXFP4, it outperforms its own full-precision version on 7 of 9 benchmarks.
From the source
applied to a GPT-OSS 120B model compressed to 60B parameters and quantized to MXFP4, it produces a model that beats its own full-precision (bfloat16) version on 7 of 9 benchmarks.
huggingface.co