Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
Hugging Face announces upgrades to the transformers library, including zero-build kernels downloadable from the Hub, MXFP4 quantization, tensor parallelism, expert parallelism, dynamic sliding window layer and cache, continuous batching and paged attention, and faster model loading, inspired by OpenAI's GPT-OSS release. These features are designed to work across major models in transformers.
From the source
To enable the release of gpt-oss through transformers, we have upgraded the library considerably. The updates make it very efficient to load, run, and fine-tune the models.
huggingface.co