From the source
Lead story
Ask your AI
Top stories
Models & availability
Latest
Lead story
Top stories
Models & availability
Latest
From the source
Transformers now supports running GGUF quantized models efficiently on Apple Silicon Macs
Hugging Face added support for running GGUF-quantized models efficiently in transformers through the familiar transformers APIs.
Users can load GGUF checkpoints from the Hub using from_pretrained and generate locally on their machines, with initial focus on Apple Silicon and the Qwen3.5 architecture.
From the source
We're adding support for running GGUF models efficiently in transformers, so you can use checkpoints sized for your laptop's memory through the familiar transformers APIs. Pick a GGUF from the Hub, load it with from_pretrained, and start generating on your own machine.
huggingface.co