# Hugging Face — Transformers now runs llama.cpp quants

- Company: Hugging Face (huggingface.co)
- Announced: 2026-09-22
- Category: capability-change
- Coverage: not counted
- Announcement: yes
- Group: announcements
- Source: https://huggingface.co/blog/transformers-llama-cpp-quants
- Record: https://forck.live/items/13000-transformers-now-runs-llama-cpp-quants
- Subject: Platform
- Models affected: Qwen3.5

Hugging Face added support for running GGUF-quantized models efficiently in transformers through the familiar transformers APIs. Users can load GGUF checkpoints from the Hub using from_pretrained and generate locally on their machines, with initial focus on Apple Silicon and the Qwen3.5 architecture.

## Evidence

Verbatim from https://huggingface.co/blog/transformers-llama-cpp-quants:

> We're adding support for running GGUF models efficiently in transformers, so you can use checkpoints sized for your laptop's memory through the familiar transformers APIs. Pick a GGUF from the Hub, load it with from_pretrained, and start generating on your own machine.

---

Record: https://forck.live/items/13000-transformers-now-runs-llama-cpp-quants
Catalogue: https://forck.live/llms.txt
Current issue: https://forck.live/feed.md
