# Hugging Face — Smaller is better: Q8-Chat, an efficient generative AI experience on Xeon

- Company: Hugging Face (huggingface.co)
- Announced: 2023-05-16
- Category: not stated
- Coverage: not counted
- Announcement: no
- Group: routine
- Source: https://huggingface.co/blog/generative-ai-models-on-intel-cpu
- Record: https://forck.live/items/2060-smaller-is-better-q8-chat-an-efficient-generative-ai-experience-on-xeon
- Subject: Platform
- Models affected: OPT 2.7B, OPT 6.7B, LLaMA 7B, Alpaca 7B, Vicuna 7B, BloomZ 7.1B, MPT-7B-chat

Intel and Hugging Face show that SmoothQuant quantization reduces model size by ~2x and enables efficient LLM inference on Intel CPUs, as demonstrated with models like OPT, LLaMA, Alpaca, Vicuna, BloomZ, and MPT-7B-chat.

## Evidence

Verbatim from https://huggingface.co/blog/generative-ai-models-on-intel-cpu:

> As a consequence, SmoothQuant produces smaller, faster models that run well on Intel CPU platforms.

---

Record: https://forck.live/items/2060-smaller-is-better-q8-chat-an-efficient-generative-ai-experience-on-xeon
Catalogue: https://forck.live/llms.txt
Current issue: https://forck.live/feed.md
