# Hugging Face — Accelerate StarCoder with 🤗 Optimum Intel on Xeon: Q8/Q4 and Speculative Decoding

- Company: Hugging Face (huggingface.co)
- Announced: 2024-01-30T00:00:00+00:00
- Category: capability-change
- Subject: Platform
- Models affected: StarCoder, StarCoder-15B
- Source: https://huggingface.co/blog/intel-starcoder-quantization
- Record: https://forck.live/items/1950-accelerate-starcoder-with-optimum-intel-on-xeon-q8-q4-and-speculative-decoding

Hugging Face and Intel present optimizations for StarCoder-15B on 4th gen Xeon, achieving over 7x inference acceleration through 8-bit and 4-bit quantization combined with assisted generation (speculative decoding). The work uses SmoothQuant for INT8 quantization and demonstrates significant speedup without accuracy loss on the HumanEval benchmark.

## Evidence

Verbatim from https://huggingface.co/blog/intel-starcoder-quantization:

> we show more than 7x inference acceleration of StarCoder-15B model on Intel 4th generation Xeon by integrating 8bit and 4bit quantization with assisted generation.

---

Record: https://forck.live/items/1950-accelerate-starcoder-with-optimum-intel-on-xeon-q8-q4-and-speculative-decoding
Catalogue: https://forck.live/llms.txt
Feed: https://forck.live/feed.md
