# Hugging Face — Accelerating Qwen3-8B Agent on Intel® Core™ Ultra with Depth-Pruned Draft Models

- Company: Hugging Face (huggingface.co)
- Announced: 2025-09-29T00:00:00+00:00
- Category: model-update
- Subject: Platform
- Models affected: Qwen3-8B, Qwen3-0.6B
- Source: https://huggingface.co/blog/intel-qwen3-agent
- Record: https://forck.live/items/1611-accelerating-qwen3-8b-agent-on-intel-core-ultra-with-depth-pruned-draft-models

Intel and Hugging Face demonstrate accelerating Qwen3-8B agent inference on Intel Core Ultra using speculative decoding with a depth-pruned draft model (Qwen3-0.6B), achieving ~1.4× speedup over baseline.

## Evidence

Verbatim from https://huggingface.co/blog/intel-qwen3-agent:

> By using speculative decoding and applying a simple pruning process to the draft, we pushed the speedup even further to ~1.4×

---

Record: https://forck.live/items/1611-accelerating-qwen3-8b-agent-on-intel-core-ultra-with-depth-pruned-draft-models
Catalogue: https://forck.live/llms.txt
Feed: https://forck.live/feed.md
