Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
Intel and Hugging Face demonstrate accelerating Qwen3-8B agent inference on Intel Core Ultra using speculative decoding with a depth-pruned draft model (Qwen3-0.6B), achieving ~1.4× speedup over baseline.
From the source
By using speculative decoding and applying a simple pruning process to the draft, we pushed the speedup even further to ~1.4×
huggingface.co