# Hugging Face — Speculative Decoding for 2x Faster Whisper Inference

- Company: Hugging Face (huggingface.co)
- Announced: 2023-12-20
- Category: capability-change
- Coverage: not counted
- Announcement: yes
- Group: announcements
- Source: https://huggingface.co/blog/whisper-speculative-decoding
- Record: https://forck.live/items/1964-speculative-decoding-for-2x-faster-whisper-inference
- Subject: Platform
- Models affected: Whisper, large-v3, large-v2, tiny, tiny.en, medium.en

Hugging Face demonstrates how Speculative Decoding can reduce Whisper inference time by a factor of 2 while ensuring exactly the same outputs, making it a drop-in replacement for existing Whisper pipelines.

## Evidence

Verbatim from https://huggingface.co/blog/whisper-speculative-decoding:

> we demonstrate how Speculative Decoding can be employed to reduce the inference time of Whisper by a factor of 2, while mathematically ensuring exactly the same outputs are achieved from the model.

---

Record: https://forck.live/items/1964-speculative-decoding-for-2x-faster-whisper-inference
Catalogue: https://forck.live/llms.txt
Current issue: https://forck.live/feed.md
