From the source
Hugging Face demonstrates how Speculative Decoding can reduce Whisper inference time by a factor of 2 while ensuring exactly the same outputs, making it a drop-in replacement for existing Whisper pipelines.
From the source
From the source

Hugging Face demonstrates how Speculative Decoding can reduce Whisper inference time by a factor of 2 while ensuring exactly the same outputs, making it a drop-in replacement for existing Whisper pipelines.
From the source
we demonstrate how Speculative Decoding can be employed to reduce the inference time of Whisper by a factor of 2, while mathematically ensuring exactly the same outputs are achieved from the model.
huggingface.co