Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
Hugging Face and Intel labs announce dynamic speculative decoding, a new method for accelerating text generation up to 2.7x, which becomes the default in Transformers version 4.45.0. The method adjusts the number of draft tokens based on confidence, outperforming the heuristic approach.
From the source
dynamic speculative decoding—a novel method developed by Intel labs and Hugging Face that accelerates text generation by up to 2.7x, depending on the task. This method is the default operational mode for assisted generation starting from Transformers🤗 release 4.45.0
huggingface.co