Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
Hugging Face and Intel have adapted and optimized assisted decoding (speculative sampling) for Intel Gaudi processors, now integrated into Optimum Habana to accelerate text generation.
From the source
We adapted and optimized it for Intel Gaudi, which delivers similar performance as Nvidia H100 GPUs as shown in a previous post, while its price is in the same ballpark as Nvidia A100 80GB GPUs. This work is now part of Optimum Habana, which extends various Hugging Face libraries like Transformers and Diffusers so that your AI workflows are fully optimized for Intel Gaudi processors.
huggingface.co