Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
Hugging Face and Intel present optimizations for StarCoder-15B on 4th gen Xeon, achieving over 7x inference acceleration through 8-bit and 4-bit quantization combined with assisted generation (speculative decoding). The work uses SmoothQuant for INT8 quantization and demonstrates significant speedup without accuracy loss on the HumanEval benchmark.
From the source
we show more than 7x inference acceleration of StarCoder-15B model on Intel 4th generation Xeon by integrating 8bit and 4bit quantization with assisted generation.
huggingface.co