From the source
Intel and Hugging Face show that SmoothQuant quantization reduces model size by ~2x and enables efficient LLM inference on Intel CPUs, as demonstrated with models like OPT, LLaMA, Alpaca, Vicuna, BloomZ, and MPT-7B-chat.
From the source
From the source

Intel and Hugging Face show that SmoothQuant quantization reduces model size by ~2x and enables efficient LLM inference on Intel CPUs, as demonstrated with models like OPT, LLaMA, Alpaca, Vicuna, BloomZ, and MPT-7B-chat.
From the source
As a consequence, SmoothQuant produces smaller, faster models that run well on Intel CPU platforms.
huggingface.co