Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
This blog post explains how to run the Phi-2 small language model locally on a laptop with an Intel Meteor Lake CPU by applying 4-bit quantization using Intel OpenVINO and Optimum Intel, enabling local inference with benefits like privacy, lower latency, offline work, and cost savings.
From the source
In this post, we'll leverage all of the above. Starting from the Microsoft Phi-2 model, we will apply 4-bit quantization on the model weights, thanks to the Intel OpenVINO integration in our Optimum Intel library. Then, we will run inference on a mid-range laptop powered by an Intel Meteor Lake CPU.
huggingface.co