Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
This blog post provides a step-by-step guide on optimizing and deploying Hugging Face Transformers models using Optimum-Intel and OpenVINO GenAI, focusing on edge and client-side deployment with C++ and Python. It covers environment setup, exporting models to OpenVINO IR, weight-only quantization (including INT4/INT8 with AWQ and scale estimation), and deployment using the OpenVINO GenAI API.
From the source
This blog will guide you through optimizing and deploying Hugging Face Transformers models using Optimum-Intel and OpenVINO™ GenAI, ensuring efficient AI inference with minimal dependencies.
huggingface.co