Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
Hugging Face announces native integration of Intel Gaudi hardware support into Text Generation Inference (TGI), enabling deployment of LLMs on Gaudi accelerators with features like multi-card inference, vision-language models, and FP8 precision, and supporting models such as Llama 3.1, Mixtral, Mistral, and more.
From the source
We're excited to announce the native integration of Intel Gaudi hardware support directly into Text Generation Inference (TGI), our production-ready serving solution for Large Language Models (LLMs).
huggingface.co