From the source
Hugging Face announces Optimum-NVIDIA, an inference library that accelerates LLM inference on NVIDIA GPUs with a simple API change.
It claims up to 28x faster inference and 1,200 tokens/second, supports FP8, and provides code examples for LLaMA models.




