# Hugging Face — Introducing multi-backends (TRT-LLM, vLLM) support for Text Generation Inference

- Company: Hugging Face (huggingface.co)
- Announced: 2025-01-16T00:00:00+00:00
- Category: capability-change
- Subject: Platform
- Source: https://huggingface.co/blog/tgi-multi-backend
- Record: https://forck.live/items/1769-introducing-multi-backends-trt-llm-vllm-support-for-text-generation-inference

Hugging Face announces multi-backend support for Text Generation Inference (TGI), allowing integration with vLLM, TensorRT-LLM, llama.cpp, AWS Neuron, and Google TPU backends through a unified frontend layer.

## Evidence

Verbatim from https://huggingface.co/blog/tgi-multi-backend:

> we are excited to introduce the concept of TGI Backends. This new architecture gives the flexibility to integrate with any of the solutions above through TGI as a single unified frontend layer.

---

Record: https://forck.live/items/1769-introducing-multi-backends-trt-llm-vllm-support-for-text-generation-inference
Catalogue: https://forck.live/llms.txt
Feed: https://forck.live/feed.md
