Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
The transformers vLLM backend now achieves native or better throughput compared to custom vLLM implementations for many LLM architectures, using torch.fx and AST to apply runtime fusions.
From the source
The transformers vLLM backend is now as fast (or faster) than custom vLLM implementations for many LLM architectures.
huggingface.co