Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
Hugging Face published a blog post explaining asynchronous batching for continuous batching in LLM inference, which separates CPU and GPU workloads to improve GPU utilization, and has been implemented in the transformers library.
From the source
we explain how to separate CPU and GPU workloads to get a massive performance boost for inference.
huggingface.co