Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
The article discusses the impact of concurrent request processing on LLM latency, throughput, and GPU utilization, focusing on the prefill and decode phases.
From the source
In this second part, we will now focus on the concurrent processing of requests, and how it impacts relevant metrics such as latency and throughput as well as GPU resource utilization.
huggingface.co