Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
This is a blog post that explains continuous batching for LLM inference, starting from attention mechanisms and KV caching, and deriving the technique by optimizing for throughput.
From the source
TL;DR: in this blog post, starting from attention mechanisms and KV caching, we derive continuous batching by optimizing for throughput.
huggingface.co