From the source
Lead story
Ask your AI
Top stories
Models & availability
Latest
Lead story
Top stories
Models & availability
Latest
From the source
Amazon SageMaker Inference introduced prefix-aware routing, a new routing strategy that sends requests sharing the same prompt prefix to the same instance to improve KV cache reuse.
In benchmarks on Llama 3.1 70B, it reduced P50 TTFT by up to 77% and increased throughput by up to 16%.
From the source
Today, Amazon SageMaker Inference introduces prefix-aware routing. It is a new routing strategy that looks at the beginning of each request and consistently sends requests with the same beginning to the same instance.
aws.amazon.com