# Amazon — Reduce LLM latency with prefix-aware routing on Amazon SageMaker Inference

- Company: Amazon (amazon.com)
- Announced: 2026-09-10T21:58:09+00:00
- Category: capability-change
- Coverage: not counted
- Announcement: yes
- Group: announcements
- Source: https://aws.amazon.com/blogs/machine-learning/reduce-llm-latency-with-prefix-aware-routing-on-amazon-sagemaker-inference/
- Record: https://forck.live/items/9973-reduce-llm-latency-with-prefix-aware-routing-on-amazon-sagemaker-inference
- Subject: Bedrock / Nova
- Models affected: Llama 3.1 70B

Amazon SageMaker Inference introduced prefix-aware routing, a new routing strategy that sends requests sharing the same prompt prefix to the same instance to improve KV cache reuse. In benchmarks on Llama 3.1 70B, it reduced P50 TTFT by up to 77% and increased throughput by up to 16%.

## Evidence

Verbatim from https://aws.amazon.com/blogs/machine-learning/reduce-llm-latency-with-prefix-aware-routing-on-amazon-sagemaker-inference/:

> Today, Amazon SageMaker Inference introduces prefix-aware routing. It is a new routing strategy that looks at the beginning of each request and consistently sends requests with the same beginning to the same instance.

---

Record: https://forck.live/items/9973-reduce-llm-latency-with-prefix-aware-routing-on-amazon-sagemaker-inference
Catalogue: https://forck.live/llms.txt
Current issue: https://forck.live/feed.md
