# Amazon — Tiered KV cache for large LLMs on Amazon SageMaker HyperPod with Curvine

- Company: Amazon (amazon.com)
- Announced: 2026-08-12T13:42:48+00:00
- Category: product-launch
- Subject: Bedrock / Nova
- Models affected: Qwen, Llama, DeepSeek
- Source: https://aws.amazon.com/blogs/machine-learning/tiered-kv-cache-for-large-llms-on-amazon-sagemaker-hyperpod-with-curvine/
- Record: https://forck.live/items/3755-tiered-kv-cache-for-large-llms-on-amazon-sagemaker-hyperpod-with-curvine

Amazon SageMaker HyperPod introduces a tiered KV cache architecture using Curvine, a distributed cache filesystem, to extend KV cache across GPU, CPU, and shared NVMe, enabling cross-replica cache reuse and reducing time-to-first-token.

## Evidence

Verbatim from https://aws.amazon.com/blogs/machine-learning/tiered-kv-cache-for-large-llms-on-amazon-sagemaker-hyperpod-with-curvine/:

> On a test deployment, this achieved up to a 100 percent cross-Pod cache hit rate, up to a 2.7x TTFT improvement, and cross-node L2 read latency of about 56 ms for a approximately 1,900-token prompt.

---

Record: https://forck.live/items/3755-tiered-kv-cache-for-large-llms-on-amazon-sagemaker-hyperpod-with-curvine
Catalogue: https://forck.live/llms.txt
Feed: https://forck.live/feed.md
