# Amazon — Reduce inference cold starts on Amazon SageMaker HyperPod with model caching

- Company: Amazon (amazon.com)
- Announced: 2026-09-10T21:37:49+00:00
- Category: capability-change
- Coverage: not counted
- Announcement: yes
- Group: announcements
- Source: https://aws.amazon.com/blogs/machine-learning/reduce-inference-cold-starts-on-amazon-sagemaker-hyperpod-with-model-caching/
- Record: https://forck.live/items/9970-reduce-inference-cold-starts-on-amazon-sagemaker-hyperpod-with-model-caching
- Subject: Bedrock / Nova

Amazon launched model caching for Amazon SageMaker Inference on HyperPod, which pre-loads model weights and container images onto cluster nodes to reduce inference cold starts from tens of minutes to seconds.

## Evidence

Verbatim from https://aws.amazon.com/blogs/machine-learning/reduce-inference-cold-starts-on-amazon-sagemaker-hyperpod-with-model-caching/:

> Today we’re launching model caching for Amazon SageMaker Inference on HyperPod. Model caching pre-loads model weights and container images onto cluster nodes before pods need them.

---

Record: https://forck.live/items/9970-reduce-inference-cold-starts-on-amazon-sagemaker-hyperpod-with-model-caching
Catalogue: https://forck.live/llms.txt
Current issue: https://forck.live/feed.md
