From the source
Lead story
Ask your AI
Top stories
Models & availability
Latest
Lead story
Top stories
Models & availability
Latest
From the source
Model caching reduces SageMaker HyperPod inference cold starts from 25-30+ minutes to seconds.
Amazon launched model caching for Amazon SageMaker Inference on HyperPod, which pre-loads model weights and container images onto cluster nodes to reduce inference cold starts from tens of minutes to seconds.
From the source
Today we’re launching model caching for Amazon SageMaker Inference on HyperPod. Model caching pre-loads model weights and container images onto cluster nodes before pods need them.
aws.amazon.com