# Amazon — Introducing Amazon SageMaker HyperPod Inference Gateway

- Company: Amazon (amazon.com)
- Announced: 2026-09-18T13:08:34+00:00
- Category: developer-tool-release
- Coverage: 1 outlet
- Announcement: yes
- Group: covered
- Source: https://aws.amazon.com/blogs/machine-learning/introducing-amazon-sagemaker-hyperpod-inference-gateway/
- Record: https://forck.live/items/12030-introducing-amazon-sagemaker-hyperpod-inference-gateway
- Subject: Bedrock / Nova

Amazon announced SageMaker HyperPod Inference Gateway, a Kubernetes-native GPU-aware routing system that deploys as an EKS managed addon to optimize inference request placement on GPU clusters. The system reduces first-token latency by up to 82% through real-time GPU metrics and intelligent routing algorithms that consider KV cache utilization, queue depth, LoRA adapter residency, and prefix cache hit rates, with no changes required to model servers or client applications.

## Evidence

Verbatim from https://aws.amazon.com/blogs/machine-learning/introducing-amazon-sagemaker-hyperpod-inference-gateway/:

> It is a Kubernetes-native, GPU-aware routing system that deploys as a single EKS managed addon on your existing HyperPod infrastructure. It uses real-time GPU signals to place every inference request on the best-suited pod, delivering lower latency with no changes to your model servers or client applications.

## Around this story

Outlets this record can name and link:

- Unite — AWS Launches SageMaker HyperPod Inference Gateway for GPU-Aware Routing: https://unite.ai/aws-launches-sagemaker-hyperpod-inference-gateway-for-gpu-aware-routing

---

Record: https://forck.live/items/12030-introducing-amazon-sagemaker-hyperpod-inference-gateway
Catalogue: https://forck.live/llms.txt
Current issue: https://forck.live/feed.md
