# Amazon — Right-size generative AI endpoints with concurrency sweeps on Amazon SageMaker AI

- Company: Amazon (amazon.com)
- Announced: 2026-09-22T15:35:53+00:00
- Category: developer-tool-release
- Coverage: not counted
- Announcement: yes
- Group: announcements
- Source: https://aws.amazon.com/blogs/machine-learning/right-size-generative-ai-endpoints-with-concurrency-sweeps-on-amazon-sagemaker-ai/
- Record: https://forck.live/items/13081-right-size-generative-ai-endpoints-with-concurrency-sweeps-on-amazon-sagemaker
- Subject: Bedrock / Nova
- Models affected: NVIDIA Nemotron-3 Nano 30B

Amazon published a guide demonstrating how to use concurrency sweeps in Amazon SageMaker AI Inference Recommendations to right-size generative AI endpoints. The walkthrough covers deploying the NVIDIA Nemotron-3 Nano 30B model, configuring workload profiles, running automated benchmarks via the CreateAIBenchmarkJob API, and using results to optimize instance selection and capacity planning.

## Evidence

Verbatim from https://aws.amazon.com/blogs/machine-learning/right-size-generative-ai-endpoints-with-concurrency-sweeps-on-amazon-sagemaker-ai/:

> Concurrency sweeps are built into Amazon SageMaker AI Inference Recommendations, so there's no custom load-testing infrastructure to build or maintain.

---

Record: https://forck.live/items/13081-right-size-generative-ai-endpoints-with-concurrency-sweeps-on-amazon-sagemaker
Catalogue: https://forck.live/llms.txt
Current issue: https://forck.live/feed.md
