From the source
Lead story
Ask your AI
Top stories
Models & availability
Latest
Lead story
Top stories
Models & availability
Latest
From the source
Amazon published a guide demonstrating how to use concurrency sweeps in Amazon SageMaker AI Inference Recommendations to right-size generative AI endpoints.
The walkthrough covers deploying the NVIDIA Nemotron-3 Nano 30B model, configuring workload profiles, running automated benchmarks via the CreateAIBenchmarkJob API, and using results to optimize instance selection and capacity planning.
From the source
Concurrency sweeps are built into Amazon SageMaker AI Inference Recommendations, so there's no custom load-testing infrastructure to build or maintain.
aws.amazon.com