From the source
Lead story
Ask your AI
Top stories
Models & availability
Latest
Lead story
Top stories
Models & availability
Latest
From the source
G7 GPU instances deliver 60.8% higher throughput and 37.6% lower latency than G6
Amazon announced benchmarking of small LLM inference on SageMaker AI, comparing G7 instances (NVIDIA Blackwell GPUs) against G5 and G6 instances for two 30B MoE models: Qwen3-Coder-30B and NVIDIA Nemotron-3-Nano-30B-A3B-NVFP4.
The benchmarks demonstrate that G7 instances deliver measurable gains in throughput, latency, and cost-per-token.
From the source
In this post, we benchmark two representative 30B Mixture-of-Experts (MoE) models across three GPU instance families on Amazon SageMaker AI Inference using Amazon SageMaker AI Generative AI inference recommendation.
aws.amazon.com