# Amazon — Benchmarking small LLM inference on SageMaker AI: G7 vs G5 and G6

- Company: Amazon (amazon.com)
- Announced: 2026-09-08T16:21:57+00:00
- Category: capability-change
- Coverage: not counted
- Announcement: no
- Group: routine
- Source: https://aws.amazon.com/blogs/machine-learning/benchmarking-small-llm-inference-on-sagemaker-ai-g7-vs-g5-and-g6/
- Record: https://forck.live/items/9280-benchmarking-small-llm-inference-on-sagemaker-ai-g7-vs-g5-and-g6
- Subject: Bedrock / Nova
- Models affected: Qwen3-Coder-30B, NVIDIA Nemotron-3-Nano-30B-A3B-NVFP4

Amazon announced benchmarking of small LLM inference on SageMaker AI, comparing G7 instances (NVIDIA Blackwell GPUs) against G5 and G6 instances for two 30B MoE models: Qwen3-Coder-30B and NVIDIA Nemotron-3-Nano-30B-A3B-NVFP4. The benchmarks demonstrate that G7 instances deliver measurable gains in throughput, latency, and cost-per-token.

## Evidence

Verbatim from https://aws.amazon.com/blogs/machine-learning/benchmarking-small-llm-inference-on-sagemaker-ai-g7-vs-g5-and-g6/:

> In this post, we benchmark two representative 30B Mixture-of-Experts (MoE) models across three GPU instance families on Amazon SageMaker AI Inference using Amazon SageMaker AI Generative AI inference recommendation.

---

Record: https://forck.live/items/9280-benchmarking-small-llm-inference-on-sagemaker-ai-g7-vs-g5-and-g6
Catalogue: https://forck.live/llms.txt
Current issue: https://forck.live/feed.md
