# Amazon — Deploying Qwen3.8-2.4T-A95B on Amazon SageMaker HyperPod with vLLM

- Company: Amazon (amazon.com)
- Announced: 2026-09-09T22:26:29+00:00
- Category: open-weight-release
- Coverage: not counted
- Announcement: no
- Group: routine
- Source: https://aws.amazon.com/blogs/machine-learning/deploying-qwen3-8-2-4t-a95b-on-amazon-sagemaker-hyperpod-with-vllm/
- Record: https://forck.live/items/9623-deploying-qwen3-8-2-4t-a95b-on-amazon-sagemaker-hyperpod-with-vllm
- Subject: Bedrock / Nova
- Open weights: yes
- Models affected: Qwen3.8-2.4T-A95B
- Context window: 262K tokens (extensible to 1M)

Alibaba's Qwen team released Qwen3.8-2.4T-A95B, the first open-weight Qwen-Max-class model with 2.4 trillion total parameters and 95 billion activated per token. The model features a hybrid linear-plus-full-attention architecture, native context up to 262K tokens extensible to 1M, and is designed for agentic and reasoning workloads. Amazon published a guide on deploying this model on SageMaker HyperPod using vLLM on ml.p6-b300 instances.

## Evidence

Verbatim from https://aws.amazon.com/blogs/machine-learning/deploying-qwen3-8-2-4t-a95b-on-amazon-sagemaker-hyperpod-with-vllm/:

> On August 12, 2026, Alibaba's Qwen team released Qwen3.8-2.4T-A95B . This is the first time a Qwen-Max-class model has been made available as open weights.

---

Record: https://forck.live/items/9623-deploying-qwen3-8-2-4t-a95b-on-amazon-sagemaker-hyperpod-with-vllm
Catalogue: https://forck.live/llms.txt
Current issue: https://forck.live/feed.md
