# Alibaba — Qwen2.5-Max: Exploring the Intelligence of Large-scale MoE Model

- Company: Alibaba (alibaba.com)
- Announced: 2025-01-28T15:00:04+00:00
- Subject: Qwen
- Models affected: Qwen2
- Source: https://qwenlm.github.io/blog/qwen2.5-max/
- Record: https://forck.live/items/2296-qwen2-5-max-exploring-the-intelligence-of-large-scale-moe-model

The post discusses the benefits of scaling data and model size, notes limited experience in scaling large models, references DeepSeek V3, and states that Alibaba is developing Qwen2.

## Evidence

Verbatim from https://qwenlm.github.io/blog/qwen2.5-max/:

> It is widely recognized that continuously scaling both data size and model size can lead to significant improvements in model intelligence. However, the research and industry community has limited experience in effectively scaling extremely large models, whether they are dense or Mixture-of-Expert (MoE) models. Many critical details regarding this scaling process were only disclosed with the recent release of DeepSeek V3. Concurrently, we are developing Qwen2.

---

Record: https://forck.live/items/2296-qwen2-5-max-exploring-the-intelligence-of-large-scale-moe-model
Catalogue: https://forck.live/llms.txt
Feed: https://forck.live/feed.md
