# Hugging Face — DeepSeek-V4: a million-token context that agents can actually use

- Company: Hugging Face (huggingface.co)
- Announced: 2026-04-24T00:00:00+00:00
- Category: new-model
- Subject: Platform
- Models affected: DeepSeek-V4-Pro, DeepSeek-V4-Flash
- Context window: 1M-token context window
- Source: https://huggingface.co/blog/deepseekv4
- Record: https://forck.live/items/1512-deepseek-v4-a-million-token-context-that-agents-can-actually-use

DeepSeek released two new MoE models, DeepSeek-V4-Pro (1.6T parameters, 49B active) and DeepSeek-V4-Flash (284B parameters, 13B active), both with a 1M-token context window. The models are designed for efficient long-context inference and agentic tasks, featuring hybrid attention (CSA and HCA) that reduces FLOPs and KV cache memory, plus post-training improvements for multi-turn tool use.

## Evidence

Verbatim from https://huggingface.co/blog/deepseekv4:

> DeepSeek released V4 today. Two MoE checkpoints are on the Hub: DeepSeek-V4-Pro at 1.6T total parameters with 49B active, and DeepSeek-V4-Flash at 284B total with 13B active. Both have a 1M-token context window.

---

Record: https://forck.live/items/1512-deepseek-v4-a-million-token-context-that-agents-can-actually-use
Catalogue: https://forck.live/llms.txt
Feed: https://forck.live/feed.md
