Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
DeepSeek released two new MoE models, DeepSeek-V4-Pro (1.6T parameters, 49B active) and DeepSeek-V4-Flash (284B parameters, 13B active), both with a 1M-token context window. The models are designed for efficient long-context inference and agentic tasks, featuring hybrid attention (CSA and HCA) that reduces FLOPs and KV cache memory, plus post-training improvements for multi-turn tool use.
From the source
DeepSeek released V4 today. Two MoE checkpoints are on the Hub: DeepSeek-V4-Pro at 1.6T total parameters with 49B active, and DeepSeek-V4-Flash at 284B total with 13B active. Both have a 1M-token context window.
huggingface.co