DeepSeek V4 Preview: MIT Open Weights and 1M Context

DeepSeek released V4-Pro and V4-Flash in preview on April 24, 2026, with MIT-licensed weights and a one-million-token context window.

Two models, one new base architecture

DeepSeek released preview versions of two models on April 24, 2026: DeepSeek-V4-Pro, a mixture-of-experts model with 1.6 trillion total parameters and 49 billion active per token, and DeepSeek-V4-Flash, a smaller sibling with 284 billion total parameters and 13 billion active. Both ship as open weights on Hugging Face under the MIT license, both default to a one million token context window, and both run in a thinking or non-thinking mode through the same API. [1][3]

DeepSeek described the release as an upgrade path rather than a parallel option: existing API callers only need to change the model name to deepseek-v4-pro or deepseek-v4-flash, and the older deepseek-chat and deepseek-reasoner endpoints are scheduled for retirement on July 24, 2026. [1]

A new attention mechanism built for the million-token context

The architectural change centers on the attention mechanism. Where V3.2 added DeepSeek Sparse Attention on top of V3's Multi-head Latent Attention, the V4 models replace both with a hybrid scheme the Hugging Face model card describes as Compressed Sparse Attention paired with Heavily Compressed Attention, along with Manifold-Constrained Hyper-Connections strengthening the residual path across layers and a switch to the Muon optimizer for training stability. [3]

DeepSeek's efficiency claim for the new mechanism is specific to long context: at the million-token mark, V4-Pro needs 27 percent of the single-token inference compute and 10 percent of the key-value cache that V3.2 required. V4-Pro trained on more than 32 trillion tokens, with the report describing separate post-training passes for different domains that were later merged into one model. [3]

Back near the front of the open-weight pack, with a caveat

Independent benchmarking from Artificial Analysis put V4-Pro at 52 on its Intelligence Index, the second-highest score among open-weight reasoning models at the time, behind only Moonshot AI's Kimi K2.6. V4-Flash scored 47, in the range of Anthropic's Claude Sonnet 4.6 at maximum reasoning effort. On GDPval-AA, an agentic benchmark, V4-Pro's score of 1554 led every open-weight model tested, ahead of Kimi K2.6's 1484. [2]

Atlas interpretation: The same testing turned up a sharp limitation: both models answer even when they do not know, at a hallucination rate of 94 percent for V4-Pro and 96 percent for V4-Flash on the questions Artificial Analysis used to measure it. A model can score well on reasoning benchmarks and still be an unreliable source of unverified facts, and this release's own numbers illustrate that gap directly rather than as a general caveat about the category. [2]

Priced under the frontier labs, covered as a return to form

API pricing undercuts the closed-source frontier: Artificial Analysis lists V4-Pro at $1.74 per million input tokens and $3.48 per million output tokens, and V4-Flash at $0.14 and $0.28. Both models are text-only, with no image, audio or video input or output, unlike some closed frontier competitors. [2]

TechCrunch covered the release as DeepSeek narrowing, though not closing, the gap with frontier labs, citing DeepSeek's own comparisons showing V4-Pro ahead of GPT-5.2 and Gemini 3.0 Pro on some reasoning tasks, roughly level with GPT-5.4 on coding competitions, and trailing GPT-5.4 and Gemini 3.1 Pro on knowledge tests by what the outlet described as three to six months of development. [4]

Atlas interpretation: This is DeepSeek's first release to draw sustained comparison to the open-weight leaders again since R1 set off a market selloff in January 2025. The V3.2 release in between introduced the sparse-attention line this architecture builds on but did not reset where DeepSeek ranked against competing open-weight labs the way V4 did. [4]

Sources

  1. DeepSeek V4 Preview Release

    DeepSeek · Apr 24, 2026

  2. DeepSeek is back among the leading open weights models with V4 Pro and V4 Flash

    Artificial Analysis · Apr 24, 2026

  3. deepseek-ai/DeepSeek-V4-Pro

    Hugging Face · Apr 24, 2026

  4. DeepSeek previews new AI model that 'closes the gap' with frontier models

    TechCrunch · Apr 24, 2026