A 671 billion parameter model that mostly stays out of its own way
DeepSeek published V3's weights and a technical report on December 26, 2024. It is a Mixture-of-Experts model with 671 billion total parameters, of which only 37 billion activate for any given token. Context length is 128,000 tokens. [1][2]
The model carries forward Multi-head Latent Attention and DeepSeekMoE from DeepSeek-V2, adds an auxiliary-loss-free method for balancing load across experts, and trains with a multi-token prediction objective that also speeds up inference through speculative decoding. Pretraining ran on 14.8 trillion tokens using an FP8 mixed-precision framework, and the report states the run had no irrecoverable loss spikes and needed no rollbacks. [2]
The $5.576 million figure, and what it leaves out
DeepSeek's report puts total pretraining at 2.664 million H800 GPU hours, plus 0.124 million for context extension and post-training, for 2.788 million GPU hours overall. Priced at an assumed rental rate of $2 per H800 GPU hour, that comes to $5.576 million. [2]
Atlas interpretation: DeepSeek's own report flags the limit of that number: it covers only the official training run, not the prior research and ablation work on architecture, algorithms, and data that made the run possible. The figure is a real accounting of one line item, not a claim that a comparable model can be produced from scratch for that price. [2]
Released on Boxing Day, noticed a few weeks later
DeepSeek reported that V3 beat Meta's Llama 3.1 405B, OpenAI's GPT-4o, and Alibaba's Qwen 2.5 72B on several of its own benchmark comparisons, including coding evaluations. Coverage at the time noted the model performing strongly on the Aider Polyglot coding benchmark and cited engineer Andrej Karpathy's comment that it reached frontier-level results on what he called a modest compute budget. [4]
Atlas interpretation: The release landed during the Christmas week news lull and drew little attention outside people already tracking open-weight models. The scrutiny and the market reaction associated with DeepSeek arrived with R1, the reasoning-focused model the company released about a month later, built on this same V3 base. [4][2]
Open weights, not a permissive license
The repository code is MIT licensed, but the model weights themselves ship under DeepSeek's own model license, which permits commercial use rather than granting the fewer-restrictions terms of a license like MIT or Apache 2.0. DeepSeek moved to a plain MIT license for the model itself starting with R1. [3]
Sources
- Introducing DeepSeek-V3
DeepSeek · Dec 26, 2024
- DeepSeek-V3 Technical Report
arXiv · Dec 27, 2024
- deepseek-ai/DeepSeek-V3
GitHub · Dec 26, 2024
- DeepSeek's new AI model appears to be one of the best 'open' challengers yet
TechCrunch · Dec 26, 2024