What developers received
DeepSeek released R1 on January 20, 2025, with downloadable model weights, a chat service, and API access. Its announcement offered MIT-licensed code and weights and explicitly allowed distillation: using a model's outputs to train another model. The paper's first arXiv submission followed on January 22. [1][5]
The main model was large: 671 billion parameters, with 37 billion active per token. It used DeepSeek V3's base model. The six smaller releases used Qwen and Llama models ranging from 1.5 billion to 70 billion parameters. They were separate derivatives, not different download sizes of the same R1 weights. [4]
Learning to work through a problem
DeepSeek first tested R1-Zero by applying reinforcement learning to a pretrained model without an initial supervised fine-tuning stage. Correct math answers and passing code tests supplied rewards. The model developed longer problem-solving responses, but repetition and mixed languages made its output hard to use. [5]
R1 added example-based training before reinforcement learning and further training stages afterward. The smaller derivatives learned from generated examples rather than repeating the full training process. The distinction matters: “learned through reinforcement learning” does not mean R1 was trained from scratch without examples or prior language-model training. [5]
What “near o1” meant
DeepSeek reported 79.8% for R1 versus 79.2% for o1-1217 on AIME 2024 competition mathematics, but 71.5% versus 75.7% on GPQA Diamond science questions. These were developer-reported results, not an independent verdict. For sampled evaluations, DeepSeek generated 64 answers per question to estimate single-answer success, with up to 32,768 generated tokens. That is different from counting a problem solved whenever any one of 64 attempts succeeds. [4]
The original report also identified weaknesses in function calling, structured JSON output, and multi-turn interaction relative to V3. A strong math result did not establish that R1 was the better model for every application. [5]
Why the release mattered
Atlas interpretation: OpenAI's o1 preview had already made spending more computation before answering a visible product direction. R1 added downloadable weights and a published training account to that comparison. Researchers could examine a working model and test adaptations, rather than study only the responses of a remote service. Access to weights still did not make the largest model cheap to operate or reproduce its training. [6][1][4]
The release also illustrates Hugging Face's distribution role: DeepSeek linked its downloadable models there. The host and the model's creator are different organizations. [4]
Sources
- DeepSeek-R1 Release
DeepSeek · Jan 20, 2025
- DeepSeek-R1 Release
DeepSeek · Jan 20, 2025
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
arXiv · Jan 22, 2025
- DeepSeek-R1
DeepSeek · Jan 21, 2025
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
DeepSeek-AI · Jan 22, 2025
- Introducing OpenAI o1-preview
OpenAI · Sep 12, 2024