DeepSeek R1: Open Weights, Distilled Models & Benchmarks

DeepSeek released R1 with downloadable weights and six distilled models in January 2025. See its reinforcement-learning method, reported o1 comparisons, and limitations.

What developers received

DeepSeek released R1 on January 20, 2025, with downloadable model weights, a chat service, and API access. Its announcement offered MIT-licensed code and weights and explicitly allowed distillation: using a model's outputs to train another model. The paper's first arXiv submission followed on January 22. [1][5]

The main model was large: 671 billion parameters, with 37 billion active per token. It used DeepSeek V3's base model. The six smaller releases used Qwen and Llama models ranging from 1.5 billion to 70 billion parameters. They were separate derivatives, not different download sizes of the same R1 weights. [4]

Learning to work through a problem

DeepSeek first tested R1-Zero by applying reinforcement learning to a pretrained model without an initial supervised fine-tuning stage. Correct math answers and passing code tests supplied rewards. The model developed longer problem-solving responses, but repetition and mixed languages made its output hard to use. [5]

R1 added example-based training before reinforcement learning and further training stages afterward. The smaller derivatives learned from generated examples rather than repeating the full training process. The distinction matters: “learned through reinforcement learning” does not mean R1 was trained from scratch without examples or prior language-model training. [5]

What “near o1” meant

DeepSeek reported 79.8% for R1 versus 79.2% for o1-1217 on AIME 2024 competition mathematics, but 71.5% versus 75.7% on GPQA Diamond science questions. These were developer-reported results, not an independent verdict. For sampled evaluations, DeepSeek generated 64 answers per question to estimate single-answer success, with up to 32,768 generated tokens. That is different from counting a problem solved whenever any one of 64 attempts succeeds. [4]

The original report also identified weaknesses in function calling, structured JSON output, and multi-turn interaction relative to V3. A strong math result did not establish that R1 was the better model for every application. [5]

Why the release mattered

Atlas interpretation: OpenAI's o1 preview had already made spending more computation before answering a visible product direction. R1 added downloadable weights and a published training account to that comparison. Researchers could examine a working model and test adaptations, rather than study only the responses of a remote service. Access to weights still did not make the largest model cheap to operate or reproduce its training. [6][1][4]

The release also illustrates Hugging Face's distribution role: DeepSeek linked its downloadable models there. The host and the model's creator are different organizations. [4]

Sources

  1. DeepSeek-R1 Release

    DeepSeek · Jan 20, 2025

  2. DeepSeek-R1 Release

    DeepSeek · Jan 20, 2025

  3. DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

    arXiv · Jan 22, 2025

  4. DeepSeek-R1

    DeepSeek · Jan 21, 2025

  5. DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

    DeepSeek-AI · Jan 22, 2025

  6. Introducing OpenAI o1-preview

    OpenAI · Sep 12, 2024