A trillion parameters, 32 billion of them awake at once
Moonshot AI released Kimi K2 Thinking on November 6, 2025, with open weights published on Hugging Face under a modified MIT license. The model is a mixture-of-experts design with 1 trillion total parameters and 32 billion active per token, spread across 384 experts with 8 selected per token plus one shared expert, and a 256K token context window. [4]
The release used native INT4 quantization applied through quantization aware training in post-training, which Moonshot says roughly doubles generation speed in low-latency mode without a corresponding accuracy loss. Moonshot describes the model as a thinking agent built to interleave chain-of-thought reasoning with function calls across long sequences of tool calls, rather than a chat model that occasionally calls a tool. [4]
Strong on tool-heavy benchmarks, unremarkable on plain reasoning
On Moonshot's own published tables, Kimi K2 Thinking scored 44.9 on Humanity's Last Exam with tools and 60.2 on BrowseComp with tools, both ahead of the GPT-5 (high) and Claude Sonnet 4.5 (thinking) figures Moonshot lists for comparison. On tool-free reasoning benchmarks such as AIME25 and GPQA, its scores were close to but not above those same comparison models, and Moonshot credits the gap between tool-free and tool-assisted scores to the model's tolerance for hundreds of consecutive tool calls without losing track of the goal. [4]
Atlas interpretation: Those figures are Moonshot's own benchmark run against models it chose to compare against, not an independent audit, so the size of Kimi K2 Thinking's lead on agentic search tasks should be read as a vendor's best case rather than a settled ranking. [4]
A government evaluation found a narrower lead
On December 12, 2025, the US Center for AI Standards and Innovation (CAISI) published its own evaluation of Kimi K2 Thinking. Across cybersecurity, software engineering, scientific knowledge and mathematical reasoning tasks, CAISI found the model only a modest improvement on DeepSeek V3.1, and its agentic cyber and software engineering performance behind leading US models, including GPT-5 and Claude Opus 4. [2]
CAISI also reported that the model is heavily censored on Chinese-language prompts, at rates comparable to DeepSeek's R1-0528, while showing substantially less censorship in English, Spanish and Arabic. The evaluation noted that despite its capability, Kimi K2 Thinking had not seen wide adoption, reaching only about 10 percent of DeepSeek R1's download volume a month after release. [2]
Atlas interpretation: The two assessments are not really in conflict. Moonshot's benchmarks emphasize tasks built around extended tool use, where Kimi K2 Thinking's tolerance for long call chains stands out. CAISI's broader evaluation, covering tasks less dependent on that specific strength, found it trailing the closed frontier. The model's open-weight standing and its absolute standing against GPT-5 or Claude are separate claims, and only the first one is uncontested. [4][2]
Sources
- Kimi K2 Thinking
Moonshot AI · Nov 6, 2025
- CAISI Evaluation of Kimi K2 Thinking
NIST · Dec 12, 2025
- Alibaba-backed Moonshot releases new AI model Kimi K2 Thinking
CNBC · Nov 6, 2025
- moonshotai/Kimi-K2-Thinking
Hugging Face · Sep 8, 2026