A trillion parameters, thirty-two billion at a time
Moonshot AI is a Beijing lab founded in March 2023 by Yang Zhilin, a Tsinghua University computer science graduate who earned a machine learning PhD at Carnegie Mellon in 2019. Alibaba led a funding round of more than a billion dollars in February 2024 that valued the company at $2.5 billion, and a Series B that August pushed the figure to roughly $3 billion. Moonshot had already built an audience with the Kimi Chat assistant before K2, but K2 was its first release aimed squarely at developers building on open weights rather than at consumer chat. [7][1]
K2 is a mixture-of-experts model: one trillion parameters exist in the checkpoint, but each token only routes through a subset of them, 32 billion worth. Moonshot's published specification lists 61 layers (60 of them mixture-of-experts, one dense), 384 experts with 8 selected per token plus a shared expert, 64 attention heads, and a 160,000-token vocabulary. The practical effect of the mixture-of-experts design is that inference and fine-tuning cost track the 32 billion active figure much more closely than the trillion-parameter total, which is what makes a model this large runnable at all outside a big cloud provider's cluster. [1][3]
15.5 trillion tokens without a spike
Moonshot's technical report says K2 was pretrained on 15.5 trillion tokens using an optimizer the team calls MuonClip, an extension of an existing method called Muon, and reports that the run completed with what it describes as zero loss spikes across the entire pretraining schedule. That claim comes from Moonshot's own paper; no outside lab has reproduced the training run to check it, since reproducing a trillion-parameter pretraining run is not something an outside lab can casually do. [3]
Atlas interpretation: "Training instability" at this scale usually means the internal attention scores a model uses to weigh which earlier tokens matter, its attention logits, grow without bound partway through a run. Left unchecked, that can silently corrupt the model's weights and force engineers to roll back to an earlier checkpoint and restart, burning the compute spent since the last save. Moonshot's fix, which it calls QK-clip, rescales the query and key projection matrices that produce those attention scores so the logits stay bounded, without changing what the rest of Muon's update rule does for training efficiency. Whether that specific technique generalizes beyond Moonshot's own runs is untested outside the company; what is verifiable is that a model this size shipped on schedule with a coherent, non-degenerate result, which is itself evidence the approach worked well enough for one lab, once. [3]
Open weights, not open source
Moonshot published K2's weights under what it calls a Modified MIT License: the standard MIT permissions to use, copy, modify and sell, plus one added clause. If a product or service built on the software (or a derivative of it) reaches more than 100 million monthly active users, or more than $20 million in monthly revenue, the license requires that product to prominently display "Kimi K2" on its user interface. Below those thresholds the license imposes no extra obligation beyond MIT's usual copyright notice. [4]
Atlas interpretation: That license makes K2 easy to redistribute and easy to build a business on, which is a meaningfully more open position than a closed API. It is still not the same claim as fully open source. Moonshot published the weights, an inference-serving recipe, and a technical report describing the architecture, the optimizer and the broad composition of the pretraining data (web text, code, mathematics and general knowledge, by the paper's own categories). It did not publish the training data itself or the tooling used to filter and mix it. A researcher can run K2, fine-tune it, and inspect what it does, but cannot inspect what it was built from or exactly reproduce the run. "Open-weight" is the more precise term, and the distinction is not pedantic: it determines whether outside researchers can audit what went into the model or only what comes out of it. [4][3]
Reception, and a second "DeepSeek moment"
Moonshot's own numbers, reported at launch and not yet independently reproduced at the time, put K2-Instruct at 65.8% on SWE-bench Verified (agentic software-engineering tasks, using bash and file-edit tools in a single attempt), 53.7% on LiveCodeBench, 66.1 on Tau2-Bench retail (a tool-use benchmark that scores an agent acting as a customer-service bot), and 76.5 on ACEBench for tool calling. Moonshot framed these as wins against other non-reasoning open models rather than against reasoning models that spend extra inference-time compute working through a problem step by step, since K2 at launch did not have that mode. [3]
The reception was fast and visible outside benchmark tables. Global Times reported Hugging Face downloads rising from about 76,000 to 145,000 within days of the July 11 release, nearly doubling by the following Monday, and the model climbing to the top of the LMSYS Chatbot Arena leaderboard. Nature's news desk, in a piece China Daily and other state outlets reprinted, called K2 "another DeepSeek moment," the second time in six months a Chinese lab's open release had reset expectations for what an open model could do. [2][5]
Atlas interpretation: Nathan Lambert, a researcher at the Allen Institute for AI writing on his Interconnects newsletter, pushed back on the comparison while agreeing with the substance: he called K2 "the new best-available open model by a clear margin" but argued the January DeepSeek R1 moment had two ingredients K2 lacked, a training cost low enough to shock (DeepSeek's disclosed roughly $5 million) and a visible chain of reasoning users could read. K2's impact, on his account, would land more slowly and mostly with enterprises building on the API rather than with a mass consumer audience discovering exposed reasoning traces. That is a useful corrective to the headline framing: the two releases mattered for different reasons even when the press treated them as the same kind of event. [6]
Where it sits in 2025's open-weight race
Atlas interpretation: K2 landed three months after the two biggest open-weight releases from Western and other Chinese labs that spring. Meta's Llama 4 Scout and Maverick topped out at 400 billion total parameters with 17 billion active, and landed to a reception Meta itself had not planned for. Alibaba's Qwen3 family topped out at 235 billion total parameters with 22 billion active, shipped under a permissive Apache 2.0 license, and spread through the ecosystem partly for that license. K2's one trillion total parameters made it several times larger than either by that measure, while its 32 billion active parameters sat in the same range as Qwen3's biggest model, which is the number that mostly determines what it costs to run. [1]
Atlas interpretation: K2 also was not the end of the story it opened. On November 6, 2025, Moonshot released Kimi K2 Thinking, a reasoning version built on the same trillion-parameter, 32-billion-active architecture, adding the ability to interleave chain-of-thought reasoning with tool calls across 200 to 300 sequential steps without losing track, and released under the same license structure as the original. Read together with DeepSeek's R1 release in January and V3 before it, the 2025 pattern was not one surprising Chinese lab but a sequence of them, each shipping open weights that Western labs' closed models had to be measured against within days of release. [8]
Sources
- moonshotai/Kimi-K2-Instruct
Hugging Face · Jul 11, 2025
- World's first trillion-parameter open-source model, Kimi K2, doubles downloads in a single weekend
Global Times · Jul 22, 2025
- Kimi K2: Open Agentic Intelligence
Moonshot AI (Kimi Team), arXiv · Jul 28, 2025
- Kimi-K2/LICENSE
Moonshot AI · Sep 8, 2026
- Chinese AI model Kimi K2 marks 'another DeepSeek moment': Nature
China Daily · Jul 21, 2025
- Kimi K2 and when "DeepSeek moments" become normal
Interconnects (Nathan Lambert, Allen Institute for AI) · Jul 14, 2025
- Moonshot AI: $3 billion valuation overshadowed by legal dispute with 5 key investors
TechNode · Dec 11, 2024
- Alibaba-backed Moonshot releases new AI model Kimi K2 Thinking
CNBC · Nov 6, 2025