A trillion parameters, 32 billion active, built on native vision
Kimi K2.5 is a mixture-of-experts model with 1 trillion total parameters and 32 billion active per token, drawn from 384 routed experts plus one shared expert, with 8 routed experts selected per token across 61 layers using multi-head latent attention. A 400-million-parameter MoonViT vision encoder gives the model native image and video input, and the model card lists a 256,000-token context window. [3]
Moonshot built K2.5 as continual pretraining on top of Kimi K2's base model, adding roughly 15 trillion mixed visual and text tokens on top of the existing text pretraining rather than training the vision capability from scratch. Moonshot released the weights under its own Modified MIT License, the same license family it used for K2 and K2 Thinking. [3]
Atlas interpretation: The 32 billion active parameters, not the 1 trillion total, is the number that determines how expensive a single forward pass actually is. Building the multimodal capability as continual pretraining on an existing base model, rather than a fresh run, is also why Moonshot could add native vision without changing the active-parameter footprint that made the earlier Kimi K2 line cheap to serve. [3]
Agent Swarm: up to 100 sub-agents, trained with PARL
K2.5's Agent Swarm mode, in beta at launch, lets the model decompose a complex task and dynamically instantiate up to 100 domain-specific sub-agents to work on pieces of it in parallel, coordinating up to 1,500 tool-call steps through a built-in orchestration engine. Moonshot says it trained this capability with a method it calls Parallel-Agent Reinforcement Learning, or PARL, in which the model learns to direct the swarm rather than being hand-coded to split tasks a fixed way. [2][1]
Moonshot reports that Agent Swarm cuts the minimum critical steps needed to hit a target performance level by 3x to 4.5x compared to running the same task with a single agent in a wide-search scenario, which it translates to up to a 4.5x reduction in wall-clock time from the parallelization alone. [2]
Atlas interpretation: The speedup figure describes wall-clock time on tasks that parallelize well, mainly broad search and research-style workloads where many sub-agents can explore different branches at once. It is not a claim that the model reasons better with a swarm, only that a swarm reaches the same answer faster when the task can be split. Moonshot's own framing, and the beta label, both signal this is a workload-shaped tool rather than a general capability upgrade. [2]
A self-reported top score, and a qualified win over GPT-5.2
Moonshot reports K2.5 scores of 31.5 percent (text) and 21.3 percent (image) on the full Humanity's Last Exam set without tool use, rising to 51.8 percent and 39.8 percent with tools. SiliconANGLE's coverage characterizes this as the highest score on HLE-Full, a 2,500-question set spanning math and physics among other fields, among the models Moonshot compared against. [2][1]
Across the roughly two dozen other benchmarks Moonshot published, SiliconANGLE reports that K2.5 came within a few percentage points of GPT-5.2 and Claude 4.5 Opus, and beat GPT-5.2 on several of them, without giving exact margins for those individual comparisons. [1]
Atlas interpretation: Every number here is Moonshot's own reporting from its release materials. No independent lab result appears in either source, and the HLE and cross-model comparisons both depend on which prompting setup and tool access Moonshot chose for each competing model, a choice the vendor controls in its own favor. Treat the top score and the wins over GPT-5.2 as a vendor's benchmark, not a verified ranking, until independent evaluations confirm them. [2][1]
Sources
- Moonshot AI releases open-source Kimi K2.5 model with 1T parameters
SiliconANGLE · Jan 27, 2026
- Kimi K2.5: Visual Agentic Intelligence
Moonshot AI · Aug 20, 2026
- moonshotai/Kimi-K2.5
Hugging Face · Sep 9, 2026