A new number on the open-weights leaderboard
Artificial Analysis put Zhipu AI's GLM-5.2 at the top of its Intelligence Index v4.1 with a score of 51, eleven points above its predecessor GLM-5.1 and ahead of the next-best open-weights models: MiniMax-M3 and DeepSeek V4 Pro in max mode, both at 44, and Kimi K2.6 at 43. [1]
GLM-5.2 pairs 744 billion total parameters with 40 billion active per token, a mixture-of-experts design, and Zhipu released the weights under an MIT license. Context length grew to 1 million tokens, up from 200,000 tokens in GLM-5.1. Zhipu's API prices it at $1.40 per million input tokens and $4.40 per million output tokens, with a lower rate for cache hits. [1]
One index, not a consensus
On Artificial Analysis's GDPval-AA v2 benchmark, which scores performance on economically valuable knowledge work, GLM-5.2 reached 1524, ahead of MiniMax-M3 (1418) and DeepSeek V4 Pro in max mode (1328) and effectively level with GPT-5.5 running its highest reasoning setting, which scored 1514. Getting there used more output tokens per task, about 43,000, than any of the other leading open-weights models the article compared it against. [1]
Atlas interpretation: Artificial Analysis builds its own suite of tasks and weighs them by its own judgment, so one evaluator naming a new leader is not the same as the field agreeing on one. What the numbers do establish is the size of the jump: eleven points in the five weeks since GLM-5.1, a larger single move than this index had shown from the open-weights field before, which is the more interesting fact than the rank itself. [1]
Ahead of Claude, on one benchmark, without its scaffolding
Two weeks later, the security company Semgrep ran GLM-5.2 against its own dataset of real, open-source applications on an IDOR (insecure direct object reference) detection task, scoring results by F1. GLM-5.2 scored 39%, ahead of Claude Code on Opus 4.6 at 37% and on Opus 4.7/4.8 at 28%, and did it for roughly $0.17 per vulnerability found. Semgrep held the dataset, the scoring method and the system prompt constant and varied only the model, and gave the open-weights models a basic prompt without the endpoint-discovery scaffolding built into Semgrep's own product. [2]
Atlas interpretation: Semgrep's own writeup undercuts the headline before anyone else can: "this is one task, one dataset, one run," the authors wrote, and warned results could look different on other vulnerability types such as SSRF. They also reported that their proprietary scaffolding, tested separately, scored 53 to 61% F1, well above any single model in the comparison, including GLM-5.2. The finding that survives that caveat is narrower than the summary above: on this one detection task and without either vendor's own tooling, an open-weights model outscored Claude at a fraction of the cost, not that it outperforms Claude generally. [2]
The third Chinese lab in seven months
Atlas interpretation: GLM-5.2 is not the first model to take this particular lead in 2026, or even the first from a Chinese lab. Moonshot's Kimi K2 Thinking topped the open-weight field in November 2025, and DeepSeek's V4 Pro reclaimed it in April 2026 before GLM-5.2 passed it in June. Each of those releases beat the previous one by a modest margin and held the position only until the next lab shipped. The pattern says more about the pace of the open-weights race than about any single model: whoever is on top on a given week has usually lost that spot within a few months. [1]
Sources
- GLM-5.2 is the new leading open weights model on the Artificial Analysis Intelligence Index
Artificial Analysis · Jun 16, 2026
- GLM 5.2 beats Claude in our cyber benchmarks
Semgrep · Jun 28, 2026