A downloadable model on top of a code-repair benchmark
Zhipu AI's GLM-5.1 is a 754 billion parameter mixture-of-experts model, released under an MIT license with weights downloadable from Hugging Face. Its model card reports a score of 58.4 on SWE-Bench Pro, a benchmark of real-world code repair tasks, ahead of the scores it lists for GPT-5.4, Claude Opus 4.6 and Gemini 3.1 Pro on the same benchmark. The model card also reports 95.3% on AIME 2026, 86.2% on GPQA-Diamond and 63.5% on Terminal-Bench 2.0. [1]
The model ships with a 200,000 token context window and a 128,000 token maximum output, and Zhipu describes the architecture as combining mixture-of-experts routing with dynamic sparse activation, trained with asynchronous reinforcement learning. It can be self-hosted through SGLang, vLLM, xLLM, Transformers or KTransformers, or accessed through Z.ai's API with OpenAI SDK compatibility. [1]
Built for long, unsupervised runs
MarkTechPost's coverage the following day described GLM-5.1 working autonomously on a single task for up to 8 hours, carrying it through planning, execution, testing, fixing and delivery without intervention. The cited examples include building a Linux desktop environment from scratch, running 178 autonomous iterations on a vector database task, and optimizing a CUDA kernel from a 2.6x to a 35.7x speedup over the unoptimized version. [2]
Atlas interpretation: That framing matters more than the SWE-Bench Pro number by itself. A model that plateaus after its first attempt at a hard problem needs a bigger context window; a model meant to hold up over hundreds of tool calls needs to keep making sound judgment calls as errors accumulate. The Linux-desktop and CUDA-kernel examples are demonstrations chosen by Zhipu and its coverage, not an independent long-horizon benchmark, so they show what the model can do under a favorable run rather than how reliably it does it. [2]
A vendor's own scoreboard
Atlas interpretation: The SWE-Bench Pro comparison to GPT-5.4, Claude Opus 4.6 and Gemini 3.1 Pro comes from Zhipu's own model card, repeated in MarkTechPost's writeup without an independent third party rerunning the benchmark against all four models under matching conditions. That does not make the number wrong, but it is a self-reported result rather than an audited one, and it is one benchmark: SWE-Bench Pro measures real-world code repair, not the full range of coding, reasoning or agentic tasks these models are used for. What is well established is that a 754 billion parameter model, released with open weights under an MIT license, posted a competitive score on a benchmark that closed frontier models had led until this release. [1][2]
Sources
- zai-org/GLM-5.1
Z.ai · Apr 7, 2026
- Z.AI Introduces GLM-5.1: An Open-Weight 754B Agentic Model That Achieves SOTA on SWE-Bench Pro and Sustains 8-Hour Autonomous Execution
MarkTechPost · Apr 8, 2026