GLM-5.3: Post-Training Gains, CyberGym and Weights Plan

Z.ai kept GLM-5.2's base model and reported coding gains and an 84.5 percent CyberGym score; open weights were promised roughly two weeks later, after a safety review.

The base model did not change

Z.ai's own framing for GLM-5.3 was specific: it runs on the same base model as GLM-5.2 from two months earlier, and every reported gain came from a longer, larger post-training run rather than a new pretrain. The company described that run as covering more task environments, more environment types, and more training time, without naming the specific reinforcement learning recipe. [2]

Z.ai reported building the post-training infrastructure on two open-source pieces it credits by name, an environment framework called slime and an asynchronous reinforcement learning system called SAO, and said the sandbox tasks used for training ran on developer workstations for days at a stretch to push the model toward longer-horizon work. [1]

Atlas interpretation: That a fixed base model can move this far on benchmarks through post-training alone is itself the notable part. It says Zhipu is treating GLM-5.2's pretrain as a reusable asset and iterating on the training loop wrapped around it, the same bet several other labs have been making with their own model lines this cycle. [2]

Coding gains concentrate on long tasks

The clearest jumps were on benchmarks built around long-running agentic coding rather than single-function completion. Z.ai reported Terminal-Bench 3.0 moving from 4.6 to 28.3 and DeepSWE v1.1 moving from 46.2 to 66.9 versus GLM-5.2. On its own internal coding-agent benchmark it reported a 50 percent improvement, with the model scoring 31.4 percent at roughly 50,000 output tokens, a proxy for how far a task can run before it falls apart. [2]

These are vendor-reported figures from Z.ai's own benchmark suite, not an independent evaluation, and the comparison baseline is Z.ai's own prior model. On harder public coding suites the reporting places GLM-5.3 behind GPT-5.6 Sol and Claude Fable 5, so the claim is a lead among open coding models on long-horizon tasks specifically, not a general coding-benchmark lead. [2]

A cybersecurity result the company did not expect

On CyberGym, a benchmark for finding exploitable vulnerabilities in real codebases, GLM-5.3 scored 84.5 percent, up from 77.2 percent for GLM-5.2 and just ahead of Anthropic's Claude Mythos 5 at 83.8 percent. Z.ai also reported the model had been used to find more than 2,400 vulnerabilities across 269 projects, with roughly half rated medium severity or worse. [1][2]

The lead did not hold on every cybersecurity measure. On ExploitBench, which scores full exploitation chains rather than vulnerability discovery alone, GLM-5.3 more than doubled its prior score to 54.4 percent but still trailed Mythos 5's 78 percent, and reporting notes the gap widens the deeper a benchmark task goes into actually exploiting a found flaw rather than just spotting it. [2]

Available now to subscribers, weights held for a safety pass

At announcement, GLM-5.3 was reachable only through Z.ai's GLM Coding Plan subscription. Open weights were promised on Hugging Face under an open-source license roughly two weeks later, once what the company called safety evaluation and hardening finished. [2][1]

Atlas interpretation: Holding weights back for a security pass reads differently for a model whose own headline result is vulnerability discovery: the delay doubles as an argument that the release process took its own CyberGym-flavored risk seriously, whether or not that was the intent. [1]

Sources

  1. Z.ai debuts GLM-5.3 with long-horizon coding, cybersecurity upgrades

    SiliconANGLE · Aug 14, 2026

  2. Z.ai Ships GLM-5.3 Without Retraining the Base Model: Better at Complex Coding and Long-Horizon Tasks

    MarkTechPost · Aug 14, 2026

  3. CAISI’s Assessment of Z.ai’s GLM-5.3 Cyber Capabilities

    NIST · Sep 17, 2026