Z.ai has unveiled GLM-5.3: a breakthrough in agentic programming and cybersecurity

Chinese startup Z.ai has officially unveiled its GLM-5.3 language model, betting on significant improvements in coding, agentic scenarios, and vulnerability discovery. This update is not just an iteration but a strategic move that redefines the company's positioning in the AI development race.
Key improvements and benchmarks
Contrary to market expectations, the performance gains were achieved solely through post-training, without changing the underlying architecture. This approach allows Z.ai to rapidly scale improvements while keeping computational costs at the same level. On the internal Code Bench test, which evaluates coding skills, GLM-5.3 showed a 50% improvement over its predecessor GLM-5.2.
The results on specialized benchmarks are particularly impressive:
- Terminal Bench 3.0 — 28.3 points versus 4.6 for GLM-5.2 (a 6-fold increase);
- DeepSWE v1.1 — 66.9 versus 46.2;
- Agents' Last Exam — 28.5 versus 23.8;
- CyberGym — 84.5% versus 77.2%;
- ExploitBench — 54.4% versus 24.4% (more than a twofold increase).
Breakthrough in vulnerability discovery
ExploitGym deserves special attention, where the model demonstrated outstanding efficiency: within a two-hour budget, it completed 105 tasks versus 29 for GLM-5.2, and with a six-hour limit — 130 tasks instead of 39. These figures indicate a qualitative leap in autonomous security analysis.
In real-world tests on codebases conducted jointly with Chinese security teams, the system identified 2,436 vulnerabilities across 269 projects, of which 1,097 were of medium and high criticality. Only 53 findings were publicly disclosed, with the remaining 2,383 kept under embargo. The average "age" of discovered vulnerabilities is 26.6 years, with the oldest dating back to 1981. For transparency, Z.ai has launched the public Z.ai Security Disclosure Ledger, tracking the status and criticality of findings.
Product integration and context
GLM-5.3 is already deployed for GLM Coding Plan subscribers, where it becomes the primary engine for agentic programming and long-running tasks. Requests to older models are automatically routed to the new version, simplifying user migration. Open weights will be released two weeks after security assessment and hardening.
It's worth recalling that in July, Anthropic accused Z.ai of unauthorized distillation of GLM-5.2 using responses from Claude and GPT. This context adds intrigue: the success of GLM-5.3 could be the result of both proprietary research and controversial methods, raising questions about the boundaries of innovation in the industry.
My take: Z.ai demonstrates that post-training is an underestimated lever for rapidly improving models, especially in niche domains like cybersecurity. However, such an aggressive focus on ExploitBench may signal a market shift toward offensive AI tools, which will inevitably spark new ethical and regulatory debates.