Z.ai releases GLM-5.3: a breakthrough in coding and cybersecurity

Chinese startup Z.ai has officially unveiled its GLM-5.3 language model, betting on significant improvements in programming, agentic scenarios, and vulnerability discovery. This update is positioned as a response to the growing demands on AI systems in software development and cybersecurity.
Key improvements and benchmarks
Z.ai emphasizes that the performance gains were achieved solely through post-training, without changing the underlying architecture, which sets this release apart from previous iterations. On the internal Code Bench test, the new model showed a 50% improvement over GLM-5.2, indicating serious algorithm optimization.
In the release table, Z.ai cites impressive results across a range of specialized tests:
- Terminal Bench 3.0: 28.3 points versus 4.6 for GLM-5.2 — nearly a sixfold increase.
- DeepSWE v1.1: 66.9 versus 46.2 — a significant leap in solving development tasks.
- Agents' Last Exam: 28.5 versus 23.8 — improved agentic capabilities.
- CyberGym: 84.5% versus 77.2% — enhanced effectiveness in cyber environments.
- ExploitBench: 54.4% versus 24.4% — doubled results in exploit discovery.
ExploitGym stands out in particular: the model completed 105 tasks within a two-hour budget versus 29 for its predecessor, and 130 versus 39 with a six-hour budget. This demonstrates not only speed but also depth of analysis.
Practical application and security
Z.ai also conducted testing on real codebases in collaboration with security teams in China. After verification and deduplication, the system identified 2,436 vulnerabilities across 269 projects, including 1,097 issues of medium and high severity. Only 53 findings were publicly disclosed, while the remaining 2,383 remain under embargo. The average "lifespan" of these vulnerabilities is estimated at 26.6 years, with the oldest dating back to 1981 — underscoring the scale of the legacy code problem.
For transparency, the Z.ai Security Disclosure Ledger has been launched — a public registry tracking status, affected projects, and severity of findings. This is an important step for an industry where openness is often lacking.
Product integration
GLM-5.3 has already been deployed for GLM Coding Plan subscribers, becoming the primary engine for agentic programming and long-running tasks. Requests to older models are automatically routed to the new one, simplifying migration for users. Open weights will be released two weeks after security assessment and hardening.
It's worth recalling that in July, Anthropic accused Z.ai of unauthorized distillation of GLM-5.2 using Claude and GPT responses, adding context to this ambitious release.
My analysis: GLM-5.3 is not just an incremental update but a targeted leap in niche yet critically important areas. The twofold improvement in ExploitBench and CyberGym suggests that Z.ai is seriously investing in security, which could become a key competitive advantage amid global demand for AI tools to protect code. However, questions of ethics and technology provenance remain open, and this could hinder the model's adoption in the West.