"As we scaled post-training, cyber capability developed faster than we expected. GLM-5.3 is state of the art on CyberGym for vulnerability discovery, and its gains are largest further up the exploitation chain," the company wrote.
Zhipu’s GLM-5.3: a deliberate bet on automated vulnerability discovery
Chinese AI company Zhipu last week announced GLM-5.3 and positioned the model explicitly as a step-change in automated bug-finding. In its announcement the company said the model “did not simply become better at identifying isolated flaws: it began to reason across multiple stages of exploitation, forming coherent plans for complete exploitation chains.” That language frames GLM-5.3 not only as a tool for locating defects but as one that can connect multiple findings into end-to-end exploit sequences.
CyberGym results: beating Fable 5 and GPT-5.6 Sol on a cyber benchmark
Zhipu published benchmark data showing GLM-5.3 outperforming two western models — Fable 5 and GPT-5.6 Sol — on CyberGym, a test the company describes as assessing a model’s ability to solve real-world cybersecurity challenges. The company emphasized that “gains are largest further up the exploitation chain,” signaling better performance on multi-step, linked exploitation tasks rather than isolated bug identification alone.

This site is the portfolio.
OSINTSights runs on Cloudflare Workers, D1, R2, and Vectorize, with an AI pipeline on Hetzner ARM. Nubivance designed, built, and operates it. We do the same for clients.
See what we buildReal-world testing: 2,436 vulnerabilities across 269 projects
Zhipu said it worked with Chinese companies to evaluate GLM-5.3 against real-world codebases. According to the announcement, the model identified 2,436 vulnerabilities across 269 projects, including 1,097 issues the company classified as medium-to-high severity. The affected codebases span “system kernels, operating systems, browser engines, open-source infrastructure, web applications, and network protocols,” and Zhipu noted some problems had “remained unnoticed for years or even decades, with the oldest dating back roughly 40 years.”
Mixed comparative performance: strong on CyberGym, weaker elsewhere
The announcement and reporting also made a point of nuance: GLM-5.3 performed worse than western models on other security and coding benchmarks. That mixed outcome—dominant on one cyber-focused benchmark but lagging on many others—frames the result as a targeted advance rather than an unambiguous, across-the-board superiority.
How security teams, Western AI developers, and Chinese software owners are affected
- Security and incident-response teams: The scale of Zhipu’s reported findings—2,436 vulnerabilities and more than 1,000 medium-to-high severity issues—suggests organizations that permit automated model-assisted scanning will need processes to triage large volumes of machine-identified issues and to verify exploitability when models link multiple flaws into exploitation chains.
- Western AI developers and model evaluators: The CyberGym result — GLM-5.3 beating Fable 5 and GPT-5.6 Sol on that benchmark — is a datapoint that western teams and benchmark designers will likely examine. The company’s own disclosure that GLM-5.3 is weaker on other security and coding tests implies benchmark- and task-specific competency, not a blanket lead.
- Chinese companies and open-source maintainers: Zhipu’s collaboration with domestic firms produced a large tally of findings spanning core system and application layers. Those owners now face remediation and prioritization choices for flaws that, according to Zhipu, in some cases have persisted for decades.
The story as released places GLM-5.3 at an inflection point: a model that, by the company’s measures, markedly accelerates post-training cyber capability and excels on a real-world exploitation benchmark, yet does not dominate across every security evaluation. The announcement also frames the development as fast-following the debut of Anthropic’s Mythos, a sequence the article characterizes as eroding any single-nation advantage in offensive-quality bug hunting. Whether GLM-5.3’s gains translate into widespread defensive improvement, new automated red-team workflows, or unforeseen risks from scaled exploit planning remains a concrete and immediate question for organisations that handle the affected codebases.
Read the original report: https://www.theregister.com/security/2026/08/17/chinese-ai-company-zhipu-claims-its-new-is-a-better-bug-finder-than-anthropic-openai/5288203




