Skip to main content
AI & Machine LearningQuantum Computing

Zhipu Unveils AI Model Rivalling Western Bug-Finders

Lab setting with computer screen displaying code, researcher working in background.

“GLM-5.3 is state of the art on CyberGym for vulnerability discovery,” Zhipu wrote — and backed the claim with a concrete tally: 2,436 vulnerabilities located across 269 projects, including 1,097 of medium-to-high severity.

Zhipu’s claim and the CyberGym headline result

Chinese AI company Zhipu last week announced a new model, GLM-5.3, and published benchmark data claiming it outperforms two prominent Western systems on CyberGym, a test designed to measure a model’s ability to solve real‑world cybersecurity challenges. Zhipu’s announcement says GLM-5.3 beats Fable 5 and GPT‑5.6 Sol on that benchmark and that, “as we scaled post‑training, cyber capability developed faster than we expected.”

How Zhipu describes the model’s reasoning across exploitation chains

In its announcement the company emphasized not just isolated flaw‑finding but cross‑stage reasoning. “The model did not simply become better at identifying isolated flaws: it began to reason across multiple stages of exploitation, forming coherent plans for complete exploitation chains,” Zhipu wrote. The company highlighted that the model’s gains were “largest further up the exploitation chain,” framing the advance as a step beyond single‑point detection toward multi‑step attack planning.

Real‑world testing: 2,436 vulnerabilities across kernels to web apps

Zhipu said it worked with Chinese companies to exercise GLM‑5.3 against real codebases. The company reports it found 2,436 vulnerabilities spanning 269 projects, and classified 1,097 of those as medium‑to‑high severity. Zhipu listed the affected areas as system kernels, operating systems, browser engines, open‑source infrastructure, web applications, and network protocols. The announcement adds a chronological note: “Many had remained unnoticed for years or even decades, with the oldest dating back roughly 40 years.”

Benchmark nuance: strengths on CyberGym, weaker elsewhere

Zhipu’s release does not present an unqualified superiority claim. It notes that GLM‑5.3 “also performed worse than western models on other security and coding benchmarks.” The company frames its CyberGym performance as a notable capability while acknowledging mixed results across the broader set of evaluations it ran.

How technologists, vendors, and policymakers are named in the claim

  • Technologists and security teams: Zhipu’s report stresses the model’s ability to find vulnerabilities across low‑level system software and user‑facing applications, which suggests the kinds of codebases — system kernels, operating systems, browser engines, open‑source infrastructure, web applications, and network protocols — that it was applied to.
  • Vendors and software maintainers: the company’s finding that many flaws “had remained unnoticed for years or even decades” highlights long‑standing latent issues in diverse codebases the announcement says GLM‑5.3 can surface.
  • Policymakers and international competitors: the release explicitly positions GLM‑5.3 in a geopolitically framed contest. The piece notes that the model’s rapid development “very quickly” followed the debut of Anthropic’s Mythos and that, in Zhipu’s framing, “Any advantage the US felt it had as the home of Anthropic has therefore dissipated.”

That last point is presented in the source as an assessment of competitive parity: the company’s CyberGym result and its reported real‑world discoveries are offered as evidence that China “is not far behind” in AI‑driven vulnerability discovery.

Zhipu’s announcement couples a headline benchmark victory on CyberGym with a large, concrete count of discovered vulnerabilities and a candid acknowledgment of weaker outcomes on other tests. Together those elements set a specific agenda for readers: the model claims advanced exploit reasoning and produced a large vulnerability haul in collaborative testing, but its superiority is not uniform across benchmarks.

Whether GLM‑5.3’s CyberGym lead will translate into consistent superiority outside that benchmark remains an explicit part of the story itself: the company reports both strong CyberGym performance and weaker showings elsewhere. The announcement also leaves on the record the scale of the vulnerabilities it says it found — 2,436 total, 1,097 medium‑to‑high severity, across 269 projects — and the industries and code types those findings touched.

Original story