Topic overview
Briefly
- Zhipu's GLM-5.3 model found 2,436 vulnerabilities in 269 real-world software projects.
- The model can now reason across multiple stages of exploitation, not just find isolated flaws.
- The US is drafting plans to make countries choose sides in the US-China AI competition.
What happened
Chinese AI startup Zhipu has announced its new GLM-5.3 model, claiming it has surpassed Anthropic's Mythos 5 on a critical cybersecurity benchmark called CyberGym. This benchmark tests a model's ability to solve real-world cybersecurity challenges, and Zhipu's achievement marks a significant moment in the ongoing AI competition between China and the West. The company, also known as Z.ai, stated that as they scaled post-training, the model's cyber capability developed faster than expected, becoming state-of-the-art for vulnerability discovery. The model did not just get better at finding isolated flaws; it began to reason across multiple stages of exploitation, forming coherent plans for complete exploitation chains. This rapid iteration comes as the White House is reportedly drafting plans to force countries to choose sides in the US-China AI race, demanding they not participate in Beijing's framework.
Beyond benchmark tests, Zhipu has collaborated with Chinese companies to test GLM-5.3 on real-world codebases. The results were substantial, with the model finding 2,436 vulnerabilities across 269 projects. Of these, 1,097 were classified as medium-to-high severity issues. The findings spanned a wide range of software, including system kernels, operating systems, browser engines, open-source infrastructure, web applications, and network protocols. The company noted that many of these vulnerabilities had remained unnoticed for years or even decades, with the oldest dating back roughly 40 years. This practical application demonstrates that the model's capabilities extend beyond theoretical benchmarks and into tangible, impactful security research.
While GLM-5.3 excelled in the CyberGym benchmark, it is important to note that it performed worse than Western models on other security and coding benchmarks. However, the specific capability of being a highly effective bug finder is strategically crucial. It signals that China is not far behind in its ability to probe and find weaknesses in its rivals' software, a capability that developed very quickly after the debut of Anthropic's Mythos. This rapid progress suggests that any perceived advantage the United States felt it had as the home of leading AI companies like Anthropic has dissipated. The development underscores a narrowing gap in AI capabilities, particularly in applications with direct national security implications.
The announcement from Zhipu is part of a broader narrative of Chinese open-weight models claiming parity with the Western AI frontier. The company is iterating fast, with this release following its last one in June. The benchmark data provided by Zhipu claims that GLM-5.3 beats not only Mythos 5 but also Fable 5 and GPT-5.6 Sol on the CyberGym test. This specific focus on cybersecurity vulnerability discovery and exploitation represents a new front in the AI competition, moving beyond general language and coding tasks to highly specialized, offensive security applications. The fact that the model can reason across multi-stage exploitation chains points to a sophisticated level of planning and execution that was previously a hallmark of top-tier human security researchers.

Comprehensive report
Full story,
in detail.
Trace the developments that led here, see how the story evolved, and understand the forces and wider context surrounding it.
