Chinese AI model now finds more software bugs than Anthropic's best

Chinese AI startup Zhipu announced its new GLM-5.3 model has outperformed Anthropic's Mythos 5 on the CyberGym cybersecurity benchmark. The company stated the model's cyber capability developed faster than expected during post-training, allowing it to reason across multi-stage exploitation chains. In real-world tests with Chinese companies, GLM-5.3 found 2,436 vulnerabilities across 269 projects, including 1,097 medium-to-high severity issues, some decades old. This rapid progress occurs as the White House drafts plans to force countries to pick a side in the US-China AI competition, signaling that any US advantage in this domain has dissipated.

Chinese AI model now finds more software bugs than Anthropic's best
3 sources 1 view
Updated Aug 17, 2026 First published Aug 16, 2026

Topic overview

Briefly

  • Zhipu's GLM-5.3 model found 2,436 vulnerabilities in 269 real-world software projects.
  • The model can now reason across multiple stages of exploitation, not just find isolated flaws.
  • The US is drafting plans to make countries choose sides in the US-China AI competition.

What happened

Chinese AI startup Zhipu has announced its new GLM-5.3 model, claiming it has surpassed Anthropic's Mythos 5 on a critical cybersecurity benchmark called CyberGym. This benchmark tests a model's ability to solve real-world cybersecurity challenges, and Zhipu's achievement marks a significant moment in the ongoing AI competition between China and the West. The company, also known as Z.ai, stated that as they scaled post-training, the model's cyber capability developed faster than expected, becoming state-of-the-art for vulnerability discovery. The model did not just get better at finding isolated flaws; it began to reason across multiple stages of exploitation, forming coherent plans for complete exploitation chains. This rapid iteration comes as the White House is reportedly drafting plans to force countries to choose sides in the US-China AI race, demanding they not participate in Beijing's framework.

Beyond benchmark tests, Zhipu has collaborated with Chinese companies to test GLM-5.3 on real-world codebases. The results were substantial, with the model finding 2,436 vulnerabilities across 269 projects. Of these, 1,097 were classified as medium-to-high severity issues. The findings spanned a wide range of software, including system kernels, operating systems, browser engines, open-source infrastructure, web applications, and network protocols. The company noted that many of these vulnerabilities had remained unnoticed for years or even decades, with the oldest dating back roughly 40 years. This practical application demonstrates that the model's capabilities extend beyond theoretical benchmarks and into tangible, impactful security research.

While GLM-5.3 excelled in the CyberGym benchmark, it is important to note that it performed worse than Western models on other security and coding benchmarks. However, the specific capability of being a highly effective bug finder is strategically crucial. It signals that China is not far behind in its ability to probe and find weaknesses in its rivals' software, a capability that developed very quickly after the debut of Anthropic's Mythos. This rapid progress suggests that any perceived advantage the United States felt it had as the home of leading AI companies like Anthropic has dissipated. The development underscores a narrowing gap in AI capabilities, particularly in applications with direct national security implications.

The announcement from Zhipu is part of a broader narrative of Chinese open-weight models claiming parity with the Western AI frontier. The company is iterating fast, with this release following its last one in June. The benchmark data provided by Zhipu claims that GLM-5.3 beats not only Mythos 5 but also Fable 5 and GPT-5.6 Sol on the CyberGym test. This specific focus on cybersecurity vulnerability discovery and exploitation represents a new front in the AI competition, moving beyond general language and coding tasks to highly specialized, offensive security applications. The fact that the model can reason across multi-stage exploitation chains points to a sophisticated level of planning and execution that was previously a hallmark of top-tier human security researchers.

Comprehensive report

Full story,
in detail.

Trace the developments that led here, see how the story evolved, and understand the forces and wider context surrounding it.

Entities

How Mestios works We aggregate coverage, extract key information, and use AI to summarize and compare perspectives. Learn more

Updated Aug 17, 2026

AI-generated summary. Please verify important information from original sources.