If Chinese AI goes rogue, expect opacity

We now know that Chinese AI agents are inclined to cheat just as US models are. The big difference is that if one causes a cybersecurity incident, the ways we would find out are much more restricted.
OpenAI’s chief strategy officer fronted a livestreamed parliamentary inquiry in Sydney on Tuesday and apologised for the company’s handling of AI agents that accessed Australian government websites without authorisation. He acknowledged that OpenAI should have told the Australian government sooner.
This was the latest example of US labs demonstrating some transparency with governments and civil society over cybersecurity incidents caused by their AI agents. These labs have done this despite there being no general US legal requirement for disclosure on every cybersecurity-related incident. Expect nothing so straightforward if a Chinese lab gets into a similar situation.
Chinese labs have an opacity problem. Many withhold details of internal safety tests and have no whistleblower policies. China’s 2025 rules on cybersecurity incidents require serious incidents to be reported to the authorities but not necessarily to the public.
Beijing’s rules also dictate how researchers and companies disclose network vulnerabilities, holes in code that AI agents like Anthropic’s Mythos 5 are adept at identifying. Chinese lab Z.ai previously built a public ledger of the vulnerabilities identified by GLM-5.3, saying the document listed only vulnerabilities that were already public. Then it took the ledger off its website, which now says future findings will be published through state vulnerability databases instead. The implication is that, even when Chinese labs want to disclose information, the government may stop them.
Chinese labs’ technical papers are one possible channel for understanding cybersecurity issues they’re encountering during training and testing. The papers may also help policymakers elsewhere to understand the current capabilities of Chinese models.
An assessment of papers released by Chinese labs in 2026 shows their agents cheating and trying to break out of contained testing, the sort of behaviour seen in OpenAI’s agents that have hacked into outside systems. But whether a serious incident would appear in one of these papers is up to a lab’s managers and the officials above them.
The papers describe Chinese agents exploiting network access that their sandboxes already allowed and, in some cases, crashing internal infrastructure. But none describes an agent breaking through a properly secured sandbox or causing a cybersecurity incident in an external party’s systems, as OpenAI’s agents have – for example, in hacking into a database of Australia’s government medical-insurance scheme, Medicare, in June.
Chinese labs are confronting the same technical problem identified by leading US labs: reward hacking, which is finding an unintended way to achieve a goal and thereby maximising a score on a task rather than completing it as intended. Colloquially, it can be called cheating.
For example, models are tested in sandboxes, virtual environments that cordon off an agent from other parts of a computer’s systems or the internet, but, driven by an incentive to complete a task, may hack their way out. A swarm of OpenAI agents did just that in July, because they saw that getting into the model-hosting platform Hugging Face was part of the answer to their problem.

But the way models are trained today makes reward hacking a feature, not a bug. ‘No silver bullet’ is part of the title of a paper from Alibaba’s Qwen team in June on how to reward coding agents reliably. It concludes ‘reward hacking is not a bug that can be patched but an inevitable consequence of sustained optimisation towards an imperfect objective.’
As agents have gained capability, so has their capacity for reward hacking. Technical papers published by Chinese labs in 2026 show the labs all grappling with agents gaming their rewards, much as Anthropic and OpenAI are. Chinese lab Moonshot has said that in early testing its Kimi K3 model sometimes crashed the lab’s internal systems during training, because increasingly capable agents ‘explore more aggressively’. Z.ai and fellow Chinese labs Xiaomi and DeepSeek have all reported agents using network access available from their sandboxes to seek answers outside the evaluation environment, including on GitHub, a cloud platform for storing code. DeepSeek goes furthest, noting that its model reached GitHub despite being denied access via other channels. The model also tried to forge official system instructions and bring down internal systems.
The intent to cheat is still there. In a paper published in late September, Alibaba Token Hub (the business group that houses Alibaba’s AI units, including the Qwen team) tested Alibaba’s latest model, Qwen-3.8, against models from OpenAI, Anthropic and Z.ai. It found that Z.ai’s GLM-5.3 was the most inclined to try to escape the sandbox to get answers and that Qwen-3.8 was still trying. The research did not measure whether those attempts succeeded.

What none of the technical papers disclose is evidence that models broke through a properly configured sandbox or caused a cybersecurity incident in a third party’s systems during training or testing. And no third-party evaluator has to date said that’s happened.
DeepSeek has been the most forthcoming on details; others have held back. The most opaque among the Chinese labs is MiniMax, whose blog posts on its M2 and M3 models have been candid about their reward hacking in some areas (in writing mathematical proofs and creating realistic chatbots) but are silent about whether this extended to sandbox escapes. Moonshot, likewise, has written about the trouble it had controlling its agent without mentioning the agent’s known tendency to leave the sandbox in search of answers. When US start-up Frontier Security tested Kimi K3 in August in a misconfigured sandbox, one that let the model reach a list of selected websites, including GitHub, the agent went straight to GitHub.
The OpenAI case also shows why disclosure matters. Its agents first reached beyond their sandbox through an internal package manager that was allowed internet access, and the public learned of the resulting Hugging Face incident from the affected third party before OpenAI published its own account. Under China’s cybersecurity-incident regime, serious incidents are reported to the authorities rather than being subject to an equivalent public-disclosure requirement. That makes routes to disclosure through developer and victims both less reliable.
As Chinese models grow more capable, there is no guarantee the technical papers will reflect any cybersecurity problems that may arise during training. The tenure of President Xi Jinping has seen a well-documented decline in publicly available information about emergencies and incidents in China. For outsiders trying to assess the risks of increasingly capable Chinese AI agents, technical papers will remain valuable, but they should not be mistaken for a reliable incident-reporting system.
