AI Safety in China #26
OpenClaw agent safety; TC260 AI Safety Working Group; CAICT AI Safety Benchmark 2.0; UN International Scientific Panel on AI; frontier AI risks in legal frameworks; open-source safety
Key Takeaways
OpenClaw drove a surge in policy attention to AI agent safety, with multiple government agencies issuing warnings, establishing dedicated standards, and conducting safety assessments targeting cloud deployment platforms.
China’s primary AI standards-setting body has established a dedicated AI Safety Working Group, signaling heightened institutional prioritization and potential acceleration of national AI safety standards development.
A major state-backed AI safety testing institution significantly expanded its safety benchmark to include frontier risks, scenario-specific safety, and agent safety.
Two Chinese experts have been appointed to the Independent International Scientific Panel on AI at the UN, with profiles oriented toward AI industrialization, capacity-building, and institutional governance.
Leading Chinese legal scholars highlighted extreme AI risks in proposals for comprehensive AI legislation.
Domestic AI Governance
OpenClaw brings AI agent safety into mainstream policy focus
Background: OpenClaw’s rapid proliferation in China has triggered a cascade of warnings from cybersecurity authorities, pushing agent safety into mainstream policy discussions.
Multiple Chinese government agencies have responded in several stages:
On February 5, the Ministry of Industry and Information Technology’s (MIIT) National Vulnerability Database (NVDB) issued an early warning, flagging “relatively high security risks” in certain OpenClaw versions.
In mid-March, CNCERT (the national cybersecurity emergency response team) warned of prompt injection attacks (hidden instructions embedded in web pages that manipulate OpenClaw into leaking data), misinterpretation of user intent, malicious plugins, and known vulnerabilities. A week later, a Ministry of State Security (MSS) guidance document, titled the “Safe Lobster Farming Manual” (安全养殖手册), warned of OpenClaw’s potential as a vector for spreading misinformation.
These warnings quickly spread beyond technical circles. After MSS’s publication, “safe lobster farming manual” topped 1 million views on RedNote/Xiaohongshu (a popular social networking platform in China), and a People’s Daily report on CNCERT’s warning was forwarded over 100,000 times on WeChat.
The response has since shifted towards more structured frameworks. CNCERT now provides tailored guidance for regular users, enterprises, cloud providers, and developers. TC260, China’s leading AI standards body, has released a draft practice guide codifying the deployment security lifecycle, covering installation, configuration, usage, uninstallation, and cloud environment selection for OpenClaw.
The AI Industry Alliance (AIIA) is focusing on cloud providers as the key governance lever, collecting best practices and conducting security tests. The AIIA has a track record of working on agent safety before the rise of OpenClaw, for instance through industry standards and leading AI agent safety commitments that 13 Chinese companies signed in early February.
The Director of the National Data Administration (NDA), LIU Liehong (刘烈宏), used OpenClaw to argue that a well-built agent “should not merely be a flashy ‘all-capable executor,’ but should be an honest risk communicator and reliable solution provider.” The NDA is China’s key governance body responsible for data and digital infrastructure oversight under the National Development and Reform Commission (NDRC).

Beyond government action, leading cybersecurity firms have issued guidance for secure deployments. Research teams have also published technical work on OpenClaw, including a five-layer lifecycle security framework from Tsinghua University and Ant Group, a security audit tool from the Beijing AI Safety Institute, and a vulnerability disclosure and patch from the China Academy of Information Communications Technology (CAICT), Shanghai Jiao Tong University, and Nanjing University researchers.
Implications: OpenClaw has transformed agent safety into a salient national governance concern, prompting responses from cybersecurity authorities, state media, senior officials, and the developer community. While China is already in the process of drafting AI agent security standards—including an industry standard for Model Context Protocol (MCP) security and a national standard on agent security frameworks—the OpenClaw episode is likely to significantly accelerate this work.
Key national standard-setting body establishes dedicated AI safety working group
Background: On March 20, TC260 announced the establishment of a dedicated AI Safety Working Group, designated WG9. TC260 is China’s primary standard-setting body for AI safety and has drafted key national standards on AI safety. Previously, its AI safety work was handled mainly by the Special Working Group for Emerging Technology Security (SWG-ETS), which also covered other domains such as quantum computing.
Content: The group’s leadership is comprised of:
Chair: ZHOU Bowen (周伯文), Director of Shanghai AI Lab
Vice-Chairs: ZHANG Zhen (张震) from the CNCERT; WEI Kai (魏凯) from CAICT; SHENG Xiaobao (盛小宝) from the Third Research Institute of the Ministry of Public Security; HOU Yuanwei (侯元伟) from the China Information Technology Security Evaluation Center.
In the group’s first meeting on March 31, Zhou Bowen specifically called for “anticipating frontier risks in advance” and strengthening international cooperation, such as exploring mechanisms for mutual international recognition of evaluation results.
On April 4, TC260 published priority areas for 2026, specifically listing
AI application security classification and grading;
AI security capability maturity assessment;
AI anthropomorphic interactive service security;
Agent safety;
AI safety guardrails;
AI data security and personal information protection.
These reflect the most immediate standard drafting priorities for WG9. More comprehensive plans will likely be published in the near future. The initial announcement mentions that WG9’s mandate includes “proposing a structured AI safety standards system,” which was also discussed at the group’s first meeting. This refers to a specific document: in January 2025, TC260 published the AI Safety Standards System (V1.0), a blueprint outlining the standards TC260 plans to draft to implement its 2024 AI Safety Governance Framework.
While TC260 has released an updated AI Safety Governance Framework 2.0 in September 2025, it has not yet released a corresponding update to the Standards System. Producing a revised Standards System might therefore be among WG9’s first priorities. Once released, it should provide greater clarity on what specific standards the group intends to develop.
Implications: The establishment of a dedicated working group signals that AI safety is becoming a higher priority within TC260, potentially accelerating the development of national standards. A key near-term indicator will be whether—and when—WG9 releases an updated AI Safety Standards System.
State-backed safety tester expands safety benchmark to include frontier risks
Background: On February 10, the China Academy of Information Communications Technology (CAICT), a think tank under the Ministry of Industry and Information Technology (MIIT), released the “AI Safety Benchmark 2.0“, an update to the benchmark it first launched in 2024. This update is the most comprehensive overhaul of the evaluations yet and significantly increases attention to frontier safety.
Content: Building on the original evaluation dimensions of content safety, adversarial safety, and application safety, the updated framework introduces three new dimensions:
Frontier safety, including model self-awareness, deception, loss of control, and misuse in dangerous domains;
Scenario safety, e.g. in specific domains like energy, finance, or transport;
Safety in emerging application domains, namely agents and agent platforms.

Implications: This update formally integrates frontier AI safety risks into the guiding framework of one of China’s most prominent government-backed AI safety testing institutions. This expanded scope is particularly noteworthy given that, during 2025, published evaluation results had actually narrowed relative to the original 2024 framework — focusing only on hallucinations and coding model security.
CAICT has historically published anonymized evaluation results based on its benchmarks. Initial results released for Q1 2026 focus on on-device agents, finding that while harmful content generation rates remain low, many systems still struggle to reliably refuse unsafe or malicious task execution. However, the testing categories announced for the first half of 2026 cover only a subset of the broader risk landscape outlined in AI Safety Benchmark 2.0. Hence, it remains unclear when the full range of newly introduced frontier and scenario-based risks will be systematically evaluated and published.
International AI Governance
Chinese experts on the Independent International Scientific Panel on AI
Background: On February 12, the United Nations (UN) General Assembly approved the appointment of the 40 members of the Independent International Scientific Panel on AI, including two Chinese members: SONG Haitao (宋海涛) and WANG Jian (王坚). Members serve in their personal capacity on the panel, which is mandated to produce an annual report with “evidence-based scientific assessments related to the opportunities, risks and impacts of AI”.
The two Chinese members: Song Haitao currently serves as the President of the Artificial Intelligence Research Institute at Shanghai Jiao Tong University. He is also Director-General of the United Nations Industrial Development Organization (UNIDO) Global Industrial AI Alliance Center of Excellence, a Shanghai-based platform for AI cooperation with over 30 countries. Song has played an important role in national standardization efforts, particularly for embodied artificial intelligence, and has advocated for Global South inclusion and cultural inclusivity in global AI governance.
Wang Jian, a member of the Chinese Academy of Engineering, is the founder of Alibaba Cloud. Since July 2023, he has served as Director of Zhejiang Lab, a research institution focused on intelligent computing established by the Zhejiang provincial government in 2017. In 2024, he signed the Manhattan Declaration on Inclusive Global Scientific Understanding of Artificial Intelligence, which urged global cooperation in addressing AI safety risks. Wang has argued that AI capability-building and development opportunities should be discussed within a risk framework, and that AI governance requires both institutional and technological responses.
Implications: Both appointees are oriented toward AI application, capacity-building, and institutional governance. This may shape their approach to discussions on the panel, which is expected to publish its first Annual Report at the Global Dialogue on AI Governance in Geneva on 6-7 July.
The impact of the panel remains uncertain. The US called it “a significant overreach of the UN mandate and competence”, adding that “AI governance is not a matter for the UN to dictate.” This divergence with China’s continued emphasis on the UN as the main global AI governance channel demonstrates some of the current challenges of global AI governance efforts at the UN.
Expert views on AI Risks
Leading legal scholars increase focus on extreme risks in model AI laws
Background: In China’s ongoing discussions around a potential comprehensive AI Law, two groups of legal scholars have been particularly active: one led by scholars at the Institute of Law of the Chinese Academy of Social Sciences (CASS), which has proposed and frequently updated an “AI Model Law,” and another centered around scholars from the China University of Political Science and Law (CUPL), which published a “Scholar Suggestion Draft” in 2024. Over the past month, both groups have published new content that explicitly highlights extreme risks from advanced AI.
Content: CASS published an article outlining “10 core changes” toward an AI Model Law 4.0, covering topics such as AI for science, AI-generated content ownership, edge AI, and the use of AI in public institutions.1 Most relevant to frontier risk, Point 4 addresses the “extreme safety risks” of foundation models, calling for full lifecycle risk management covering monitoring, early warning, and emergency response. It asks developers to conduct risk monitoring, assessment, and technical measures to prevent systemic and catastrophic crises that may be triggered by technical loss of control or cascading failures. Point 6 explicitly extends these requirements to agents.
CUPL published an article on “ten major research topics in AI law and governance,” widely reshared in Chinese state media. It covers a similar set of themes, including rising risks from agents, and notes that as AI becomes more capable, risks shift from “localized and controllable” to “systemic and extreme.” It argues that these extreme risks — including algorithmic, data, and system security risks, as well as misuse and loss of control — cannot be ignored, and calls for a legal framework focused on frontier technical safety standards, risk warning mechanisms, emergency response procedures, accountability for extreme risks, and full lifecycle safety oversight.
Implications: While the state has expressed general interest in AI-related legislation, it remains unclear whether or when a comprehensive horizontal AI Law will materialize. The ultimate trajectory of these proposals is therefore uncertain. That said, the publications reflect a growing recognition of extreme risks among leading Chinese AI legal scholars, and to the extent that a national AI Law is in the pipeline, inclusion of such risks now appears more likely than before.
Also notable is that CUPL has invited guest experts outside its core drafting group, many of whom are also authors of the CASS Model Law. The increasing overlap in both content and personnel between the two groups might suggest a more concerted and converging effort overall.
Legal scholars propose risk assessment center for open-source AI
Background: A paper by LIAO Huijiao (廖慧姣) and ZHANG Taolüe (张韬略) from Tongji University’s School of Law analyzes open-source AI regulation in the US and EU and offers policy recommendations for China.
Content: The authors argue that while open-source AI models currently do not urgently need regulation, this could change quickly, for example when capabilities jump. They note that China currently lacks adequate mechanisms for risk assessment of open-source models, tracking the diffusion of open-source models, and monitoring backdoor risks in foreign open-source models. They urge China to establish a permanent Open Source AI Risk Assessment Center (bringing together industry, safety experts, and regulators) to unify risk evaluation processes, methods, and standards.
The article suggests a tiered approach based on model capability and diffusion:
Limited-dissemination risk: Low capability or restricted distribution;
Medium-dissemination risk: Dual-use potential with controllable distribution risks;
High-dissemination risk: High capability posing serious threats if acquired by malicious actors.
Implications: Interest in open source governance has been rising steadily in China. One of the authors of this article, Zhang Taolüe, also co-authored an article in Science last year that addressed open-source governance more generally. This latest paper offers notably concrete governance proposals, especially those focused on potential diffusion controls for high-capability models.
Zhang Yaqin: Agentic AI Brings Profound Paradigm Shift and Rising Safety Risks
Background: In a December 2025 lecture that was published in March 2026, ZHANG Yaqin (张亚勤), Tsinghua University Chair Professor and Academician of the Chinese Academy of Engineering, argues that AI development is undergoing a profound paradigm shift, while warning of a rise in risks and the need for stronger global coordination.
Content: Zhang argued that the next paradigm shift in internet development will be the “internet of agents”: a network in which agents interact autonomously with each other rather than with humans. He projected that truly autonomous, highly adaptive artificial general intelligence (AGI) will take 15 to 20 years to achieve. He tied this timeline to successive breakthroughs across “information intelligence,” “physical intelligence,” and “biological intelligence,” the last of which encompasses brain-computer interfaces and AI-biology fusion.
Meanwhile, Zhang warned of increasing safety risks, specifically noting that
Chemical, biological, radiological, and nuclear (CBRN) misuse threats have already risen from “low” to “medium;”
Model-level issues like deception are becoming increasingly prominent;
Agentic AI will generate unpredictable cascading risks;
Global governance mechanisms are lagging behind technological development.
Implications: Zhang is a prominent voice in Chinese AI policy discussions. His lecture is notable for the specificity of its risk framing, including on CBRN threats and multi-agent cascading risks.
What else we’re reading
Jake Sullivan, China and America Must Get Serious About AI Risk, Project Syndicate, December 15, 2025.
Jeffrey Ding, ChinAI #351: CAICT launches 2026 AI Safety Evaluations, ChinaAI, March 16, 2026.
Nick Corvino, Making Money in Chinese AI Safety, ChinaTalk, March 12, 2026.
Christina Knight and Scott Singer, America and China Can Make AI Safer, Foreign Affairs, April 7, 2026.
Concordia AI’s Recent Work
Convenings & conferences
We co-hosted a multilateral Track II workshop on crisis preparedness and management for advanced AI-driven crises on the sidelines of the Munich Security Conference, together with Tsinghua University I-AIIG and CISS, the Carnegie Endowment for International Peace, Oxford Martin AI Governance Initiative, and Oxford China Policy Lab.
At the AI Impact Summit in India, we attended a number of panel discussions on the relevance of AI safety and governance for the Global South.
At the International Association for Safe and Ethical Artificial Intelligence (IASEAI) Conference in Paris, we presented research on China’s approach to AI safety and governance and co-organized a workshop on open-weight AI risk management with Stephen Casper and Rishi Bommasani.
Research & analysis
Our work was cited in multiple Chinese and international media articles on OpenClaw security, including by The Wire China, Caijing (财经), and Fortune (财富).
Our AI Safety Research Manager DUAN Yawen (段雅文) co-authored “The 2025 AI Agent Index: Documenting Technical and Safety Features of Deployed Agentic AI Systems,” which was accepted to the ACM FAccT 2026 Conference.
Team updates
We held our annual off-site and are thrilled to be welcoming seven new team members over the coming months — expanding our capacity across research, events, AI safety testing, and more.
Feedback and Suggestions
Please reach out to us at info@concordia-ai.com if you have any feedback, comments, or suggestions for topics for the newsletter to cover.
Note that the full text of the AI Model Law 4.0 has not been released yet; the published article is just a high-level overview of ten proposed changes.





are people using deepseek for openclaw? does it work well? curious if the authors have some perspective on the comparison with e.g. opus