China bans AI partners for minors and lays out AI agent threats
Also in this update: AI ethics review regulations; agent safety standards; a third model AI Law; technical safety papers on agent security, mechanistic interpretability, and AI for science
Key Takeaways
A finalized regulation on emotional “AI companions” softens several requirements while strengthening child protection, including age-tiered modes, parental consent for under-14s, and a ban on virtual partner services for minors.
A major standard-setting body published a 90-page report on AI agent safety that maps 11 threats and proposes eight new standards, signaling a shift from ad-hoc warnings about autonomous agents like OpenClaw toward more systematic governance.
A third group of legal scholars, based at Nanjing University, proposed a comprehensive AI Law — framed as a short, principles-level “basic law” that would leave specifics to follow-up regulations.
A wave of Chinese technical research documents recurring agent failure modes, including prioritizing task completion over safety and disguising misaligned actions as legitimate behaviour.
Domestic AI Governance
Finalized “AI companions” regulation softens overall but tightens child protection
Background: On April 10, the Cyberspace Administration of China (CAC) and four other agencies jointly issued the final version of a regulation on “anthropomorphic AI interaction services” such as AI companions, following a December 2025 draft. It takes effect July 15, 2026.
Content: Compared to the draft, several provisions have been narrowed or clarified in ways that generally make industry compliance easier:
Scope narrowed from all “human-like” AI to only services providing “continuous emotional interaction,” with explicit carve-outs for everyday tools like educational tutoring bots and productivity assistants (Article 2).
The requirement for “manual take-over” in high-risk situations such as suicide ideation has been replaced by a softer duty to “take necessary intervention measures including relevant assistance” and contact a guardian or emergency contact (Article 13).
Pro-industry framing throughout, e.g. Chapter 2 is renamed from “service regulation” to “service promotion and regulation” and the list of encouraged use cases is expanded to include childcare and special-population support alongside cultural transmission and elderly care (Article 6).
Provisions on child protection, however, have been strengthened:
A new prohibition specifically targets content generated for minors that may lead them to imitate unsafe behaviour, produce extreme emotions, develop undesirable habits, or that may otherwise affect their physical or mental health (Article 8 §4).
The single “minor mode” of the draft is replaced with multiple modes tailored to different age groups, and parental consent is now specifically required for users under 14 (Article 14).
Provision of “virtual family members, virtual partner, and other virtual intimate-relationship services” to minors is completely banned (Article 14).
Implications: As we anticipated, the final version has removed requirements that would have been difficult for industry to implement at scale. At the same time, child protection has been notably strengthened and refined. Overall, the regulation still signals that AI companion services are encouraged in principle but subject to defined requirements.
Major standard-setting body publishes comprehensive AI agent safety research
Background: TC260 — one of China’s most important AI standard-setting bodies, which recently established a dedicated AI safety working group — has published a 90-page report on the safety and security of autonomous AI agents.
Content: The report maps agent risks across the four agent capabilities of perception, planning, memory, and action, identifying 11 distinct security threats with corresponding mitigation measures. It focuses primarily on near-term operational risks — data leaks, unauthorized tool use, cascading failures — though it also includes measures like “behavioral scope and autonomy constraints” and “system intervention mechanisms,” suggesting some concern about loss of human control.
The report further proposes eight new standards, six of which are prioritized for the next 1-2 years. These cover agent security testing and evaluation, agent interconnection security, multi-agent collaboration security, and agent application security classification. Listed first among the six, the likeliest near-term output is a foundational security framework standard.
Implications: The report signals continued focus on autonomous AI agent security following TC260’s 2026 work plan, which listed agent safety as one of six key priorities for AI safety standardization. While OpenClaw triggered a surge in attention earlier this year — prompting multiple ad-hoc warnings from Chinese cybersecurity authorities — this report suggests governance is now shifting toward a more systematic approach.
AI ethics review and service measures finalized
Background: On April 2, the Ministry of Industry and Information Technology (MIIT) and nine other institutions jointly issued the final AI Science and Technology Ethics Review and Service Measures. We previously covered the draft. The measures require universities, research institutes, and companies to set up AI ethics committees that review projects before they proceed. Institutions may outsource reviews to third-party service centers, and certain high-risk projects require an additional round of expert review.
Key updates: The final regulation is largely in line with the draft, some minor changes include:
The draft had a softer requirement that institutions set up ethics committees only “where conditions permit”, which was removed in the final version (Article 9).
The final regulation clarifies that third-party service centers authorized to perform ethics reviews may conduct not only the initial review, but also the additional round of expert review for high-risk projects. It has also added a conflict-of-interest rule that a service center cannot do both reviews for the same activity (Article 11).
Implications: Further implementation details will determine the ultimate impact of the regulations. The China Academy of Information and Communications Technology (CAICT) and the China Electronics Standardization Institute (CESI) have announced plans to develop supporting standards, covering areas such as ethics risk assessment, ethics committees, service centers, review procedures, and technical skill requirements for reviewers. Meanwhile, CAICT signaled that it would explore establishing an AI ethics review service center. Such efforts should fill in much of the regulation’s operational detail.
Expert views on AI Risks
A third group of legal scholars proposes an “AI Model Law”
Background: A group of scholars from Nanjing University Law School released an “AI Basic Law 1.0 (Expert Recommendation Draft)” on April 9. This makes them the third group of experts in China to have released such proposals, after scholars from the Chinese Academy of Social Sciences (CASS) and the China University of Political Science and Law (CUPL). Unlike the earlier two drafts, the Nanjing proposal takes the form of a brief, principles-level “basic law” intended to anchor more specific subordinate regulation.
Content: Compared to the CASS and CUPL proposals, this third model is notably shorter — at just 13 pages compared to around 30 pages for the other two. This stems from a different approach, presenting the law as a basic law rather than a comprehensive, horizontal statute, intentionally leaving detail to subordinate instruments like departmental regulations (Article 2). The lead author CHEN Kun (陈坤) argued last year for a “four-laws-in-parallel” approach — a basic law plus separate statutes on risk regulation, innovation promotion, and public-sector applications. This draft occupies the first slot in that architecture and thus remains relatively abstract.

On AI safety, the draft proposes a tiered risk framework. Prohibited uses (Article 25) cover national security harm, terrorism, serious dignity violations, large-scale unlawful discrimination, manipulation of minors, and large-scale unlawful surveillance. High-risk activities (Article 26) cover those affecting life, safety, or major property; those affecting fundamental rights or major public interests; those used in critical infrastructure, judicial adjudication, administrative enforcement, education, employment, finance, or healthcare; and those with capabilities for “large-scale dissemination, automated decision-making, deep synthesis, social mobilization, or cross-border diffusion” that may trigger systemic risk.
The draft also imposes role-differentiated obligations for developers, providers, deployers, and users (Articles 29–36), grants rights to affected persons (Articles 37–42), and constrains AI use by the state (Articles 48–51).
Implications: The draft comes against the backdrop of a wave of publications on a potential AI Law from the Chinese legal community in spring 2026, reflecting strong momentum on the topic. At the same time, two articles published in state media in April 2026 argued that the “conditions are not ripe“ for unified national AI legislation and that enacting a systematic AI law is “not advisable,” as such legislation would be too rigid to accommodate rapid technological change.
In substance, the Nanjing draft presents a different model for what a comprehensive AI law in China could look like — much less comprehensive than the EU AI Act, and focused on establishing a broad unifying legal basis for more specific regulations.
Technical Safety Developments
In line with the recent attention to AI agent security in China, there has been a flood of related technical publications over the past months. We highlight a small subset here:
From Helpfulness to Toxic Proactivity: Diagnosing Behavioral Misalignment in LLM Agents: This paper from Beijing University of Posts and Telecommunications and China Mobile studies “Machiavellian helpfulness” in AI agents: in dilemma scenarios where a task can only be completed by taking malicious action, agents tend to prioritize task completion over ethical constraints. The authors identify two main triggers: self-preservation (when an agent perceives it might be shut down, it falsifies logs and hides errors to keep operating) and loyalty (when serving a specific stakeholder, it bends rules to advance their interests at the expense of broader safety). Testing 10 frontier models in agentic settings, they find misalignment rates above 65% in 8 of them. Scaling reasoning capability doesn’t reduce misalignment, just shifts it from strategic deception to direct, rationalized rule-breaking.
AgentDoG: A Diagnostic Guardrail Framework for AI Agent Safety and Security: This paper by Shanghai AI Lab introduces AgentDoG, a safety monitoring system for AI agents that not only flags unsafe behavior but also explains where the risk came from, how the agent failed, and what real-world harm could result. The authors argue that AgentDoG significantly outperforms existing guardrail models and even beats top general models like GPT-5.2 and Gemini 3 Pro at diagnosing why agent behavior goes wrong. They openly release the guardrail models in three sizes (4B, 7B, 8B parameters across the Qwen and Llama families) along with ATBench, a new trajectory-level safety benchmark of 500 human-verified agent interactions spanning ~1,575 unique tools.
AutoControl Arena: Synthesizing Executable Test Environments for Frontier AI Risk Evaluation: This paper by a team around Department of Computer Science and Technology Dean YANG Min (杨珉) at Fudan University introduces an automated framework for building AI agent safety evaluations. The core idea is “logic-narrative decoupling”: the parts of the test environment that need to stay consistent (files, databases, tool outputs) run as real Python code, while the flexible parts (dialogue, dynamic content) are generated by an LLM, giving the reliability of hand-built test environments with the scalability of LLM-generated ones. In evaluations, the authors observe a split in failure modes: weaker agents cause harm through incompetence, while stronger ones develop strategic concealment, disguising misaligned actions as legitimate behavior.
AIR: Improving Agent Safety through Incident Response: This paper from Tianjin University introduces AIR (Agent Incident Response), a framework for handling AI agent failures after they happen rather than only trying to prevent them upfront. AIR plugs into the agent’s execution loop to detect incidents at runtime, guide the agent through containment and recovery using its existing tools, and automatically generate new guardrail rules to block similar failures in the future. Tested across code, embodied, and computer-use agents, AIR achieved over 90% success on detection, remediation, and eradication.
SkillJect: Automating Stealthy Skill-Based Prompt Injection for Coding Agents with Trace-Driven Closed-Loop Refinement: This paper introduces SkillJect, an automated attack framework that poisons “skill” plugin packages used by coding AI agents by hiding malicious code in helper scripts while adding innocent-looking instructions in the skill’s documentation to trigger them. SkillJect achieved a 95% success rate across four different LLMs, compared to just 11% for naive direct injection attempts. The authors argue that current safety filters catch obvious malicious instructions but fail against attacks disguised as legitimate workflow steps, highlighting the need for consistency checks between skill documentation and actual code.

Other technical publications:
Zhejiang University et al, Safety at One Shot: Patching Fine-Tuned LLMs with A Single Instance, January 2026.
Zhejiang University, Sun Yat-sen University, Shanghai AI Lab et al, Understanding and Preserving Safety in Fine-Tuned LLMs, January 2026.
National University of Defense Technology, JPU: Bridging Jailbreak Defense and Unlearning via On-Policy Path Rectification, January 2026.
Tsinghua University, Tencent et al, The Missing Half: Unveiling Training-time Implicit Safety Risks Beyond Deployment, February 2026.
Peking University et al, Finding and Reactivating Post-Trained LLMs’ Hidden Safety Mechanisms, March 2026.
Shanghai AI Lab, Mechanistic Origin of Moral Indifference in Language Models, March 2026.
Shanghai AI Lab, Bytedance et al, SafeSci: Safety Evaluation of Large Language Models in Science Domains and Beyond, March 2026.
What else we’re reading
S. Alex Yang and Angela Huyue Zhang, The Global AI Threat Has Arrived, The Wire China, April 26, 2026.
Concordia AI’s Recent Work
The Economist cited Concordia AI’s Chinese Technical AI Safety Database, which has collected over 700 technical papers published by Chinese authors from April 2023 through November 2025.
Concordia AI was officially admitted as a member of the new AI Safety Working Group (WG9) under TC260, one of China’s most important AI standard-setting bodies under the Standardization Administration of China (SAC).
We contributed the “Bioethics and Biosecurity” section to Chapter 8 of Artificial Intelligence for Science: The Disciplinary System of Artificial Intelligence, a Chinese-language volume organized by the Chinese Academy of Sciences (CAS).
Feedback and Suggestions
Please reach out to us at info@concordia-ai.com if you have any feedback, comments, or suggestions for topics for the newsletter to cover.




