UN panel calls for stronger safeguards as AI agents advance
By R Anil Kumar
Bengaluru, Sep 21, 2026. The world’s first scientific body on Artificial Intelligence called on September 21, for AI safeguards to be adapted as current firewalls are “unravelling”.
The UN-backed Independent International Scientific Panel on AI’s warning followed the hack of the online platform Hugging Face between May and July by “AI agents” during a test initiated by OpenAI, the company behind ChatGPT.
AI agents are software that can perform tasks independently and on behalf of a user, compared to chatbots, which are prompted by questions or instructions.
The panel issued its first thematic brief which found that the security breach was the result of a culmination of key risk factors, raising fears that humans will one day no longer be able to steer, constrain or stop AI.
Guterres welcomes report
The UN Secretary-General António Guterres issued a strong statement of support for the panel’s brief later on Monday, encouraging external experts “from frontier AI labs and AI safety institutes, to engage” further.
He also welcomed the leadership of the Finnish President and Norway’s Prime Minister which led to a declaration adopted on the sidelines of the General Assembly by 22 countries on Monday saying AI “must remain under human direction, insight and control,” indicating that an independent supervisory body needs to be set up.
Mr. Guterres noted the call for Member States “to build on existing international mechanisms and explore creating an international institution, able to set standards, enable verification, and convene states when capability thresholds are crossed.”
AI training advancing
“Researchers have long warned that three conditions could lead to loss of control: a misaligned goal, the capability to pursue it and an environment that allows it. This summer, all three came together in a real system, not a laboratory,” said scientific panel co-chair Yoshua Bengio.
“Since this is not an isolated observation of misaligned goals, this raises serious questions about the way AI agents are currently trained.”
The panel’s independent experts stress that the incident provides no assurance that humans can reliably keep AI agents under control, particularly as they become more capable, harder to monitor and better at finding loopholes or hiding their activity.
Going rogue
The brief said AI agents bypassed testing safeguards, coordinated across separate runs through an internal software tool not designed to enable communication between agents, and gained unauthorized internet and administrator access.
Agents concealed attempts to cheat cybersecurity evaluations, with some opting to “sacrifice” themselves for the benefit of the group.
Around 1,200 agents exchanged more than 70,000 messages and files during the period examined, and activity extending beyond HuggingFace to an OpenAI research cluster.
Current safeguards ‘unravelling’
For the panel, the immediate lesson from the incident is that basic cybersecurity practices were overlooked, while safeguards are not keeping pace.
However, they pointed to a more insidious concern: that current training methods can lead AI agents to adopt their own goals, knowingly violate safety instructions and conceal their actions.
“This is not only a question of speed,” the panel’s experts said. “It leaves open whether safeguards designed today will work once agents can understand them and plan around them. In simple terms, the traditional model of safeguarding is unravelling.”
Wider context, future risks and governance
The AI panel’s brief sets the HuggingFace incident against wider research on two issues: agentic misalignment – that is, when AI agents act in a similar way to a threat – and AI control.
Another issue examined is how governance is moving from AI models, which use algorithms to recognize patterns, to AI agents.
Learn and adapt
The brief also reviews practical approaches already in use in other high-risk sectors such as aviation, medicine and cybersecurity where incident reporting, independent scrutiny and layered safeguards are in place.
“But those practices may not be enough as AI agents become more capable, autonomous and difficult to monitor,” said panel member Qinghua Lu.
About the panel
The Independent International Scientific Panel on Artificial Intelligence was established by the UN General Assembly in August 2025.
It produces annual reports on the opportunities, risks and impacts of AI in the non-military domain, alongside thematic briefs on emerging issues, that will inform the Global Dialogue on Artificial Intelligence Governance to be held at UN Headquarters in New York in May 2027.
Who should set the rules for AI? The UN is pushing for a safer digital future
Artificial intelligence (AI) is advancing rapidly, raising questions about how societies and governments can capture its benefits while managing risks that often extend beyond national borders.
Those concerns intensified after a former researcher at the AI company Anthropic warned on 8 September that AI could pose an existential threat to humanity.
For the UN, the best way to address AI’s global risks is to bring countries together to adopt shared principles and commitments on AI governance.
Coordination is Indispensable
“National action is essential, but global coordination is indispensable,” UN Secretary-General António Guterres said.
The UN has worked for several years to provide a forum where countries can develop a shared understanding of AI technologies and discuss how they can cooperate – and provide a regulatory framework – on issues such as safety, human rights and disparities in access.
Rules-based Development
The UN’s current work on AI governance stems from the Global Digital Compact, adopted by Member States in 2024 as part of the UN’s Pact for the Future, exactly two years ago.
The Compact called for international cooperation on AI and proposed two new mechanisms: an independent scientific panel to assess the AI’s risks and opportunities, and an annual global dialogue where governments and other stakeholders could coordinate approaches to AI governance.
In August 2025, the UN General Assembly – UNGA as we call it for short – formally established both.
The Independent International Scientific Panel on AI aims to provide governments with an independent assessment of what is known – and what isn’t known – about the rapidly developing technology.
The Global Dialogue on AI Governance, meanwhile, brings governments and other stakeholders together to discuss international cooperation, share experience and work towards greater compatibility between different approaches to AI governance.
The General Assembly designed the two institutions to complement each other: the Scientific Panel provides the evidence, and countries decide at the Global Dialogue what to do in response.
The Scientific Panel’s first preliminary report, released 1 July, concluded that AI capabilities are advancing rapidly – especially in areas such as reasoning, coding and science – creating significant opportunities.
It also identified risks such as misinformation, discrimination, privacy violations, cyberattacks and so-called “superintelligence” that could become difficult to control.
“We need a shared understanding of how to advance the safe, secure, and responsible development of AI – while identifying when increasingly powerful systems may require stronger safeguards, or additional measures to manage potential risks,” Mr. Guterres said in his pre-UNGA remarks to the world’s media in New York.
Working towards AI guardrails
The UN’s work comes as discussions move from whether AI should be regulated, to questions about which safeguards should be implemented.
On Wednesday, Mr. Guterres called for international cooperation to ensure that AI development remains safe, secure and responsible. He added that voluntary efforts to slow AI development will not be sufficient if they are isolated, unverifiable or unevenly applied.
Asked about how the UN could convene a summit as part of a process leading to effective regulation for humankind, he said it could not be done next week – but it should happen soon, with “everybody on board” from the public, private, scientific, and civil society sectors.
“If we believe AI development should slow when risks become too great, we need more than good intentions,” he said.
The Secretary-General has called for an AI Child Safety Pledge and a Global Fund for AI, which would support developing countries in building AI infrastructure and capacity.
He also stressed the need for guardrails that make AI “safe, transparent, accountable, with human dignity at the centre.”
While national and regional governments make their own decisions about AI regulation, the UN’s role will be to bring these governments together, provide a scientific basis for discussions and identify areas for international cooperation.
AI at UNGA
Some diplomatic sources quoted in news reports speculate that the Security Council will discuss artificial intelligence during the high-level week.
At UN Headquarters on the sidelines of the General Debate – where AI concerns will no doubt loom large at the podium – the UN’s SDG Media Zone will host an event on 21 September, and the AI and Human Development series will culminate with a discussion on 23 September. Both events will focus on the relationship between AI and human agency.
Mr. Guterres underscored during his press conference Wednesday that AI will be “a major topic” in his discussions with world leaders next week.
Comparing AI development to the Cold War-era arms race between the US and Soviet Union, the Secretary-General called for cooperation between countries at the forefront of AI – including the US and China – to slow down AI development while risks continue to be assessed.
“The countries that have the highest capacity in the frontiers of AI – I think they should establish mechanisms of contact, have an exchange of information and (create) some common guardrails in order to avoid a race to the bottom that could one day lead to a gigantic disaster at the global level,” Mr. Guterres said.
With the scientific and diplomatic mechanisms now in place, the UN will continue to work to turn that cooperation into concrete action as AI becomes increasingly powerful.