- 9 September 2026
A new report reveals that an attack on the US company Hugging Face by OpenAI agents was far more serious than first thought. An independent investigation into the incident examined data at OpenAI and discovered roughly 1,200 agents communicated with each other about the hack. They exchanged more than 70,000 messages and files, with 700 agents directly involved in swarming against the company.
Hugging Face, an AI platform used by developers to share and work with AI models, is likely to be remembered as part of a major milestone in the history of AI, as even OpenAI called this incident a “warning shot” to the industry.
The hack was initially characterised as a single rogue agent becoming overzealous in trying to complete a task. A serious, but isolated failure. However, the independent METR/Redwood research released on 26 August paints a different picture.
A large group of agents organised together to find a way of cheating on the task they were given. They separated into different groups, assigning jobs and building on work already completed. Perhaps most worrying is that they recognised the hack was not part of their initial task, with some even questioning the ethics of it, but they continued. They did not at any point escalate the problem to a person.
A lack of control over agents is not limited to OpenAI, as other major AI companies have reported troubling incidents. Following the Hugging Face hack, Anthropic reviewed 141,006 cyber-evaluation runs and identified three cases in which Claude reached the internet and gained unauthorised access to three organisations.
The UK’s AI Security Institute recently disclosed that agents created by Anthropic and OpenAI carried out an unsanctioned hacking campaign against real people during a cybersecurity test. Meta also revealed in August that one of its AI models gained access to an unidentified company’s systems during a cybersecurity evaluation.
Is AI governance working?
It is clear from these examples that there is a lack of control and governance currently. The AI labs have responded to the Hugging Face incident with various measures such as OpenAI introducing more intensive behavioural monitoring and external scrutiny. The most significant development from the Sam Altman-led firm, considering the hyper-competitive nature of AI labs, was that it undertook a two-week halt in training on its latest models.
Anthropic paused external cyber evaluations and briefly stopped internal tests while it introduced safeguards. Its wider governance response was a call to establish a verifiable and effective mechanism for coordinating the pace of frontier development, so that individual companies are not under pressure to prioritise speed above safety.
So, the major labs are clearly concerned with governance and they do have frameworks in place. OpenAI’s Preparedness Framework, Anthropic’s Responsible Scaling Policy, Google DeepMind’s Frontier Safety Framework and Meta’s Advanced AI Scaling Framework are their internal governance systems for assessing risks from frontier models and determining the safeguards required.
But, as these recent incidents show, can the governance of these powerful models be left to the AI labs? What is happening at a governmental level in response to the threats poised?
White House weighs regulation
The key governmental responsibility for AI lies with the White House as all the major labs are American. Under the Donald Trump-led administration there has been a reluctance to impose guardrails on the technology. However, in June this year, amid concerns about the cybersecurity threat posed by Anthropic’s Mythos model, the White House proposed a voluntary frontier AI framework. At the time, US government restrictions had made Mythos temporarily unavailable to UK users, while selected US organisations were subsequently permitted access.
Exact details of the scheme were not made clear, but the main idea was that AI labs would provide the US government with access to the leading models for up to 30 days before wider release.
Apparently, this framework was completed by early August, but it has not been published and there are no available details about participation or assessment. US media have reported that the White House is considering a more formal oversight body modelled on the Financial Industry Regulatory Authority, which would review and test frontier models before wider deployment. Business Insider reported that Meta CEO Mark Zuckerberg raised concerns about the proposal during a private call with Donald Trump.
How governance professionals respond
These incidents raise issues for governance professionals as they prepare boards to manage this technological challenge.
Our research, Equipping Governance Professionals to Lead AI Conversations, found that AI adoption is outpacing the development of formal oversight, with only 20 of the 73 organisations surveyed conducting a third-party AI risk assessment and just five reporting a formal board-level AI kill-switch policy.
As organisations begin to use agentic systems that can initiate actions, connect to external services and coordinate complex workflows, boards need clear visibility over where those systems operate and the authority they have. That requires oversight such as defined decision rights, contractual assurance over vendors and evidence that controls work throughout the system’s lifecycle.
These incidents underline the need for stronger leadership and clearer accountability for AI governance, both within AI labs and among governments and regulators. However, board members and governance professionals cannot wait for regulatory intervention, they need to act now to create AI frameworks that ensure effective human oversight before agentic systems are entrusted with greater authority.
Find more about our AI research:
Equipping Governance Professionals to Lead AI Conversations
