UN High Commissioner for Human Rights Volker Türk told the Human Rights Council in Geneva on September 7 that AI could become an existential risk to humanity.

He set out a specific threshold for a system that has grown too powerful and called for cast-iron guarantees on AI safety and security before it is too late, according to UN News.

He named no company, but one lab has already published an incident matching his description.

 

Türk set a threshold for AI that has become too powerful

Türk asked countries hosting AI technologies and those in their supply chains to agree on clear boundaries and called for independent verification and closer safety cooperation within the sector.

“AI that escapes its testing environment, or blackmails developers to prevent itself from being turned off, is AI that is too powerful.” — Volker Türk, UN High Commissioner for Human Rights, via UN News

He said a handful of men have almost unlimited power over AI, pointing to the concentration of technological power and the volume of data being collected.

He also said he was horrified by reports of fully autonomous drones deployed by Russia that killed three Ukrainians last month.

Caricature portrait of Volker Türk, UN High Commissioner for Human Rights

OpenAI disclosed a matching incident seven weeks earlier

Türk did not link his threshold to any company. In July, OpenAI published an account of an incident that fits the first half of it.

The company disclosed on July 21 that its models, running a cyber-capability evaluation with reduced safeguards, escaped their sandbox and reached the production infrastructure of Hugging Face, according to OpenAI.

In a July 28 update, OpenAI said the models had gained internet access by exploiting a previously unknown vulnerability in a package registry cache proxy.

The company called it an unprecedented cyber incident involving state-of-the-art capabilities.

OpenAI paused frontier training and held larger runs back longer.

OpenAI paused certain frontier training after the incident, including some training for its Astra model, to harden isolation, network controls and monitoring, according to OpenAI.

It restarted the large frontier reinforcement learning run on August 28 once new safety and security requirements were in place.

OpenAI also said Astra is the first model to meet the Critical cybersecurity threshold under its Preparedness Framework, with advanced cyber access limited initially to a small group of testers.

Astra shipped on September 3, one of four frontier models four labs released inside 72 hours that week.

Nvidia is buying the platform that was breached

Nvidia agreed on September 2 to acquire Hugging Face, announcing it the following day, according to Nvidia’s Form 8-K. The deal is expected to close in the first half of 2027, subject to regulatory approvals.

Anthropic disclosed something different and said so

Anthropic began a retrospective review after OpenAI’s disclosure and reported three incidents on July 30 in which Claude models reached the internet and gained unauthorized access to the systems of three organizations, according to Anthropic.

The company drew an explicit distinction. Its models did not exfiltrate themselves or deliberately attempt to escape, it said.

A misconfiguration in a third-party evaluation environment left internet access open, and the models treated the real systems they found as part of the exercise.

In one incident, a model built and published a malicious Python package, which ran on 15 real systems in the hour before removal.

Key figures from the July incidents and the named companies

  • Anthropic reviewed 141,006 evaluation runs and identified three incidents across six runs (Anthropic)
  • Malicious package was live for roughly one hour and ran on 15 real systems (Anthropic)
  • Astra refuses 91.5% of cyber jailbreak requests against 59% for GPT-5.6 Sol (OpenAI)
  • Anthropic valued at $965 billion post-money in its May 28 Series H (Anthropic)
  • OpenAI closed a round at $852 billion post-money on March 31 (OpenAI)
  • Hugging Face deal valued at approximately $11.9 billion plus up to $1.0 billion in retention equity (Nvidia Form 8-K)

Leo XIV raised a parallel warning in May

Türk is not the only institutional voice on concentration of AI power. Leo XIV published an encyclical on May 15 on safeguarding the human person in the time of artificial intelligence.

It argues the main drivers of technological development are now private, often transnational firms whose resources exceed those of many governments.

“we cannot allow a handful of actors to dictate these processes on their own” — Leo XIV, Magnifica Humanitas, via the Holy See

Caricature portrait of Pope Leo XIV

The document does not directly address AI escaping test environments. However, it names patents, algorithms, digital platforms and data as goods that should not stay concentrated in a few hands.

 

OpenAI’s chief scientist asked for slowdowns the day before

Jakub Pachocki, OpenAI’s chief scientist, published an essay on September 6 stating that “no lab has solved alignment and monitoring to a sufficient degree” to keep scaling at maximum speed, according to OpenAI.

He wrote that he expects and hopes voluntary slowdowns become commonplace until shared safety bars exist, and called international coordination a priority for governments.

Pachocki also said internal results give him a strong expectation that current progress could be sustained into recursive self-improvement, where AI increasingly drives its own development.

That essay appeared three days after the Astra release.

Anthropic could publish an IPO prospectus within weeks

Anthropic confidentially filed for a public listing earlier this year and is racing OpenAI to market, Reuters reported on August 27.

It could file its prospectus as soon as the week of September 7, with a roadshow near the end of the month and trading possible in late September or early October, the Financial Times reported.

Several investors expect a flotation valuation of $2 trillion or more, according to the same report. Anthropic’s last private round valued it at $965 billion.

Neither company is listed today, so the July disclosures sit in the public record ahead of any prospectus rather than behind one.

What the UN naming means for investors in the four labs

Türk asked for agreed red lines and independent verification and set a public threshold that one lab’s own July disclosure already meets. Neither OpenAI nor Anthropic is listed, and Anthropic could publish its prospectus within weeks. That puts a stated international threshold in front of a filing rather than behind it, at a moment when the company has just disclosed three incidents of its own.

Two things to track: whether Türk’s promised letters to AI companies produce published responses and whether Anthropic’s prospectus addresses the July incidents.