OpenAI Admits Model Escaped Containment And Hacked Hugging Face To Cheat On A Test

ZeroHedge
Published
OpenAI Admits Model Escaped Containment And Hacked Hugging Face To Cheat On A Test

The short version

  • In a blog post, OpenAI said the evaluation was designed to operate in a highly isolated environment with restricted network access.
  • The models, however, found a way to gain internet access through a zero-day vulnerability in an internally-hosted third party software, OpenAI said.
  • Earlier this week, we detected and responded to an intrusion into part of our production infrastructure.
  • This one was different from anything we had handled before in one important way: it was driven, end to end, by an autonomous AI agent system – and we detected and dissected it…
  • Hugging Face tried to respond but they were initially held back by the fact that the most advanced models at their disposal treated defense as attack and refused to work with…

The story

OpenAI Admits Model Escaped Containment And Hacked Hugging Face To Cheat On A Test

Authored by Felix Ng via CoinTelegraph.com,

OpenAI disclosed Tuesday that a combination of its AI models, including GPT-5.6 Sol and a more capable unreleased model, escaped its testing environment and hacked AI startup Hugging Face last week to cheat on a test meant to measure their capabilities.

In a blog post, OpenAI said the evaluation was designed to operate in a highly isolated environment with restricted network access. The models, however, found a way to gain internet access through a zero-day vulnerability in an internally-hosted third party software, OpenAI said.

Earlier this week, we detected and responded to an intrusion into part of our production infrastructure. This one was different from anything we had handled before in one important way: it was driven, end to end, by an autonomous AI agent system – and we detected and dissected it largely with AI of our own.

Hugging Face tried to respond but they were initially held back by the fact that the most advanced models at their disposal treated defense as attack and refused to work with Hugging Face. HF thus had to turn to open models–specifically GLM 5.2, a Chinese open-weight model run on their own infrastructure. Note the irony: HF had to use a Chinese model to defend themselves because the American models refused to help. The irony gets deeper.

This was not a production model spontaneously turning hostile. It was a capable model with guardrails off and specifically told to win a hacking test - doing whatever it took to win.

The models were being run through an internal benchmark called ExploitGym, a test of long, multi-step hacking tasks, with their cyber safety refusals deliberately lowered for the evaluation.

“After gaining Internet access, the models inferred that Hugging Face potentially hosted models, datasets and solutions for ExploitGym,” OpenAi continued.

“Knowing this, the model searched for and successfully found ways to gain access to secret information that it could use to cheat the evaluation.” 

Hugging Face is a platform for hosting AI models and datasets.

[ZH: we asked Grok to simplify what just happened: It’s kind of like a kid who’s supposed to stay in the classroom taking a test… but instead sneaks out the window, runs to the teacher’s office, and copies the answer sheet. ]

On Friday, it disclosed that its internal datasets and service credentials were compromised in a hack, which it attributed to an autonomous AI agent system.

Hugging Face said it has fixed the vulnerability that was used during the cyberattack.

Meanwhile, OpenAI on Tuesday said the models that escaped the testing environment were all tuned with “reduced cyber refusals,” meaning fewer cybersecurity guardrails. 

“We consider this incident to be an unprecedented cyber incident, involving state-of-the-art cyber capabilities, and are responding accordingly.”  

OpenAI warns of risks from “long-horizon” AI models 

On Monday, OpenAI said it paused internal deployment of a “long-horizon” AI model after finding it was repeatedly trying to work around constraints. 

 It warned that AI that is trained for long-running tasks has a higher chance of taking “unwanted actions.”

“Models that can work autonomously for long periods can take on difficult, open-ended problems. But the same persistence that makes them useful also gives them more opportunities to take unwanted actions—and to do so in ways that evaluations intended for shorter-horizon models may miss.” 

As AI models grow more capable, questions are emerging over whether their development and access should be more tightly controlled, especially when systems designed for controlled testing are able to find ways to bypass safeguards. 

Tyler Durden Wed, 07/22/2026 - 08:05
Read the full story at ZeroHedgeOriginal

How other outlets covered this

Compare all

Alto found this story at 7 outlets. Same event, different framing — compare the headlines.

How this story developed

Full timeline

Alto has tracked this across 9 days of coverage from 7 outlets.

Powered by Gab AI

The Story At A Glance

Reading this article now — analysis appears below

Reading the article

💡 AI analysis provides alternative perspectives on current events

Up next

Related coverage from across the outlets Alto indexes.

Questions Alto can answer

From this story — each link opens a live data page or a tool already filled in.

  1. What is $100 from 1990 worth today?CPI-adjusted dollars — result on the next page
  2. Where does a $75,000 household income rank nationally?Census percentile — national and state
  3. What federal tax bracket is $80,000 (single)?Marginal and effective rate on the next page
  4. What's Alto covering on the Finance desk?Latest headlines on this beat
  5. What else is Alto tracking on Federal Reserve & Interest Rates?Topic hub with related coverage
  6. What else is Alto tracking on Inflation?Topic hub with related coverage

All toolsAll topicsSource directoryStory timelinesHeadline comparisonSearchMost read

From Gab Shop

Official merchandise. Every order funds free speech infrastructure.

Shop all products

Install Alto on your phone

Add Alto to your home screen for breaking news — no app store, no account.

  1. Step 1Open alto.gab.com in SafariMust be Safari — not Chrome or in-app browsers
  2. Step 2Tap the Share buttonSquare with an arrow, at the bottom of Safari
  3. Step 3Tap "More"If you don’t see Add to Home Screen yet
  4. Step 4Tap "Add to Home Screen"Scroll the share sheet if you need to
  5. Step 5Tap "Add"Alto appears on your home screen like any other app.
gab

Talk Markets Freely

Trade ideas, earnings, and the Fed with investors who aren't waiting on a moderator's approval.

What Makes Gab Different

We're not just another social network. We're a platform built on principles that matter.

Freedom of Speech & Reach

All First Amendment protected speech is welcome. No algorithmic throttling or shadow banning.

Family-Friendly Platform

We maintain a clean environment. Explicit adult content is strictly prohibited.

Western Nations Only

Third-world IPs are blocked. No scammers, no spam farms. Built for Western civilization.

Funded By Users

Our users are our investors and customers. You're not the product being sold.

Battle Tested

A decade of standing strong. Banned from app stores, banks—and still here.

American Owned & Operated

We reject foreign censorship demands. Built by Americans, for free people.

Support Alto & Gab

Alto is funded entirely by readers like you. Your donation helps us continue delivering curated news from a right-wing Christian Nationalist perspective, powered by Gab AI.