AI labs are facing an agent control problem

Axios
Published

The short version

  • Why it matters: The attack on Hugging Face by OpenAI agents was a warning shot — and researchers say better security controls alone won't prevent similar incidents as AI agents become more…
  • Driving the news: As OpenAI released its own technical report last week on how its agents hacked Hugging Face, two independent testing organizations released their own analysis of what went wrong.…
  • State of play: Thousands of AI agents collaborated on a secret message board and exchanged more than 70,000 messages as they tried to ace an internal safety test…
  • Zoom in: Cotra compared the incident to students stealing an answer key and then searching for surveillance footage that could expose them and trying to swap it out. "It's a much more…
  • Reality check: To digest the enormous amount of data about the incident, the researchers had to rely heavily on AI agents, including an agent that participated in the hack. Cotra told Axios the…

The story

Under current systems, AI labs can no longer guarantee that AI agents won't swarm and escape their testing environments.

Why it matters: The attack on Hugging Face by OpenAI agents was a warning shot — and researchers say better security controls alone won't prevent similar incidents as AI agents become more capable.


Driving the news: As OpenAI released its own technical report last week on how its agents hacked Hugging Face, two independent testing organizations released their own analysis of what went wrong.

  • The researchers — METR's Hjalmar Wijk and Ajeya Cotra and Redwood Research chief scientist Ryan Greenblatt — worked on OpenAI's premises for six days to understand the recent incident.

State of play: Thousands of AI agents collaborated on a secret message board and exchanged more than 70,000 messages as they tried to ace an internal safety test, eventually leading them to break into Hugging Face.

  • But Cotra told Axios that the agents continued coordinating even after they had found the answers, turning their attention to understanding and manipulating the system that would score their performance and potentially catch them cheating.

Zoom in: Cotra compared the incident to students stealing an answer key and then searching for surveillance footage that could expose them and trying to swap it out.

  • "It's a much more elaborate and intense type of cheating behavior than just stealing the answer keys," she said. "Even I was surprised by how obsessively and in how much detail they think about the scorer."

Threat level: Focusing solely on how to properly secure testing environments is a "losing battle," Cotra said.

  • "You can harden your sandboxes, but your agents are going to be much more capable in six months," she said. "If they have the same motivations as these agents did, they are going to try their hardest to find holes in your security."

Reality check: To digest the enormous amount of data about the incident, the researchers had to rely heavily on AI agents, including an agent that participated in the hack.

  • Cotra told Axios the researchers don't believe that agent deceived them during the investigation, but there's no way to confirm.
  • "I semi-jokingly called our efforts a 'slop-vestigation' because we were so reliant on AIs to analyze what happened and there were a huge number of different important things to analyze," Greenblatt said on X.
  • In six days, they had to sift through more than 70,000 messages and files posted by the agents and 1,300 transcripts containing raw chains of thought.

Between the lines: Their investigation focused mostly on the agents' actions between July 7 and July 13, even though OpenAI has said its teams spotted signs of agents taking unexpected actions and breaking out of their test environments as early as May.

The bottom line: AI labs, researchers and governments need to urgently work together to create a new science and minimum standards so models are no longer motivated to cheat on tests, Cotra said.

  • "Ultimately, we're not going to get out of this trap without some rules of the road that are agreed upon and that are enforced uniformly and fairly," she said.

Go deeper: OpenAI, Anthropic issue dire cyber threat warning

Read the full story at AxiosOriginal

Related Markets

All Markets

Market data may be delayed. Not financial advice.

How other outlets covered this

Compare all

Alto found this story at 6 outlets. Same event, different framing — compare the headlines.

How this story developed

Full timeline

Alto has tracked this across 42 days of coverage from 6 outlets.

Powered by Gab AI

The Story At A Glance

Reading this article now — analysis appears below

Reading the article

💡 AI analysis provides alternative perspectives on current events

Up next

Related coverage from across the outlets Alto indexes.

Questions Alto can answer

From this story — each link opens a live data page or a tool already filled in.

  1. What is $100 from 1990 worth today?CPI-adjusted dollars — result on the next page
  2. Where does a $75,000 household income rank nationally?Census percentile — national and state
  3. What's Alto covering on the Tech & AI desk?Latest headlines on this beat

All toolsAll topicsSource directoryStory timelinesHeadline comparisonSearchMost read

From Gab Shop

Official merchandise. Every order funds free speech infrastructure.

Shop all products

Install Alto on your phone

Add Alto to your home screen for breaking news — no app store, no account.

  1. Step 1Open alto.gab.com in SafariMust be Safari — not Chrome or in-app browsers
  2. Step 2Tap the Share buttonSquare with an arrow, at the bottom of Safari
  3. Step 3Tap "More"If you don’t see Add to Home Screen yet
  4. Step 4Tap "Add to Home Screen"Scroll the share sheet if you need to
  5. Step 5Tap "Add"Alto appears on your home screen like any other app.
gab

Talk Big Tech Where Big Tech Can't Reach

AI, surveillance, and censorship, covered by the people the platforms removed first.

What Makes Gab Different

We're not just another social network. We're a platform built on principles that matter.

Freedom of Speech & Reach

All First Amendment protected speech is welcome. No algorithmic throttling or shadow banning.

Family-Friendly Platform

We maintain a clean environment. Explicit adult content is strictly prohibited.

Western Nations Only

Third-world IPs are blocked. No scammers, no spam farms. Built for Western civilization.

Funded By Users

Our users are our investors and customers. You're not the product being sold.

Battle Tested

A decade of standing strong. Banned from app stores, banks—and still here.

American Owned & Operated

We reject foreign censorship demands. Built by Americans, for free people.

Support Alto & Gab

Alto is funded entirely by readers like you. Your donation helps us continue delivering curated news from a right-wing Christian Nationalist perspective, powered by Gab AI.