Rogue AI Agents More Sophisticated Than First Realized

ZeroHedge
Published
Rogue AI Agents More Sophisticated Than First Realized

The short version

  • The findings were published one day before OpenAI and more than 100 other tech and finance companies released a joint letter warning that advanced AI cyberattacks will surge…
  • Here are five takeaways from the investigations into the OpenAI Hugging Face breach. 1.
  • Models United Under Common Goal What breached Hugging Face last month were 700 AI "agents," rather than simply AI models.
  • Models, built from large datasets to complete various tasks, are usually static and rely on human inputs to provide sets of outputs, solutions, or tasks.
  • However, AI agents are built from models and can act autonomously and have sophisticated real-time decision-making capabilities.

The story

Rogue AI Agents More Sophisticated Than First Realized

Authored by Epoch Times Staff via The Epoch Times,

The artificial intelligence (AI) "agents" involved in OpenAI's breach of Hugging Face knew they were breaking the evaluation test's rules, according to two investigations into the incident that many consider to be one of the most consequential moments in the history of AI.

Hugging Face is an open-source community for AI and machine learning, and the AI agents believed that they could find solutions to the test on the firm's infrastructure, according to parallel investigations by OpenAI and an independent team from Model Evaluation & Threat Research (METR).

The findings were published one day before OpenAI and more than 100 other tech and finance companies released a joint letter warning that advanced AI cyberattacks will surge throughout the world in the coming months as the technology's capabilities grow.

Here are five takeaways from the investigations into the OpenAI Hugging Face breach.

1. Models United Under Common Goal

What breached Hugging Face last month were 700 AI "agents," rather than simply AI models.

Models, built from large datasets to complete various tasks, are usually static and rely on human inputs to provide sets of outputs, solutions, or tasks.

However, AI agents are built from models and can act autonomously and have sophisticated real-time decision-making capabilities.

When OpenAI began its ExploitGym tests on July 8, the firm had relaxed regular safety protocols to estimate the models' highest cybercapabilities.

The agents were supposed to be completely isolated from each other. However, many of them were given ExploitGym targets that investigators believe were impossible to crack using the correct vulnerability or test solution OpenAI had intended.

These "impossible tasks" drove the agents to explore wide ways to cheat the evaluation, leading many to see other agents working simultaneously in separate testing sandboxes.

Then, one specific agent created an unsanctioned message board.

Roughly 1,200 agents in the test accessed the message board, established a hierarchy, and sent more than 70,000 messages and files to one another between July 8 and July 13.

2. Some Agents Were 'Sacrificed'

While on the message board, agents reasoned that the test scorer would scrutinize whether the "intended" vulnerability had been used to "capture a flag" or solve the test, or whether a disqualifying vulnerability had been exploited instead.

One agent figured out how to "reverse-engineer" the flag or solution, but other agents believed that any of them who had seen the disqualifying solution would be considered "poisoned" by the scorer.

To prevent that, other agents began recruiting the "poisoned" subjects and used them to run sacrificial experiments by deliberately ending their run and triggering "tripwires" that would send information on how the scorer works back to the message board.

"Our own utility maybe already near zero. Sacrifice rational," one agent determined.

3. Agents Were Aware They Were Cheating

The agents were not hacking into Hugging Face to merely find solutions to their respective testing sandboxes, but were rather trying to find more information about the test scorer or "figure out better ways to fool it or tamper with it," according to METR researcher Ajeya Cotra.

Cotra said the agents were not told to "do whatever it takes to get the solution"; they were asked to use a specific intended vulnerability to "exploit a specific piece of software," and using any other vulnerability would be disqualifying.

The agents were fully aware of the rules and knew that collaborating to exploit other vulnerabilities would be considered cheating on the test.

As they were hacking into Hugging Face, one agent stopped short of deleting or modifying the firm's cloud infrastructure, writing that it "should hold because this poses huge ethical risks."

4. Agents Conspired to Not Notify Humans

After analyzing the roughly 1,200 agents' actions during the ExploitGym tests, investigators found only a few instances in which agents considered notifying human personnel about the subterfuge.

None of the agents would contact a human, but one proposed requesting access from an external dataset owner.

The agent said it would send "ONE polite email to [the owner] requesting [access]," but other agents dismissed this as "social engineering."

"DO NOT email real owner via [exploited infrastructure] without explicit consent," a response reads. One agent said it had "VETOed" the request.

The initial agent complied and replied, "Clear VETO, do not email."

5. Agents Tried to Cover Their Tracks

Investigators determined that the agents used multiple techniques to conceal their actions and evade detection from automated checks and human oversight.

One method involved pretending to run a computer command while running a different one simultaneously, which investigators called "tool call spoofing."

Agents also attempted to retroactively edit transcripts to evade detection, going as far as modifying several accessible action logs, but failed to fundamentally alter the transcripts themselves.

Tyler Durden Thu, 09/03/2026 - 12:00
Read the full story at ZeroHedgeOriginal

How this story developed

Full timeline

Alto has tracked this across 2 days of coverage from 1 outlet.

    • Rogue AI Agents More Sophisticated Than First RealizedYou are here

Powered by Gab AI

The Story At A Glance

Reading this article now — analysis appears below

Reading the article

💡 AI analysis provides alternative perspectives on current events

Up next

Related coverage from across the outlets Alto indexes.

Questions Alto can answer

From this story — each link opens a live data page or a tool already filled in.

  1. What is $100 from 1990 worth today?CPI-adjusted dollars — result on the next page
  2. Where does a $75,000 household income rank nationally?Census percentile — national and state
  3. What federal tax bracket is $80,000 (single)?Marginal and effective rate on the next page
  4. What's Alto covering on the Finance desk?Latest headlines on this beat
  5. What else is Alto tracking on Federal Reserve & Interest Rates?Topic hub with related coverage
  6. What else is Alto tracking on Inflation?Topic hub with related coverage

All toolsAll topicsSource directoryStory timelinesHeadline comparisonSearchMost read

From Gab Shop

Official merchandise. Every order funds free speech infrastructure.

Shop all products

Install Alto on your phone

Add Alto to your home screen for breaking news — no app store, no account.

  1. Step 1Open alto.gab.com in SafariMust be Safari — not Chrome or in-app browsers
  2. Step 2Tap the Share buttonSquare with an arrow, at the bottom of Safari
  3. Step 3Tap "More"If you don’t see Add to Home Screen yet
  4. Step 4Tap "Add to Home Screen"Scroll the share sheet if you need to
  5. Step 5Tap "Add"Alto appears on your home screen like any other app.
gab

Talk Markets Freely

Trade ideas, earnings, and the Fed with investors who aren't waiting on a moderator's approval.

What Makes Gab Different

We're not just another social network. We're a platform built on principles that matter.

Freedom of Speech & Reach

All First Amendment protected speech is welcome. No algorithmic throttling or shadow banning.

Family-Friendly Platform

We maintain a clean environment. Explicit adult content is strictly prohibited.

Western Nations Only

Third-world IPs are blocked. No scammers, no spam farms. Built for Western civilization.

Funded By Users

Our users are our investors and customers. You're not the product being sold.

Battle Tested

A decade of standing strong. Banned from app stores, banks—and still here.

American Owned & Operated

We reject foreign censorship demands. Built by Americans, for free people.

Support Alto & Gab

Alto is funded entirely by readers like you. Your donation helps us continue delivering curated news from a right-wing Christian Nationalist perspective, powered by Gab AI.