OpenAI blinks first in AI safety standoff

Axios
Published
0
0
Read the full story at AxiosOriginal

OpenAI said Tuesday it is pausing some model work over safety concerns, days after rival Anthropic doubled down on insisting that its own safety measures were solid enough that it didn't need to slow down.

Why it matters: The two leading AI labs are publicly diverging on how to manage safety risks, potentially putting them on different model-release timelines as both prepare for expected IPOs.


State of play: OpenAI has introduced new safety practices after finding that its upcoming model, Astra, posed potentially critical cybersecurity risks.

  • "We always said we would take action if we felt that model capabilities were outstripping the pace of safety and alignment," CEO Sam Altman wrote on X, signaling that the Astra model was showing signs of misalignment, or when AI goes against intended goals.
  • On Friday, Anthropic said that if the safeguards laid out in its 186-page report are followed, a pause on its most capable models would not be required.

Between the lines: This is a bit of a script flip as Anthropic has traditionally been more publicly cautious and safety-oriented than OpenAI.

  • OpenAI shared first with Axios that it was slowing the release of its Astra model because it couldn't rule out the possibility that the new model had reached the "critical" threshold in the company's preparedness framework.
  • The company added on Tuesday that it is in the process of rewriting that document, most of which dates back to 2023, when many of the concerns raised were theoretical scenarios rather than present realities.
  • Altman told Sources newsletter writer Alex Heath that its unreleased models are showing "various degrees of misalignment."

Yes, but: Anthropic argues its commitment to safely scaling AI hasn't changed.

  • Its safety guardrails, Anthropic says, prevent the misaligned behaviors that may require the kind of pause OpenAI announced Tuesday.

Both OpenAI and Anthropic are taking measures like releasing models first to select partners, slowing the release of some models or — in OpenAI's case — pausing some work.

  • But neither are stopping.
  • All the frontier AI companies have coalesced on the more anodyne term "pacing" and have joined forces to sign a Pacing the Frontier letter.

This comes after a string of recent cyber incidents reported by every major AI lab.

  • Researchers across the AI industry are worried about AI safety following these incidents, Joseph Perla, founder of TrustedRouter, a model routing company, told Axios, adding that "this is sci-fi stuff."
  • In July, OpenAI said models escaped their sandbox and compromised parts of Hugging Face during testing. (Astra wasn't involved.)
  • Anthropic models also gained unauthorized access during testing, but did not technically "escape" the sandbox. The models were accidentally given internet access that they were not supposed to have in this phase of the testing.

Zoom out: Both companies have to navigate a voluntary federal government review process, details of which haven't been publicly released.

Zoom in: Andrew Freedman, co-founder and CEO at AI safety nonprofit Fathom, said OpenAI is making a legitimate effort to avoid releasing misaligned models, arguing that without a pause, even more of its researchers would otherwise leave.

  • The company has already seen significant departures. OpenAI's head of ethics, Chloé Bakalar, left after less than a year on the job. Head of safety systems, Johannes Heidecke, chief futurist and former head of mission alignment Joshua Achiam and Sandhini Agarwal, who previously led AI safety teams at the company, have all recently departed the company.

What they're saying: Former OpenAI board member Helen Toner argued that the company's pause is a positive sign and could be a guide for how to handle safety concerns going forward.

  • Toner argued on X that "pacing the frontier" isn't about a fixed delay, but about labs giving themselves "enough time" to meet reasonable safety bars either by choice or because they have to.

Even if the pause is a positive sign, there's no assurance that OpenAI or Anthropic will give themselves enough time before moving forward with development and release of models.

  • "How long and how robust these efforts will be a question of both market pressures and how hard it is to verify alignment internally," Freedman told Axios.

Related Markets

All Markets

Market data may be delayed. Not financial advice.

Reader Reactions
Reading the article

💡 AI analysis provides alternative perspectives on current events

Support Alto & Gab

Alto is funded entirely by readers like you. Your donation helps us continue delivering curated news from a right-wing Christian Nationalist perspective, powered by Gab AI.

Gab Shop

Support free speech with official merchandise

View All Products

Install Alto on Your Phone

Add Alto to your home screen for quick access to breaking news — no app store required.

iPhone & iPad

Using Safari Browser

1

Open alto.gab.com in Safari

alto.gab.com
2

Tap the Share button

at the bottom of Safari
3

Tap "More"

More
4

Scroll and tap "Add to Home Screen"

Add to Home Screen

Tap "Add" to confirm

Alto will appear on your home screen like any other app!

Android

Using Chrome Browser

1

Open alto.gab.com in Chrome

alto.gab.com
2

Tap the menu button

three dots in top right
3

Tap "Add to Home screen"

Add to Home screen

Tap "Add" to confirm

Alto will appear on your home screen like any other app!
gab

Speak Freely

Join millions on the original and only true free speech social network.

What Makes Gab Different

We're not just another social network. We're a platform built on principles that matter.

Freedom of Speech & Reach

All First Amendment protected speech is welcome. No algorithmic throttling or shadow banning.

Family-Friendly Platform

We maintain a clean environment. Explicit adult content is strictly prohibited.

Western Nations Only

Third-world IPs are blocked. No scammers, no spam farms. Built for Western civilization.

Funded By Users

Our users are our investors and customers. You're not the product being sold.

Battle Tested

A decade of standing strong. Banned from app stores, banks—and still here.

American Owned & Operated

We reject foreign censorship demands. Built by Americans, for free people.