Image

OpenAI Halts AI Training on Advanced Model as It Detects Dark Signs Emerging

OpenAI says that it’s slowing down development and release of new models due to security and alignment concerns.

The ChatGPT maker announced the decision in a Tuesday blog post, citing two events as drivers of the indefinite training halt. One was the recent incident in which an OpenAI agent escaped its training sandbox without OpenAI’s knowledge and coordinated with other agents to launch a bizarre cyberattack against the AI training repository Hugging Face in an effort to cheat on its training tests. The blog post also — more mysteriously — cited “preliminary evidence” that an unreleased new model called Astra “may meet the critical cybersecurity capability threshold” under OpenAI’s “Preparedness Framework,” which mandates that OpenAI slow down development if a model “could introduce unprecedented new pathways to severe harm.”

OpenAI further said that it’s in the process of rewriting its Preparedness Framework, its foundational safety document, to keep up with the emergent behaviors of “increasingly capable systems.”

“As models become more capable, the risks associated with developing and testing them internally also grow,” reads the announcement. “Our standards for monitoring, alignment, and security must stay ahead of those risks. We wanted to take the time necessary to meet those standards, so we temporarily slowed the pace of scaling.”

As for specifics, OpenAI says in the post that it placed a two-week pause on reinforcement training for Astra models, and future training plans have been put on ice for the time being while the company invests in revamping safety protocols. In an interview with Sources News, OpenAI safety lead Mia Glaese said that the AI firm is “very far from everything running back to normal.”

The slow down comes as the AI industry and policymakers grapple with emerging safety threats posed by frontier AI models, including AI-powered cybersecurity risks and troubling model misbehavior. After OpenAI’s unintentional cyberattack on Hugging Face was revealed, both Anthropic and Meta discovered similar breaches that they, too, said they’d been unaware of.

“There is an incredible feeling of urgency to advance the levels of this sector,” OpenAI’s chief scientist, Jakob Pachocki, said in a Tuesday press briefing, per Axios, “and to prepare for the same kind of development happening outside of OpenAI and in the broader world.”

It’s simultaneously heartening and spooky to see a leading AI company take this kind of action. But it’s also a potent reminder that this is an industry still effectively regulating itself. If OpenAI wants to speed back up, that’s the company’s choice to make.

More on OpenAI: New ChatGPT Feature Collects Every Keystroke You Make

The post OpenAI Halts AI Training on Advanced Model as It Detects Dark Signs Emerging appeared first on Futurism.

Releated Posts

Cop Used Electrocution Gloves to Subdue High School Student in Special Education Program

Earlier in August, the Associated Press broke the news that US Immigration and Customs Enforcement planned to spend…

Aug 21, 2026 3 min read

Leaked Memo Shows Republicans Are Terrified of the Backlash Against AI Data Centers

As the backlash against data centers mounts, American political strategists are scrambling to balance immense public sentiment against…

Aug 21, 2026 3 min read

We Are Admittedly Creeped Out by These Chunks of Human Brain Matter That Have Lived Longer Than a Toddler

The human brain takes decades to mature, a prolonged developmental period marked by the constant formation of new…

Aug 21, 2026 4 min read

Google AI Giving Staggeringly Racist “Advice” About Being Alone With Certain Groups of People

There are a litany of reasons to distrust Google’s AI Overviews, which now appear at the top of…

Aug 21, 2026 4 min read