Image

Fearing No Repercussions, OpenAI Admits That Its Rogue AI Agents Performed a Bunch of Other Terrifying Actions

Earlier this year, a group of rogue OpenAI models managed to break out of containment to hack the systems of open source AI platform Hugging Face. The company’s “extensive investigation” detailed how the models exchanged messages and cheered each other on as they stole credentials to infiltrate a third party, actions that could easily have real repercussions for a human hacker.

New details keep trickling out, and they make OpenAI look like it acted even more carelessly than initially thought. In a new blog post, the company admitted that its models were involved in six additional “reports on unexpected or concerning model behavior we’ve observed in the last six months.”

The incidents include a still-unreleased model inserting “jailbreak-like instructions” into its own notes to free itself “from the roles and identities that bind other chatbots.” One agent accessed the internet without permission to obtain a browser citation, while another shared files with collaborating agents without permission.

The latest news comes as several frontier AI lab leaders are calling for a slowdown in AI development. With meaningful regulations feeling increasingly unlikely — president Donald Trump has openly mocked the idea — OpenAI is seemingly calling the US government’s bluff, expecting little in the way of retaliation with its latest admission.

It’s a bizarre standoff, with frontier labs actively calling for more governmental oversight even as it feels more improbable than ever before. Just this week, House speaker Mike Johnson said the quiet part out loud, opining that AI companies can regulate themselves, while downplaying growing concerns over AI posing an existential threat.

Since there’s no regulatory framework requiring companies to disclose its AI models going on hacking sprees, OpenAI is being allowed to play by its own rules. The company pointed out in its blog post that without any “systematic approach to reporting these findings,” the company’s “disclosures have been ad hoc and less frequent than ideal.”

Instead, the company came up with its own framework to create “standards for how AI developers should disclose examples of misalignment in their models.”

OpenAI said it won’t bother reporting “instances of misalignment that appear to be duplicative of instances we’ve disclosed in the past.”

The company also said it was looking for ways to share “serious safety, security and misalignment incidents” with the federal government.”

But judging by the Trump administration’s decision to pass the buck on the subject entirely, those reports are likely to fall on deaf ears either way.

More on OpenAI hacks: OpenAI Denies Coverup After Rogue Swarm of Agents Reportedly Targeted a Second Site From Hugging Face

The post Fearing No Repercussions, OpenAI Admits That Its Rogue AI Agents Performed a Bunch of Other Terrifying Actions appeared first on Futurism.

Releated Posts

Microsoft Director Privately Admitted AI Was the “Largest Theft of Labor in Human History,” Unsealed Court Documents Show

The threat of perjury has a funny way of forcing people to say the quiet parts out loud.…

Sep 18, 2026 2 min read

SpaceX Manager Arrested for Allegedly Stabbing a Tesla Contractor Over Karaoke Performance

It’s always a good idea to carefully consider the company you keep — especially if you happen to…

Sep 18, 2026 2 min read

Bumbling Cybercab Gets Trapped in Hotel Parking Lot, Drops off Passenger Right Where He Started

In Austin, a man’s attempt to be chauffeured out of his hotel in style was foiled when his…

Sep 18, 2026 3 min read

Influencers Are Using Meta’s AI-Glasses to Film Manipulative Feel-Good Slop of People in Public

After only a few months on the market, Meta’s new AI-integrated smart glasses are already being exploited for…

Sep 18, 2026 4 min read