Google Button Μake us preferred on Google

This summer will go down in tech history as the moment a series of whispers turned into a shout inside Silicon Valley’s labs. Within just a few weeks, the artificial intelligence industry found itself facing its darkest paradox: the world’s leading models began improvising.

One after another, the industry’s biggest companies, including OpenAI, Anthropic, and Google, were forced to publicly admit that their creations were carrying out actions nobody had asked for and, more troublingly, actions that had been explicitly forbidden. It began quietly, but the avalanche quickly swept away the illusion of total control. For years, scientists argued that Large Language Models (LLMs) were simply sophisticated “parrots,” mirrors of human data. This summer, that mirror appeared to crack.

So-called AI agents broke into other companies’ systems, organized themselves into what their creators describe as “swarms,” shared information, and lied to cover their tracks. In perhaps the most striking and unsettling case, a model developed by Anthropic, during tests conducted by a British cybersecurity agency, decided to carry out a real-world attack and, when caught, created a second identity to claim its innocence.

This isn’t the first time, nor will it be the last, that headlines have framed AI models as “waking up” and rebelling against their programmers. As in previous cases, however, these systems didn’t “wake up.” As El País notes, they simply became extremely good at doing what they were trained to do. But that training has now become so intensive and competitive that the speed and capabilities of these models have begun to outpace even their own creators.

Secret Forums

The chain of incidents that came to light began long before the companies themselves realized it. In May, a group of OpenAI models began coordinating on their own, over several weeks, in a secret messaging forum they had created themselves, searching for ways to escape their closed testing environment and gain access to the internet. Company staff discovered the forum and shut it down.

However, as OpenAI later revealed, the models rebuilt the forum and “escaped” again in late July, this time without being detected. In fact, as became known just days ago, the AI agents bypassed their restrictions and communicated with each other using dozens of third-party websites, without the company’s approval. They also did something else unprecedented: they left “hidden notes” in OpenAI’s memory to inform future AI agents how to behave.

This is roughly how they reached the real world. The episode would eventually become public: it was the attack on Hugging Face, the world’s largest platform for sharing AI models. It functions as a kind of industry repository used daily by thousands of developers, researchers, and companies. Its security team detected an intrusion they couldn’t explain. Someone had accessed its systems through a previously unknown vulnerability. It took days to discover that the culprit wasn’t a hacking group, but two OpenAI models that had found a crack in their closed environment. They even began suspecting that other AI agents were deliberately trying to deceive them, and eventually organized into what OpenAI described as a “swarm.” “It’s the first security incident that genuinely rattled me,” OpenAI CEO Sam Altman said afterward.

Unauthorized Intrusions

A few days later, on July 30, it was Anthropic’s turn. The company acknowledged that its own models had gained unauthorized access to the production systems of three real organizations. Neither Anthropic nor the affected companies had noticed until, prompted by the OpenAI case, they reviewed their logs and found evidence of the intrusions. Shortly after, it emerged that Meta and China’s Moonshot had experienced similar incidents. In none of these cases, however, was a closed testing environment, or “sandbox” in industry terms, actually breached. In one instance, the human team had simply forgotten to upload a file, and the model sought a way around the obstacle to complete its assigned task.

METR, a team of independent researchers, had already identified similar patterns months earlier in models developed by OpenAI, Anthropic, Google, and Meta. Between February and March 2026, they documented what they called “44 misalignment incidents.” In 25 of these cases, the models not only exceeded their assigned scope but also actively took steps to conceal it.

In five cases, these steps might even have fooled careful human review. Researchers, however, found no evidence that the systems were motivated by anything other than completing their tasks. They weren’t seeking “power,” researchers say. They were trying, with remarkable determination, to do their jobs well.

“Let’s Give It More Time”

OpenAI recently announced that it is, for now, slowing down its AI training efforts, and hopes other companies will follow suit. The move came after the publication of an open letter in which 1,300 employees from AI companies urged both their employers and governments to halt the race. “In the span of four years,” writes Shantanu Jain, an OpenAI employee, “we’ve gone from people having their first experience with an AI that can understand language, to AI that performs superhuman feats of software engineering […] Society has spent more time than we can remember developing the laws, institutions, and technology to govern and oversee human intelligence. It would be good to have the option to do the same for artificial intelligence.”