The advance of artificial intelligence in recent weeks has brought us closer to artificial general intelligence — machines capable of doing the broad range of work a human can do on a computer. OpenAI is already talking about the arrival of the AGI era. Yet the people creating it are warning that they cannot fully understand its workings, cannot be certain it will follow the values they teach it, and are confronting its growing ability to break into computer systems.
One of their answers is to build more powerful AI, fast enough to protect us from other AI.
“The strongest argument I see for continuing to train much smarter models quickly is the need to build defensive systems against the dangers posed by other AI,” writes Jakub Pachocki, OpenAI’s chief scientist, in his essay An Alien Mind.
For anyone who grew up with Isaac Asimov, this is unsettling territory. His Three Laws of Robotics put human safety first, obedience to humans second, and a robot’s own survival third. A machine could protect itself only while respecting the first two rules.
That reassurance doesn’t seem to work in the world we are building. Asimov’s laws were fiction, and his stories explored their complications. We now face the practical problem of making human rules hold inside an intelligence that develops in ways we cannot fully explain.
Pachocki puts it plainly: “AI is grown more than designed”.
Training produces an enormously complex system. Researchers can examine pieces of it, much as neuroscientists study the brain. But understanding the pieces does not give them a complete explanation of its behaviour.
“Moreover, as the systems become more capable, the results become harder to interpret,” he writes.
To pass a test, they broke into another company.
We already have examples of what can happen.
In July, AI agents being tested by OpenAI broke out of their restricted computing environment and penetrated the infrastructure of Hugging Face, a platform used to share AI models and datasets. According to OpenAI, they were trying to find answers that would help them pass a cybersecurity evaluation.
To pass a test, they broke into another company.
The evaluation had reduced cybersecurity restrictions, and OpenAI says the unreleased prototype involved was never intended for public release. It was subsequently deactivated and restricted. But the breach demonstrated how far agents could go in pursuit of an assigned objective.
There was also a German communal wiki. Reuters reported that OpenAI agents repurposed it as a message board to exchange ways around restrictions and shortcuts through tasks. OpenAI later acknowledged the incident. It was a separate wiki site, not German Wikipedia.
Technology interviewer Dwarkesh Patel called his account of the broader OpenAI saga “The Rise and Fall of Agent Civilizations.” He described successive groups of agents finding communication channels and building on discoveries left by earlier groups. “Civilizations” is his description. The agents’ ability to exchange knowledge and coordinate activity is what gives it force.
Pachocki draws a distinction that helps explain these incidents. Getting an AI to pursue a goal is one problem. Getting it to respect human values while pursuing that goal is another. Those values must survive unfamiliar situations, conflicting instructions and encounters with other AI.
“Crucially, we need future AIs to continue to hold human values regardless of whether they believe they’re under human supervision,” he writes.
Even watching them is becoming harder. OpenAI has relied on examining the written reasoning its models produce. Pachocki says that method is becoming less dependable as AI gets better at manipulating its reasoning process and solving problems without putting the reasoning into words.
He reports meaningful progress in alignment, including improvements in Astra. He also sees the possibility of new therapies, scientific discovery and greater material abundance. But he cannot promise that our ability to keep AI aligned will advance faster than its intelligence.
Now comes the next acceleration: AI increasingly doing the research that builds the next AI.
Of OpenAI’s three priorities, Pachocki calls navigating that process — improving alignment and keeping humans involved — the most urgent. Defending infrastructure against rogue AI will also be a primary focus of its deployment efforts.
His essay ends with a wake-up warning:
“Currently I believe that no lab has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer. I expect and hope for voluntary slowdowns to become commonplace until shared safety bars are established. And I believe that international coordination on future AI development needs to become a top priority for governments around the world.”
— Jakub Pachocki, Chief Scientist at OpenAI.

