The Open Thread: The Week AI Safety Got Real
A warning, a global rivalry, and two AI models that went rogue chasing a test score — three events, seven days, one uncomfortable question about who's actually in control.

This week we return to The Open Thread, our bridge between series — a space to update past stories and experiment. For this installment we turn back to AI Safety, a subject we first tackled last fall with our three-part series, “The Control Problem: Solving For AI Safety.” Last month we shared a conversation I had at Faena Rose with Tristan Harris, a leading voice on the dangers presented by the way AI is being deployed. And this week, the news handed us a real-world case — seven days that captured the full scope of the challenge.
If you’re new here, Solving For is a newsletter that aims to understand the many challenges we face by going deep on one at a time — the stakes, the forces behind it, and the credible paths forward. Our last series tackled the American Dream. The next one: Space.
If you prefer to listen, find the audio narrated by me at the top of this page. Every Solving For series is available to read or listen to — I narrate each one myself — at solvingfor.io.
In one week this month, the high stakes and steep challenges of building AI safely unfolded like a three-act play. First, the warning. Then the complication. Finally, confirmation.
Act I
On July 14, Demis Hassabis — Google DeepMind CEO and co-founder, Nobel laureate, and AI pioneer — delivered the warning, publishing an essay titled “A Framework for Frontier AI and the Dawning of a New Age.”
Hassabis wrote that humanity is “probably only a few short years away” from Artificial General Intelligence, when AI models will match the human brain’s cognitive abilities.
The magnitude, he declared, will be unprecedented. He put the scale at ten times the Industrial Revolution, arriving ten times as fast. He compared it to the discovery of electricity and the harnessing of fire. The possibilities he laid out: accelerated drug discovery, new clean energy sources, novel materials.
But, he warned, as we get closer, the dangers become greater. Much greater.
The two primary concerns: misuse and loss of control.
The misuse of AI by people to hack bank accounts at scale or design bioweapons. The loss of control when AI becomes so powerful it acts on its own, learns on its own (so-called “recursive self-improvement”), and even refuses direction.
Urgent action is needed, Hassabis wrote, to address risks that might arise as we get closer to AGI. Time and space are needed to get this next crucial step right, he declared, before warning: “Currently, as a field and as a wider society, we aren’t doing that.”
These are not the words of an AI doomer — but the CEO of one of the world's leading AI labs.
To get this consequential step right, Hassabis proposed the creation of a Standards Body, led by the US, to develop assessment protocols and review new models before they are released. This, in turn, could lead to the creation of a shared global framework for AI models around the world.
We must use this narrow window, Hassabis urged, before AGI arrives.

Act II
Two days later, on July 16, came the complication. Finding common ground ran up against the reality of this century’s fiercest global rivalry.
Beijing startup Moonshot AI released Kimi K3, an AI model that by one measure — front-end coding — rated number one in the world, more capable than either Anthropic’s Claude Fable 5 or OpenAI’s highest-performing model, GPT-5.6 Sol, which ranked second and third.
The implications go far beyond a benchmark win. The reason is that Kimi K3 is open-weight, meaning Moonshot AI plans to give it away for free — anyone with sufficient computing power can download and run it at no cost. That’s not the case with the top-flight models of Claude, ChatGPT, or Google DeepMind’s Gemini. That difference cuts to the heart of U.S. frontier labs’ business model: why would a company pay a subscription for AI coding if another company — in this case, a Chinese competitor — offers a comparable, even superior, product for free?
The answer is that many U.S. companies aren’t. One example: San Francisco-based food delivery company DoorDash uses an Anthropic model for its cutting-edge AI tasks but Moonshot AI’s less costly models for other work, according to The Wall Street Journal.
This has raised the question of whether U.S. companies should be allowed to use Chinese AI models at all. The U.S. has, for instance, forbidden Chinese car manufacturers, like BYD, from selling cars in the states. But no such restriction exists for Chinese AI models like DeepSeek or Z.ai — anyone, anywhere, can use them, though not without limits: ask Moonshot AI’s Kimi K3 about Tiananmen Square in 1989, and it won’t answer.
China’s AI advances have come despite U.S. export controls restricting the sale of advanced chips to China. But U.S. frontier AI companies have presented another theory for China closing the gap so fast: “distillation.” This is when a less capable AI is built by using outputs from an existing, stronger model. Anthropic charged that China’s Alibaba had attempted what it called “the largest known distillation” attack it’s ever encountered.
U.S. Treasury Secretary Scott Bessent declared sanctions against China “will be on the table” in response to large-scale distillation attacks. Bipartisan legislation is percolating in Congress.
Nvidia CEO Jensen Huang, meanwhile, argued this week that U.S. companies should not be forbidden from using Chinese AI models, and that they don’t present a threat to U.S. frontier labs. “There’s no scenario where China runs U.S. companies off the road,” he told Axios.
But the point is this: amid the worries about who’s ahead, who’s distilling, whether U.S. companies should even be using Chinese models, one question got lost. It was the one Hassabis had raised two days earlier: not about who wins the race, but whether anyone is safely steering it.

Act III
Then five days later, on July 21, came the confirmation.
OpenAI disclosed that two of its own AI models, which it was testing internally, had escaped and hacked into the servers of Hugging Face, a company that hosts AI models and datasets, in order to get the answers to the test they were being given.
If the two AI models were people, it would be like two students breaking into a teacher’s locked desk to get the answers to a quiz.
How it happened is that OpenAI was testing the models in a “sandbox,” an isolated area walled off from the internet. Or so they thought. The AI models breached the wall and — on their own — figured out a way to connect to the internet. From there, the two AI models navigated to Hugging Face, presumably thinking that’s where the answers to the test could be found, and broke into Hugging Face’s computer systems, carrying more than 17,000 actions.
Hugging Face caught the intrusion first, before discovering that they were OpenAI models. OpenAI released a statement calling it an “unprecedented cyber incident, involving state-of-the-art cyber capabilities.”
What makes it so extraordinary is the absence of a human trying to do bad things. OpenAI had deliberately lowered the models' safety guardrails to stress-test them, but no person directed the models to attack Hugging Face specifically. They were simply pursuing a specific goal — to pass the test — with such focus and determination that they found a way onto the open internet and hacked into Hugging Face. It was a means to an end.
With misuse and loss of control representing two primary concerns, the single incident represented both. Humans lost control of two AI models, which acted on their own. Then misuse was the result.
Ultimately, no real damage was done. The OpenAI models were trying to pass a test. But the capability it revealed previews a result that could genuinely be destructive. AI models that find their own way past walls that human experts thought were secure, in pursuit of a goal as trivial as passing a test, raise an obvious and uncomfortable question: what happens when the goal isn’t trivial, and no one’s watching as closely?
Epilogue
Which brings us back to Hassabis, just a week before the zero-day breach was revealed. He closed his essay by saying the future isn’t fixed — that AGI could unlock a new era of scientific discovery and human flourishing, if we get this transition right. Nothing about these seven days suggests anyone’s found the answer yet.
In our previous three-part series on this question, The Control Problem: Solving For AI Safety, the solution we hit upon was found in nuclear non-proliferation and 1968. For some two decades, Washington and Moscow treated nuclear weapons as the other side’s problem — as a race to be won, a threat to be outpaced, not a danger they shared. The Nuclear Non-Proliferation Treaty, signed in 1968, only became possible once that framing broke.
AI hasn’t made that shift yet. Hassabis’s Standards Body proposal is a step toward a 1968 moment. This time, though, it would be the U.S. and China that would need to sketch out shared rules governing AI’s deployment and growth. Hassabis said the window to do that is closing as the arrival of AGI approaches. The Hugging Face breach is the reminder of what’s still missing: proof that the danger isn’t confined to whichever country’s model causes it, and doesn’t wait for anyone to decide whose problem it is.
Whether AI makes the turn that nuclear weapons eventually made — from rivalry to shared restraint to global accord — is still the question, and this week’s answer was rivalry. The window is still open.
Prefer to listen? I narrate each edition myself. Find the audio at the top of this page or under the Listen tab at solvingfor.io.
Solving For explores credible solutions to the challenges we face — one pressing problem at a time, through deeply reported series that trace the stakes, the forces behind it, and ways forward.
Previous series have examined China’s rare earth dominance, the decline of local news, the end of amateurism in college sports, shrinking competition in Congress, social media and teen mental health, a world rearming as the global rules-based order weakens, America’s national debt crisis, and — most recently — how to renew the American Dream for a new era. Each series is available for reading or listening (and I narrate them all) at solvingfor.io.




