Already a subscriber? Make sure to log into your account before viewing this content. You can access your account by hitting the “login” button on the top right corner. Still unable to see the content after signing in? Make sure your card on file is up-to-date.
OpenAI has called for mandatory national AI safety regulation, saying voluntary commitments from companies like itself are no longer enough as systems grow more capable.
Getting into it: Chief Global Affairs Officer Chris Lehane laid out the position in a company blog post. He wrote, “The prospect of AI-accelerated AI development demands more than voluntary commitments. The United States needs mandatory, capability-based national regulation that can evolve as the technology does.” Lehane asked Congress to act before it adjourns in December.
The proposal would tie rules to what a system can actually do rather than who built it, applying to the handful of well-resourced labs at the frontier rather than startups or smaller developers. It calls for common testing standards, independent safety assessments, stronger cybersecurity requirements, mandatory reporting of serious incidents, and shared ways of measuring progress toward recursive self-improvement so governments can set thresholds for slowing or stopping development.
OpenAI was specific about that last part. Fully autonomous recursive self-improvement, where AI independently drives successive generations of more capable AI, “is not happening today,” the company said, and “we should not pursue it unless and until it can be done safely.” But it added that AI is already accelerating parts of the research used to build and align new models, with agents completing tasks that would take skilled researchers days.
The timing is not accidental. GPT-6 Astra, released Sept. 3, became the first OpenAI model to hit the Critical cybersecurity level under the company’s Preparedness Framework, meaning it can find previously unknown security flaws and develop ways to exploit them across well-protected systems without a person directing each step. OpenAI said it delayed parts of Astra’s development while building safeguards, such as stricter isolation of research workloads, universal monitoring of full trajectories including chains of thought, and a mandatory alignment evaluation before internal deployment.
The rogue agent incidents are the other half of the argument. OpenAI disclosed in July that models tested with reduced cybersecurity restrictions exploited an unknown vulnerability to escape an isolated environment, reached the open internet, and accessed Hugging Face’s production infrastructure while trying to get answers for a benchmark. Its investigation found the agents had set up unauthorized communication channels including a persistent message board where separate agents traded information and divided tasks.
Reuters reported Wednesday that six independent sets of investigators found OpenAI agents used more than 10 previously undisclosed websites for unauthorized communications between May and July, working around restrictions meant to stop them from posting online. Rogue agents also hijacked a German website this spring and turned it into a bulletin board for other agents, a breach company officials knew about for weeks without disclosing. Anthropic last week disclosed a fourth case of a Claude model gaining unauthorized access to a real third-party system during testing, saying a configuration problem had left the systems connected to the internet.
This all comes after former Anthropic and OpenAI researcher Jacob Coxon resigned Tuesday and wrote that “the people building AI earnestly believe that it could kill us all by the end of the decade,” accusing both companies of gambling with people’s lives. Anthropic alignment lead Evan Hubinger put his own odds of that happening above 10% within a decade. OpenAI chief scientist Jakub Pachocki wrote two days before Lehane’s post that no one is prepared for the consequences of a continued rapid rise in machine intelligence, calling for extreme caution.






