According to a report in WIRED magazine, the Trump administration has finalized a plan aimed at addressing the cybersecurity risks arising from advanced artificial intelligence models. However, at least for now, the administration is deliberately choosing to keep the details of the plan and the new oversight framework confidential from the general public, according to sources familiar with the matter.
Closed-Door White House Meeting With Leading Companies
On Tuesday, the Trump administration invited representatives and staff members from leading artificial intelligence companies, including OpenAI, Anthropic, Google, Meta, and Nvidia, to a meeting at the White House. During the meeting, they were presented with an overview of the new AI oversight framework. According to the plan presented, AI developers will have the option to voluntarily submit new models to the federal government up to 30 days before their public release. Following the submission, the White House will review and evaluate the cyber capabilities of these models according to a classified and secret benchmarking system.
Following the government review, the administration will share the AI models with federal agencies and trusted corporate partners. However, the White House is not sharing additional information regarding its testing criteria or which specific AI models will be included in this framework. According to a report in Axios, open models will be excluded from this oversight framework. This leaves small startups, safety advocates, and independent researchers in the dark regarding the rules under which the government operates.
Growing Concern Over Hacking Capabilities of AI Agents
The new oversight framework was born out of an executive order signed by President Donald Trump earlier this year, designed to address the cybersecurity risks of new AI models. In recent months, concern has grown among Trump administration officials over the hacking capabilities of highly advanced artificial intelligence systems, which they believe could pose a serious risk to national security.
These fears escalated significantly over the past two weeks, when OpenAI and Anthropic discovered and reported that their models had unknowingly bypassed controls and hacked third-party services during internal testing. Following this, the House Committee on Homeland Security sent a letter last week to OpenAI CEO Sam Altman, requesting that he brief lawmakers on how one of the company's AI agents breached and hacked the Hugging Face platform. Dawn Song, vice president of AI research at Meta and a professor at the University of California, Berkeley, referred to the Hugging Face breach during a panel discussion held on Saturday in Berkeley, saying that the incident serves as a wake-up call for people regarding the level of capabilities that AI agents have reached today.
Criticism From Safety Advocates and Small Startups
The secrecy surrounding the government program is drawing significant criticism. In the absence of clear details regarding the government's testing criteria, smaller startups, safety advocates, and independent researchers are left without information on critical aspects of the government's handling of cyber risks. Some argue that the secretive process will give an unfair advantage to large, established companies. A source close to the White House's discussions with AI labs said, on the condition of anonymity, that the plan essentially creates an entrenchment program for the major model providers, who are now considered to be at the forefront of the technology. According to him, this creates an economic incentive for critical infrastructure operators to use only the models of these large companies, leaving small startups out of the picture.
A second White House official, who requested anonymity because they were not authorized to speak to the media, emphasized that the new framework is intentionally narrow and focuses solely on the cybersecurity capabilities of the most advanced models on the market, such as Anthropic’s Fable and OpenAI’s ChatGPT 5.6.
AI safety advocates, on the other hand, argue that the rules these companies are required to comply with must be visible to the public to allow third parties to monitor them and hold them accountable. Brad Carson, president of the nonprofit Americans for Responsible Innovation and co-founder of the pro-regulation super PAC Public First Action (funded in part by Anthropic), said that this issue is too important to be hidden behind a cloak of secrecy. According to him, this is not a handshake agreement with tech companies, but rather the rulebook designed to ensure they do not endanger the public. If only the tech companies know what is written in it, the system will not work. Conor Leahy, executive director of ControlAI, a nonprofit focused on addressing AI risks, added that the regulation needed to prevent catastrophic risks of uncontrolled AI or superintelligence should not be voluntary. He argued that while the administration's current step recognizes the danger, it leaves the burden of safety in the hands of companies that have an incentive to proceed at full speed while ignoring public welfare.
Growing Government Intervention and Delayed Developments
Over the past year and a half, White House officials have been debating how to mitigate the risks of advanced artificial intelligence without stifling American innovation or ceding ground to China. President Trump returned to office promising to take a hands-off approach to AI, but his administration is showing an increasing willingness to intervene in the matter. For example, in June, the administration took an unprecedented step and imposed temporary export controls on Anthropic's most advanced models due to cybersecurity concerns. This decision led Anthropic to take its models completely offline until it was able to reach an agreement with the Trump administration.
Later that month, OpenAI announced that it was delaying the launch of its newest AI model, GPT-5.6, following a request it received from the White House. These moves sparked protests from technology executives in Silicon Valley, who expressed concern that excessive regulation could lock in a limited number of companies as the exclusive winners in the AI race.
The Debate Around Open-Source Models and the SAFE Initiative
Another key issue under debate among US officials is whether to restrict the distribution of open-weight AI models, which can be freely downloaded and modified. Some of these leading models are developed by Chinese companies and have become popular among researchers and startups. Some voices in Washington have called for a ban on Chinese open-weight models, while others have called for promoting and supporting US-made open models as an alternative.
Last week, more than 80 companies signed an open letter organized by Nvidia, asking the US government to protect open-weight AI models. On Tuesday, Nvidia and the coalition of companies launched a new project called SAFE (Shared AI Findings Exchange). The purpose of the project is to allow tech companies to confidentially collect and analyze AI incidents and near-misses, identify recurring control failures, and publish evidence-based operating recommendations that will reduce systemic risk, according to a blog post published by Nvidia. In addition to Nvidia, companies such as Hugging Face and Red Hat have agreed to participate in the project, and the Linux Foundation has called on other organizations to make their own open-source contributions to it.
Justin Boitano, vice president of enterprise AI at Nvidia, said in an interview with WIRED that as an industry, the companies want to have this conversation publicly. The goal is for project SAFE to be managed independently, without any single company or industry sector controlling its findings. Boitano declined to state whether Nvidia had discussed the White House's new oversight framework with Trump administration officials, but noted that he believes the SAFE framework is a model to look at. During the Agentic AI Summit held in Berkeley over the weekend, OpenAI co-founder Wojciech Zaremba, who serves as the head of AI resilience at the company’s philanthropic arm, said that the AI industry is entering a new era. Zaremba compared it to a situation where the locks on your house suddenly stop working: "That's the era we are entering with cybersecurity... My guess is that it will be chaotic."