White House Keeps AI Safety Framework Secret Amid Growing Industry Alarm
The White House is withholding its AI testing framework as experts warn of existential risks. Discover the latest on federal oversight and AI security.
POLICY WIRE — Washington, USA — The White House has yet to disclose the voluntary framework for testing frontier artificial intelligence models that was finalized in August. There is currently no timeline for when these guidelines might be made public, even as current and former staff at Anthropic raise alarms regarding the potential existential threats posed by advanced AI.
Anthropic, the firm behind the Claude model, reported on Thursday that it successfully blocked several attempts by researchers to utilize its technology for activities that could facilitate the creation of biological weapons.
The administration completed its framework in early August, following an executive order issued by President Trump in June. That order requested that major AI developers, including OpenAI and Anthropic, provide the federal government with access to their most sophisticated models up to 30 days prior to public release to bolster security. President Trump had previously delayed an initial version of the order, citing concerns that federal oversight might stifle technological innovation.
Because the evaluation criteria remain confidential, the public is unaware of the specific standards the government expects companies to meet or whether firms are reporting significant technical breakthroughs. Neither the administration nor the participating companies are under any obligation to publish review results or even confirm their involvement in the process.
OpenAI CEO Sam Altman noted last week that his company submitted its new Astra model for federal review, describing the interaction as productive. Altman told Axios that as models reach higher levels of capability, close engagement with safety institutes in the U.S., the U.K., and other nations will become increasingly vital.
📄 POLICY WIRE WHITEPAPER PUBLISHED: PAKISTAN’S NATIONAL SECURITY POLICY PRIORITIES
Internal concerns regarding the technology’s trajectory have intensified. This week, Anthropic researcher and former OpenAI staffer Jacob Coxon resigned, stating on X that the individuals developing AI earnestly believe that it could kill us all by the end of the decade. Fellow Anthropic employee Evan Hubinger echoed this sentiment, stating, “We really do earnestly believe A.I. could kill all humans!”
Advocacy groups are pushing for greater openness. The Center for Democracy and Technology recently sent a letter to the White House demanding the release of the frontier model review framework. Tim Harper, an expert on election security at the organization, argued that the public deserves to understand the government’s legal interpretation of these issues, noting that the lack of transparency reflects a broader trend within the administration.
In July, a coalition of over 1,000 AI industry employees signed a statement calling on the U.S. government to support international efforts to develop governance tools that can effectively pace the advancement of automated AI. Despite these warnings, the White House continues to prioritize the development of advanced models, driven by the strategic goal of maintaining global dominance and outpacing China.
President Trump has consistently championed AI development and the expansion of domestic data centers, recently remarking that China could not be happier about opposition to new infrastructure. The White House did not provide a response to multiple requests for comment regarding the status of the framework.
Reporting by Policy-Wire (PW)





