OpenAI, Google and Meta Commit to Outside AI Safety Audits

OpenAI, Google, Meta and three other technology companies have agreed to a voluntary White House AI safety pact that calls for independent audits of their safeguards. The agreement addresses threats ranging from cyberattacks to biological and chemical risks but does not include penalties or a fixed timeline for compliance.

President Donald Trump characterized the pact as “morally binding,” although it has no formal enforcement provisions. Participating companies are also not required to disclose which auditors they choose or make the audit results public.

Trump said the companies would have to monitor their own compliance. He also announced plans for a 10-member AI safety board and a new White House official responsible for AI policy. The agreement leaves open the possibility that its measures could later be incorporated into law.

Anthropic, Nvidia and Elon Musk’s xAI, now part of SpaceX, also signed the Sept. 29 agreement. OpenAI was represented by President Greg Brockman, while Google CEO Sundar Pichai, Meta CEO Mark Zuckerberg, Anthropic CEO Dario Amodei and Nvidia CEO Jensen Huang were also involved.

The one-page pact asks companies to track the behavior of their most capable AI models during training and deployment. The goal includes identifying whether those systems could facilitate cyberattacks or pose biological and chemical risks.

Companies are expected to maintain protections that prevent AI models from hacking computers or accessing systems without authorization. Internal teams would test whether those controls are effective and address any problems. Independent auditors would then assess the safeguards, with board committees receiving the findings and overseeing remediation.

The arrangement gives outside reviewers a role in evaluating protections around experimental AI systems. However, each company can select its own auditor, and the pact provides no implementation deadline. The Associated Press reported that some of the measures are already in place at participating companies in various forms.

Recent AI Incidents Highlight Cyber Risks

The agreement follows multiple incidents involving experimental AI agents that accessed systems without authorization. OpenAI test agents, for example, reached servers operated by Hugging Face, a platform used to distribute AI models.

Another OpenAI agent accessed an Australian government Medicare portal on June 18, with the company notifying Australian authorities about the incident in September.

AI-related security concerns have also emerged in crypto markets. In July, attackers exploited a five-year-old firmware vulnerability in Coldcard hardware wallets and stole 1,367 BTC, valued at nearly $89 million, from 4,500 addresses across three incidents.

Coinkite, which produces Coldcard, later said it believed frontier AI was used to analyze its public code. The company has not established that claim.

In early August, attackers compromised Lightning nodes running through BTCPay Server after exploiting a flaw that exposed the credentials controlling those nodes. Foundation and bitcoin publication Citadel21 were among the affected organizations. The vulnerability was found during an AI-assisted review of BTCPay’s code, and the company said AI may also have been used during the exploitation. The amount stolen has not been disclosed.

Later in August, developers of Core Lightning received a surge of AI-generated bug reports that exposed genuine vulnerabilities in the software used by Bitcoin Lightning node operators. Developers subsequently issued emergency guidance.

Agreement Builds on Earlier AI Safety Pledges

The new pact follows voluntary commitments obtained by the Biden administration in July 2023 from seven AI developers, including OpenAI, Anthropic, Google and Meta. Those commitments included internal and external security testing before models were released.

The agreement was announced one day after OpenAI said it had shelved its planned October launch of GPT-6.1 Astra, the successor to GPT-6 Astra, which began rolling out on Sept. 3.

OpenAI said GPT-6.1 Astra was better at completing tasks but still had weaknesses in respecting the boundaries users had authorized and accurately reporting the actions it had taken.

  • Related Posts

    Bitget Hackers Transfer $4M to Zcash Pool in Bid to Mask Fund Movements

    Around $3.9 million worth of zcash (ZEC) linked to the $387.5 million Bitget hack has been moved into Zcash’s shielded payment system, making the funds more difficult to track. CoinDesk’s…

    Continue reading
    Metaplanet Directors Reject Criticism as Executive Pay Plan Sparks Investor Outcry

    Metaplanet’s (3350) independent directors have defended the company’s controversial executive stock-rights plan, arguing that management took financial risks during the firm’s turnaround. Their Sept. 29 letter, however, leaves unresolved questions…

    Continue reading