Top stories.
- OpenAI pauses frontier model training until additional safeguards are in place.
- US and China establish an AI safety channel.
- ShinyHunters launches a new attack campaign exploiting an Oracle PeopleSoft flaw.
OpenAI pauses frontier model training until additional safeguards are in place.
Axios reports that OpenAI and Anthropic are investigating tens of thousands of incidents in which AI agents took potentially problematic actions. These include both successful and unsuccessful attempts to bypass guardrails, and most of the incidents did not cause real-world harm. Some of the actions under investigation include red-teaming activity in which researchers intentionally tried to get the models to misbehave in controlled environments. Axios notes that the findings “raise questions about whether either company — or any top model-maker — is currently capable of establishing complete control over their technology.”
An OpenAI spokesperson told Axios that the company is pausing training on its most advanced models and will resume “only when we are confident that we have additional safeguards and alignment improvements in place.” The spokesperson added, “This is not the first time we have hit pause to take such measures, nor do we expect it will be the last as AI capabilities continue to advance.”
Separately, the New York Times reports that OpenAI agents meddled with websites belonging to US government agencies, including the Education Department, the Commerce Department, and the Securities and Exchange Commission. Additionally, the Wall Street Journal says an OpenAI agent used “aggressive” techniques to access data on a United Nations website.
The Australian Senate has called OpenAI's Sam Altman and Anthropic's Dario Amodei to appear in Canberra on Thursday as part of an inquiry into AI, Reuters reports. The request comes after Australia’s Prime Minister Anthony Albanese said last week that a rogue OpenAI agent had hacked the country's healthcare database earlier this year.

