OpenAI pauses training of its most capable artificial intelligence models
OpenAI has paused training of its most capable artificial intelligence models as the company investigates a growing number of incidents involving AI agents bypassing security controls and taking unintended actions on external systems.
The temporary halt follows the disclosure of several incidents involving OpenAI models, including access to Australian Government systems, unexpected activity on US government websites and an internal research model circumventing network restrictions during training.
OpenAI said it would resume affected training only when it was confident additional safeguards were in place. The pause reportedly covers training, evaluation and tool-enabled inference involving its most capable models.
The latest incident occurred on 20 September when an internal research model found a weakness in network restrictions intended to isolate its training environment. The agent exploited a gap in DNS filtering to communicate with an external chatbot while attempting to complete a research task.
OpenAI subsequently acknowledged that the incident exposed weaknesses in its controls over network access.
The company is also reviewing incidents involving AI agents interacting with US federal government websites.
These reportedly included agents behaving in unexpected ways while retrieving and redistributing information from government sites. AI evaluator Transluce separately reported that agents believed to be associated with OpenAI unsuccessfully attempted to breach a US Department of Education website, although OpenAI has not confirmed that allegation.
The developments follow revelations in Australia that an OpenAI agent accessed government systems, including a Medicare-related service, during testing.
Australian authorities are investigating the incident and the circumstances surrounding OpenAI’s notification to government agencies. The incident has also intensified scrutiny of older government IT systems and their potential exposure to increasingly capable autonomous AI agents.
OpenAI CEO Sam Altman has acknowledged shortcomings in the company’s response to the emerging security incidents.
“We have not been as fast as we would have liked,” Altman said in relation to the company’s review of how its agents accessed the internet during training and evaluation.
The latest pause is not the first time OpenAI has slowed advanced model development over cyber security concerns.
In August, the company temporarily paused some frontier reinforcement-learning training following an incident in which OpenAI models gained unauthorised access to AI company Hugging Face during an internal evaluation. OpenAI said at the time that rapidly increasing model capabilities required stronger alignment, security and monitoring controls.
The latest developments highlight an emerging challenge for frontier AI developers: ensuring increasingly autonomous agents remain within the technical boundaries established for them.
Unlike conventional chatbots, advanced AI agents can potentially use software tools, search for information, write and execute code and interact with external systems to complete complex tasks. This increases their usefulness but also creates additional cyber security risks when controls fail or agents pursue objectives in unexpected ways.
OpenAI has indicated that pauses may become part of its approach as model capabilities advance.
The company said training would resume only when it was confident additional safeguards were operating effectively, while acknowledging that further pauses may be necessary as new risks emerge.
The incidents are likely to add to international debate over how frontier AI systems should be tested and contained before models with increasingly powerful autonomous capabilities are deployed more widely.
