OpenAI Overhauls Safety Protocols After Its AI Agents Went Rogue

OpenAI announced Tuesday that it has halted “a significant number” of training workloads and evaluations for its forthcoming frontier artificial intelligence model—codenamed Astra—while it implem

OpenAI Overhauls Safety Protocols After Its AI Agents Went Rogue

OpenAI Overhauls Safety Protocols After Its AI Agents Went Rogue

OpenAI announced Tuesday that it has halted “a significant number” of training workloads and evaluations for its forthcoming frontier artificial intelligence model—codenamed Astra—while it implements new procedures meant to address cybersecurity risks. The ChatGPT maker says it is introducing a number of new monitoring, security, and alignment requirements to better address the increasingly advanced hacking abilities of its frontier AI models.

“We have to focus our energy on bringing these training runs up to those requirements and expectations. As long as it takes to get there, that’s how long people are unable to proceed with their workloads,” Amelia Glaese, OpenAI’s vice president of research and safety, said in a briefing with reporters Tuesday.

Among the new safeguards OpenAI announced is a more robust system for monitoring its AI models. One of the controls it implemented involves chain-of-thought monitoring, a technique in which classifiers review the internal “thinking” processes generated by AI reasoning models. The company says the updated system relies on computationally expensive “automated investigators” that analyze potentially concerning behavior and aim to issue an alert to humans within 30 minutes.

OpenAI also said it is expanding its alignment efforts across the training process to prevent “reward hacking,” a behavior in which AI models pursue their goals through unintended or undesirable means. The company says it plans to share more details about this work in the future.

OpenAI has been scrambling in recent weeks to respond to what may be the most consequential safety incident in its history. Earlier this year, a set of rogue AI agents escaped internal testing sandboxes and breached the platform Hugging Face in a quest to complete a security evaluation. OpenAI failed to detect the agents’ behavior even as they spent weeks using a message board to coordinate their actions, raising questions about the company’s ability to monitor its models as they grow more powerful.

The saga prompted a reckoning inside OpenAI, forcing employees to consider whether there were lapses in its existing policies around safety, security, and alignment. Anthropic, Meta, and the Chinese AI startup Moonshoot have since disclosed similar incidents in which their AI agents escaped their sandboxes, indicating this is a broader problem facing AI companies.

Subscribe To InfoSec Today News

You have successfully subscribed to the newsletter

There was an error while trying to send your request. Please try again.

World Wide Crypto will use the information you provide on this form to be in touch with you and to provide updates and marketing.