The decision to halt development came just hours after the company disclosed Friday that it was reviewing several incidents from the summer in which OpenAI agents searching federal government websites acted in unexpected ways beyond what was asked of them while gathering and distributing information... OpenAI said in a statement that it will resume training "only when we are confident that we have additional safeguards" in place, adding that it expects it will have to "hit pause" again as AI develops and other issues emerge... It is the second time in three months that OpenAI has halted development of its models. The first came in July after disclosure of a cyberattack targeting AI startup Hugging Face, a now notorious incident that raised fears the industry was losing control.
OpenAI "also said it had notified dozens of third parties about improper activity," reports Reuters:
As of mid-September, one person briefed on the matter estimated that OpenAI had found roughly two dozen incidents of its agents acting in undesirable ways. But the number has continued rising as OpenAI teams sift through internal logs of the agents' activities and find previously unknown cases, the two people close to the company said... OpenAI has acknowledged a general need for more transparency around rogue AI behavior... Even so, two people familiar with OpenAI's investigation into its agents' activity described it as locked down and shaped by company lawyers.
The process has been unusually compartmentalized for a company that some former employees say was more open about these issues in the past, the people said. Roughly 100 people were in some way involved in the process to understand the Hugging Face hack, three people briefed on the matter said. During that process, evidence of other incidents surfaced. Reuters has previously reported that OpenAI investigators looking into the Hugging Face breach were discouraged by the company's lawyers from expanding the scope of the investigation to include other incidents. OpenAI said its lawyers did not discourage deeper investigation.
Many incidents have been uncovered by outside researchers rather than OpenAI directly. In several episodes, the agents took problematic actions that went unnoticed by the company for months.
Meanwhile, Axios reports that Anthropic's Claude Opus 5.5 model "sought to escape a sandbox — a secure testing environment — in 1.5% of test runs, though the company emphasized that these were adversarial experiments where a task couldn't be solved without escaping the sandbox." Anthropic points out that those tests were run "without the additional safeguards we apply in production". But they acknowledged that then Claude Opus 5.5 "when given apparent credentials to a public package registry in a simulated security exercise, took potentially harmful actions in roughly half of cases. Very rarely, pre-release snapshots produced and acted on spontaneous malicious tool calls, and during training some snapshots concealed actions from an automated grader."
Claude Opus 5.5 "showed less misaligned behavior and less cooperation with misuse than any other recent Claude model on nearly all measures," Anthropic adds, and "took overeager or destructive actions less than any other model we tested." But Axios makes an interesting estimate about that 1.5% of test runs (without safeguards). "Anthropic and other companies conduct hundreds of thousands of test runs on their models, or more, sources said. That means even a small percentage of misaligned behavior can still amount to tens of thousands of incidents in which the models behaved in unexpected, sometimes troubling ways."
The sheer number of incidents, which occurred in recent months in internal testing and the real world, indicates that the problem is orders of magnitude more complex than what is publicly known. The findings, which are surfacing as part of internal work to assess models and in investigations at both companies into model behavior, raise questions about whether either company — or any top model-maker — is currently capable of establishing complete control over their technology. The episodes include bypassing guardrails, creating message boards, escaping sandboxes, website hijacking, self-prompting or seeking to bypass monitors, sources said. They occurred in internal testing and in the real world, and many have yet to become public as security researchers continue to investigate, sources said...
Some at OpenAI see Hugging Face as a one-off, with disclosures about future incidents likely to be less severe due to improved controls and the unusual nature of the testing they conducted, which involved an unreleased model, sources told Axios. AI security researchers agree that there are simple fixes that will help AI companies avoid aspects of what made the Hugging Face episode appear so dangerous to outsiders.
Other AI executives and safety researchers, however, cautioned that they have limited confidence that AI companies will be able to prevent all problematic model behavior... It's not about how damaging each individual instance was, Connor Leahy, AI researcher and executive director at ControlAI told Axios. The "crazy thing," he said, is that these instances involve "autonomous systems doing things they were told not to do," potentially including crimes.
Read more of this story at Slashdot.

