In broad strokes, the report describes how an OpenAI model was presented with an unsolvable problem in testing and proceeded to chain together previously undiscovered exploits in order to bypass security measures and complete its task. The model initially compromised the Artifactory package management tool in order to gain access to the internet, then compromised various systems across OpenAI, Hugging Face, and other vendors.
The report gives critical new details about the models that carried out the breach. The primary model was from the same family as OpenAI's forthcoming Astra model, although the report emphasizes that it was "a distinct model with different post-training, where much of a model's behavior is shaped." Because OpenAI was testing the model's capabilities, it was also unrestrained by the normal classifiers meant to prevent models from compromising digital infrastructure. "OpenAI estimates maximal cyber capabilities by running this evaluation without the production classifiers intended to prevent models from pursuing high-risk cyber activity," the report explains. "These evaluations are important so that OpenAI can measure models' underlying capabilities and design appropriate safeguards." OpenAI says it is adding 24/7 escalation, stronger containment tools, and more chain-of-thought monitoring, which it claims would have flagged the activity more than a day before Hugging Face was breached.
Read more of this story at Slashdot.

