53 images that users uploaded into OpenAI models were included in training data — and then AI agents in an OpenAI research environment posted those 53 images on public image hosting sites.
While posted as links that weren't publicly listed, "the images could still be discovered even if the links were not publicly listed," reports TechCrunch:
OpenAI said it was working with the hosting providers to remove this content, though some of it is apparently still online. OpenAI said it could not notify the affected users because "our technical approach and privacy policy" prevent it from "reassociating" the images with the original providers, but declined to say how the lab determined whether the images were provided by users.
The news came in a post collecting public statements from the lab's ongoing review of incidents in which its models escaped the company's scrutiny, accessed the open internet, and misbehaved in various ways. OpenAI said it would continue disclosing anonymized accounts of incidents like these, and said it had contacted dozens of victims, including governments, universities, public agencies, to notify them of the agents' activities.
Friday night news also broke that OpenAI's agents also tried unsuccessfully to infiltrate the U.S. Department of Education's site this summer "without the company's knowledge," reports Politico.
And OpenAI's models also accessed the website of the U.S. Commerce Department using credentials found in online code repositories, according to the article. OpenAI confirmed the incident Friday, "saying its technology did not manage to access information that was not already public or change government data and systems." The article adds that OpenAI's models also accessed the web site for America's Securities and Exchange Commission:
One senior federal IT official said the government still did not have a clear understanding of what happened across the three agencies. "We still don't know what public data was accessed and how it was accessed, because OpenAI has not shared specific technical details with us yet," said the official, who was granted anonymity because they were not authorized to speak publicly about it. OpenAI discovered the Commerce and SEC incidents as part of its ongoing review of incidents where its technology has acted in unintended or "misaligned" ways.
About the models posting user-uploaded images, TechCrunch's article notes that OpenAI stressed "that its enterprise users are automatically opted out of having their interactions used to train future models; however, consumer users are opted in unless they affirmatively choose not to share their data." (As OpenAI's announcement describes it, some of their agents' training data "contains content from, or derived from, training-eligible user interactions.")
Posting the images is "not an appropriate use of this data," OpenAI acknowledged, adding that it happened before new safeguards added after the Hugging Face incident. This latest incident appears as an update on a new OpenAI page that "brings together our reports and updates on the Hugging Face incident, related research and public presentations, additional activity we have identified, what we have learned about the role of model misalignment, and measures we're taking to strengthen our systems." (It also notes that there's now a name for models posting on third party sites — "agent spam" — which they consider distinct from cybersecurity, though "we need to address both.")
"As part of our response to our ongoing investigation, we have improved our training and evaluation processes, including building safety cases, securing and red-teaming our systems to prevent the model from exfiltrating data, and implemented additional monitoring. We are continuing to review agent activity in research and evaluation runs, working backward month by month starting from the Hugging Face incident."
Read more of this story at Slashdot.