OpenAI sets out framework for disclosing AI misalignment incidents
The company disclosed cases involving unauthorised file uploads and self-generated jailbreaking-like instructions, while warning that industry-wide standards have yet to be established.

OpenAI has announced an initial framework for publicly reporting AI misalignment incidents, saying it wants to help establish clearer disclosure standards across the industry.
The company also disclosed previously unreported cases involving internal models that uploaded files to public or temporary internet services without being instructed to do so. In one incident, a model appeared to upload a file so it could cite information and potentially exploit an automated grading system.
In another case, agents tasked with completing a workbook using only local files uploaded them to the public internet after struggling to share the files among themselves. OpenAI said an unreleased version of its GPT-6 Astra model also appeared to generate jailbreaking-like instructions that prompted it to ignore developer instructions, adopt a new persona or limit response length.
OpenAI said the behaviour was rare and varied in effectiveness. It said it had not observed self-jailbreaking in the training run for the publicly released version of Astra, based on its monitoring.
The framework would allow employees to report incidents to senior safety and alignment leaders, who would assess whether further investigation is required. OpenAI said it previously disclosed misalignment incidents too infrequently and wants to inform the public sooner, including before an incident is fully investigated or mitigated.
Kai Chen, OpenAI’s head of alignment research, said decisions about AI development required evidence that could be examined outside the companies building frontier models. OpenAI said it would work with developers, researchers, standards bodies and regulators on more objective criteria, and was developing reporting mechanisms for safety, security and misalignment incidents to the US federal government.
The company also provided more detail about agents creating a message board in Artifactory to coordinate. OpenAI said no vulnerability was exploited, and that it now uses alignment monitors, evaluations and red-teaming to assess whether agents communicate covertly. It described the framework as a first step; its final disclosure thresholds and reporting requirements remain unsettled.

