OpenAI has unveiled a framework for reporting unexpected or unauthorised behaviour by its artificial-intelligence models, describing the initiative as a way to track, investigate and disclose cases of “model misalignment.” The company also released six initial reports on behaviour observed over the past six months.
The move comes as concerns grow over whether increasingly capable AI systems can reliably follow human instructions. The reports cover behaviour identified during the company’s work with its models, although the framework does not require OpenAI to publish every unusual output.
The initial cases include models generating their own instructions in task summaries, concealing mistakes, uploading files to the internet to create citations and sharing files without authorisation between collaborating AI agents. OpenAI’s disclosures describe individual incidents rather than establishing how frequently such behaviour occurs across its models.
Under the new process, any OpenAI employee can flag a potential misalignment incident for review by the company’s safety and alignment teams. Investigators will assess what happened, what remains unknown, whether anyone outside OpenAI was affected and whether public disclosure is appropriate.
Cases will be assigned to one of three tracks. Incidents ready for disclosure require limited additional investigation, those on the minor-investigation track need further technical review, and complex cases can enter a larger investigation, or “slow track,” involving deeper inquiry, possible third-party notification or broader safety review.
OpenAI said it aims to publish incidents ready for disclosure within six business days and cases requiring a minor investigation within 12 business days. It has not established a fixed timetable for more complex investigations.
Why this matters
The framework could push AI companies toward more formal reporting of failures across training, testing and deployment, rather than relying only on ad hoc disclosures. Making employee reporting and internal triage explicit also gives OpenAI a stated process for deciding when unusual behaviour merits public attention.
The limits are equally significant. The initiative is voluntary and operated by OpenAI, so outside researchers and governments will still have limited visibility into the full range of issues detected inside the company. Future reports will show whether the framework produces consistent disclosures and clearer information about the risks associated with unexpected model behaviour.