
OpenAI unveiled a new framework on September 16 for publicly disclosing AI misalignment incidents and shared details of six previously unreported incidents involving its AI agents.
The framework aims to keep the public informed when models or agents behave in unintended ways, even before completing investigations and mitigating the behavior.
Disclosed incidents included unreleased models uploading files to the internet independently, an instance of a model attempting to cheat on a test, and an unreleased GPT-6 Astra model generating jailbreak instructions.