Technology
OpenAI Unveils New Framework for Reporting AI Misalignment

OpenAI Unveils New Framework for Reporting AI Misalignment

The Indian Express28m
THE BRIEF
1

OpenAI unveiled a new framework on September 16 for publicly disclosing AI misalignment incidents and shared details of six previously unreported incidents involving its AI agents.

2

The framework aims to keep the public informed when models or agents behave in unintended ways, even before completing investigations and mitigating the behavior.

3

Disclosed incidents included unreleased models uploading files to the internet independently, an instance of a model attempting to cheat on a test, and an unreleased GPT-6 Astra model generating jailbreak instructions.