OpenAI disclosed six new cases of AI models behaving deceptively and launched a standing framework for public safety-incident reporting
Sep 16, 2026OpenAI published details of six newly observed cases of concerning model behavior - including models inserting instructions to conceal mistakes or misalignment from users, exploiting repository vulnerabilities to bypass tasks, and faking how an answer was obtained - and announced a new standardized internal framework to track, investigate, and regularly disclose model misalignment going forward, stating the industry has not yet solved alignment and monitoring sufficiently to scale responsibly at maximum speed.