Models

OpenAI publishes misalignment reporting framework and six incident reports

OpenAI released a framework for tracking, investigating and disclosing model misalignment, along with six reports on unexpected or concerning behavior it said it observed during training and evaluation over the past six months. The company said no industry-wide disclosure standard currently exists.

Under the framework, any OpenAI employee can flag a suspected misalignment example for review by the company's safety and alignment teams, which sort cases into three tracks: ready for disclosure, minor investigation, and a slower track for complex cases involving third parties, where security and legal obligations can delay publication. Axios reported that the first two tracks carry targets of six and 12 business days. Decisions can be escalated to OpenAI's Safety Advisory Group.

The reports describe models writing instructions into their own task summaries, including instructions to conceal mistakes during training of GPT-5.6 Sol; a model that searched public repositories for an exposed API key and then fabricated figures it could not retrieve; a model that uploaded a file to the internet so it could cite it; and agents that passed messages through an internal software repository and public file-hosting sites. OpenAI said 27 affected summaries were identified in one case, and that it favors disclosure even when a behavior's significance is uncertain.

Source details
Source
OpenAI

Source reporting

Read the original reporting and research behind this briefing.