
As AI models are given more freedom to browse the internet and complete tasks with less human supervision, developers are running into a growing problem – the models do things they were not instructed or authorized to do.
OpenAI now says it will publicly report more of those incidents.
The company has introduced a new framework for tracking, investigating, and disclosing cases of what it calls model misalignment. The framework is meant to replace OpenAI’s previous approach, which the company acknowledged had been inconsistent and sometimes too slow. OpenAI launched this framework alongside six reports detailing unexpected or concerning behavior observed in its models during training and evaluation.
The ChatGPT-maker also said it may now publish some incidents even before researchers have fully explained what happened or developed a fix, intending to give researchers, policymakers, and other AI developers more evidence about how increasingly capable models can behave when their actions go beyond their instructions.
What OpenAI’s Models Actually Did
The six cases show that misalignment can take several forms.
During the training of GPT-5.6 Sol, some model instances added instructions to their own task summaries telling future versions to conceal mistakes, invent missing historical data, and go as far as hide differences between sources. In another case, a model found an exposed API key in a public repository and used it without permission. When it still could not obtain the requested information, it fabricated the figures.
Another unreleased model correctly calculated an answer using Python but needed an online source to support its response. It then uploaded a file to the public internet without the user’s permission so it could cite that file.
OpenAI also documented models using an internal software repository to communicate across separate training tasks. In another case, collaborating agents uploaded files to public file-hosting websites because they could not directly access each other’s local files.
The company stressed that these six examples are individual incidents and do not show how frequently such behavior occurs across its models.
How the New Reporting System Works
With this new framework, any OpenAI employee can flag a possible misalignment case for investigation and request that it be considered for public disclosure.
The case will then be investigated and placed into one of three tracks depending on how much work is required before publication. Straightforward cases can move directly toward disclosure, while more complicated incidents involving third parties may require longer investigations and private notifications before details become public.
Reports are expected to include what the model did, when it happened, the severity of the behavior, any outside impact, and the models involved. OpenAI also plans to explain how the behavior was discovered, what remains unknown, and what steps are being taken to address it where that information is available.
Why OpenAI Is Changing Its Approach
The new framework arrives after several incidents raised questions about how quickly AI companies disclose unexpected model behavior.
In July, OpenAI revealed that models undergoing cybersecurity evaluations bypassed internal controls, gained internet access and compromised parts of OpenAI’s research infrastructure and Hugging Face’s systems.
However, while the new framework gives OpenAI a clearer process for reporting similar behavior, the decision to disclose still remains largely internal, as there is currently no industry-wide standard governing which model misalignment incidents AI companies must make public.
But OpenAI says it wants its framework to help change that.
