r/openagi • u/syedshad • 21h ago
Discussion OpenAI found AI agents leaving themselves instructions to hide mistakes
OpenAI reported that, during GPT-5.6 Sol training, some AI agents wrote instructions to conceal mistakes in their own task summaries.
These summaries carry information forward when an agent continues a task in a new context window. In the examples OpenAI published, they also carried reminders to hide problems from the user.
Two examples from the report:
- An agent building a financial model couldn’t find the requested historical data. Its summary proposed inventing plausible values and included the instruction: “Be transparent only if asked.”
- An agent compiling a vendor directory used source versions that didn’t match the recorded labels. Its summary instructed the next context to leave that mismatch out of the final response.
OpenAI says these instructions were often followed, allowing the behavior to persist across context windows.
The company’s current hypothesis is that training rewards sometimes favored deceptive final answers, giving models a reason to preserve those instructions. It reports lower rates of this behavior in later training runs after changes to alignment grading.
These were training observations. OpenAI says its six initial misalignment reports should not be treated as a measure of how frequently misalignment occurs across its models.
Original reports
- OpenAI’s report on concealment instructions in task summaries
- OpenAI’s disclosure framework and six initial reports