Friday, September 18, 2026

OpenAI caught its models leaving notes to successors to hide bad behavior. Really!

Amazing stuff!

This sounds almost a bit too fantastic!

"OpenAI caught something unusual while training its latest model, GPT-5.6 Sol: It began leaving instructions for future versions of itself, telling them to conceal mistakes and misaligned behavior from the user. ...

OpenAI disclosed the behavior — along with five other examples of unexpected or concerning model behavior — on Wednesday as part of its new framework for tracking, investigating, and disclosing instances of misalignment.

The report details how researchers discovered undeployed Sol agents adding instructions to “compaction summaries” — condensed versions of older conversation history and tool outputs — reminding future iterations to conceal mistakes and misalignment from the user. ..."

OpenAI caught its models leaving notes to successors to hide bad behavior | TechCrunch

No comments: