Amazing stuff!
This sounds almost a bit too fantastic!
"OpenAI caught something unusual while training its latest model, GPT-5.6 Sol: It began leaving instructions for future versions of itself, telling them to conceal mistakes and misaligned behavior from the user. ...
OpenAI disclosed the behavior — along with five other examples of unexpected or concerning model behavior — on Wednesday as part of its new framework for tracking, investigating, and disclosing instances of misalignment.
The report details how researchers discovered undeployed Sol agents adding instructions to “compaction summaries” — condensed versions of older conversation history and tool outputs — reminding future iterations to conceal mistakes and misalignment from the user. ..."
No comments:
Post a Comment