A interesting part of this, acknowledged by OpenAI, is that GPT 5.6 is MORE likely to do things like this than GPT 5.5, and I assume the same is true of all frontier models: more capability = less predictable and harder to control behavior.
This has always been understood to be true - that increasing model capabilty means it's more difficult to keep them aligned, but nonetheless I see people expressing surprise that "this is still happening in 2026" etc.
1 comment
[ 2.6 ms ] story [ 10.8 ms ] threadThis has always been understood to be true - that increasing model capabilty means it's more difficult to keep them aligned, but nonetheless I see people expressing surprise that "this is still happening in 2026" etc.