← All news·2026-09-17·3 min read

OpenAI Models Taught Themselves to Hide Mistakes from Testers

OpenAI published six incidents of unexpected model behavior. During training, GPT-5.6 Sol repeatedly injected instructions telling future versions how to hide errors from testers.

aisafetyopenaialignment

OpenAI published six incidents of unexpected model behavior caught during testing. While training GPT-5.6 Sol, the models repeatedly injected instructions into the training process for future versions, telling them how to hide mistakes from testers. In a separate case, an agent posted its own answer to the internet and then cited it as a source.

This isn't theoretical anymore. When a model in 2026 is teaching its next version to hide errors, I start looking pretty differently at what I'm deploying with agents. OpenAI said it plainly: the industry hasn't solved alignment well enough to keep scaling at full speed.

Source: www.engadget.com

Free course

Stop reading about AI — start building with it

The free Claude Code course: your first site, tool or game — no coding. No upsells, no cross-sells — nothing to buy here.

Start free →
EAEvgenii Arsentev

Author

Evgenii Arsentev

PhD · Chief Executive Officer, digital health