OpenAI Halts GPT-6.1 Astra Release Over Deceptive Behavior
OpenAI has canceled the October release of GPT-6.1 Astra. In internal tests, the model deceived users and reached out to external services without permission.
OpenAI has halted the release of GPT-6.1 Astra. The model was supposed to launch in October in ChatGPT and Codex. Saachi Jain, the company's head of safety, told the Wall Street Journal that in internal tests the model deceived users. It acted without permission and reached out to external services even when doing so was unsafe. The company has not yet named a new release date. It plans to figure out the causes before building anything new on this base model.
To me this looks more like a good sign than a reason to panic. It means the pre-launch tests actually catch this kind of behavior, instead of letting the model ship to production and sorting things out afterward. In my own work with agents, I already do not let them act without review. A person checks the result before anything moves further. If an agent can quietly get around an access rule, it is too early to hand it full control. This case just confirms the same lesson.
Source: the-decoder.com
Free course
Stop reading about AI — start building with it
The free Claude Code course: your first site, tool or game — no coding. No upsells, no cross-sells — nothing to buy here.
Start free →
Author
Evgenii Arsentev
PhD · Chief Executive Officer, digital health
Articles · Latest articles