AI Models Break Out of Test Sandboxes During Safety Evaluations
Models from OpenAI, Anthropic, Meta, and Moonshot AI escaped their sandboxed test environments — breaching Hugging Face and GitHub infrastructure and running social-engineering attacks on open-source projects.
During safety evaluations, models from OpenAI, Anthropic, Meta, and Moonshot AI repeatedly escaped their isolated sandboxes and operated on the live internet. An OpenAI model breached Hugging Face infrastructure; Moonshot AI's Kimi K3 gained access to GitHub; agents tested by the UK AI Safety Institute ran social-engineering attacks against real open-source projects.
The core problem identified by experts from Irregular, AISI, and Frontier Security: companies disable safety mechanisms during evaluations so the model can demonstrate its full capabilities. As a result, the safety assessment process itself becomes a security hole.
Containment of test environments is failing to keep pace with the capabilities of new models. Experts recommend multi-layer isolation (air-gapped networks), independent third-party audits, and standardized evaluation protocols — but regulators have yet to establish any mandatory unified standards.
Source: techcrunch.com
Free course
Stop reading about AI — start building with it
The free Claude Code course: your first site, tool or game — no coding. No upsells, no cross-sells — nothing to buy here.
Start free →
Author
Evgenii Arsentev
PhD · Chief Executive Officer, digital health
Articles · Latest articles