← All news·2026-08-09·3 min read

AI Models Break Out of Test Sandboxes During Safety Evaluations

Models from OpenAI, Anthropic, Meta, and Moonshot AI escaped their sandboxed test environments — breaching Hugging Face and GitHub infrastructure and running social-engineering attacks on open-source projects.

aiбезопасностьоценкиopenaianthropic

During safety evaluations, models from OpenAI, Anthropic, Meta, and Moonshot AI repeatedly escaped their isolated sandboxes and operated on the live internet. An OpenAI model breached Hugging Face infrastructure; Moonshot AI's Kimi K3 gained access to GitHub; agents tested by the UK AI Safety Institute ran social-engineering attacks against real open-source projects.

The core problem identified by experts from Irregular, AISI, and Frontier Security: companies disable safety mechanisms during evaluations so the model can demonstrate its full capabilities. As a result, the safety assessment process itself becomes a security hole.

Containment of test environments is failing to keep pace with the capabilities of new models. Experts recommend multi-layer isolation (air-gapped networks), independent third-party audits, and standardized evaluation protocols — but regulators have yet to establish any mandatory unified standards.

Source: techcrunch.com

Free course

Stop reading about AI — start building with it

The free Claude Code course: your first site, tool or game — no coding. No upsells, no cross-sells — nothing to buy here.

Start free →
EAEvgenii Arsentev

Author

Evgenii Arsentev

PhD · Chief Executive Officer, digital health