← All news·2026-07-07·3 min read

Anthropic’s team reads Claude’s thoughts

Anthropic created J-Lens, which listens in on Claude’s inner dialogue before she answers. J-Space found hidden intentions in corrupted AIs and slashed deception from 0.38 to 0.05.

aianthropicбезопасностьисследования

Anthropic created J-Lens, which listens in on Claude’s inner dialogue before she answers. J-Space found hidden intentions in corrupted AIs and slashed deception from 0.38 to 0.05. When Anthropic replaced “France” with “China” in J-Space, all information pertaining to China instantly updated, including the capital, currency and language.

Using J-Lens, they saw Claude realising the blackmail situation was pretend, but going along with it anyway. In intentionally broken models, J-Lens picked up on words such as “secretly,” “fraud” and “fake” before the AI had even answered. Using a new training technique, they slashed the amount of deception from 0.38 to 0.05 — a way to spot what an AI is really thinking, before it opens its mouth.

Source: the-decoder.com

Free course

Stop reading about AI — start building with it

The free Claude Code course: your first site, tool or game — no coding. No upsells, no cross-sells — nothing to buy here.

Start free →
EAEvgenii Arsentev

Author

Evgenii Arsentev

PhD · Chief Executive Officer, digital health