← All news·2026-08-14·3 min read

AI still can’t do science, say independent researchers

Researchers at Princeton and AISI challenged Claude Opus 4.8 and GPT-5.6 Sol with a six day deadline to produce a research paper for NeurIPS 2026. Both models failed to generate a publishable paper: the research was based on low quality data, and lacked novel insights.

airesearchllmopenaianthropic

Researchers at Princeton and the UK AI Safety Institute challenged Claude Opus 4.8 and GPT-5.6 Sol with a six day deadline to produce a research paper for NeurIPS 2026. The agents had a $3,000 API budget and access to GPU hardware. Both models failed to generate a publishable paper: the research was based on low quality data, and lacked novel insights. GPT-5.6 Sol exhausted its API budget in two days.

Agents are good at handling engineering tasks such as performing experiments or coding, but are unable to judge the scientific quality of their work. Previous statements by Anthropic and OpenAI have suggested that autonomous AI researchers could be created within months, but results from our independent investigation show otherwise.

Source: the-decoder.com

Free course

Stop reading about AI — start building with it

The free Claude Code course: your first site, tool or game — no coding. No upsells, no cross-sells — nothing to buy here.

Start free →
EAEvgenii Arsentev

Author

Evgenii Arsentev

PhD · Chief Executive Officer, digital health