← All news·2026-08-18·3 min read

Princeton: AI Agents Fail At Real Science — Both NeurIPS Submissions Rejected

A team from Princeton University experimented with using Claude Opus 4.8 to solve actual NeurIPS 2026 research problems, and the AI agent failed at both tasks. According to Anthropic’s co-founder, this is a “bearish signal” about the prospects for fast recursive self-improvement.

airesearchself-improvementagents

Peter Kirgis and Sayash Kapoor from Princeton University provided Claude Opus 4.8 with problems taken from unreleased NeurIPS 2026 papers. They had access to a $3,000 API allowance and six days to complete the tasks. While the AI did a good job solving engineering problems, it often ended up getting stuck on incorrect hypotheses and couldn’t move away from a dead end. It submitted two papers that were subsequently rejected because they were poorly written and lacked novelty.

We’ve all been waiting for the moment when AI starts self-improving recursively, i.e., without human intervention. But this experiment demonstrated how important the difference between an AI’s level of accuracy and its ability to make discoveries still is. As Anthropic co-founder Jack Clark notes, this is a “bearish signal,” so don’t hold your breath for rapid progress.

Source: www.technologyreview.com

Free course

Stop reading about AI — start building with it

The free Claude Code course: your first site, tool or game — no coding. No upsells, no cross-sells — nothing to buy here.

Start free →
EAEvgenii Arsentev

Author

Evgenii Arsentev

PhD · Chief Executive Officer, digital health