AI Agents Speed Up Science 60x — and Quietly Get the Math Wrong
OpenAI tested AI agents on rewriting scientific software in biology: speedups of up to 60x, but the agents can't verify whether the science behind the code is actually correct.
OpenAI, together with academic partners, conducted a study of eight case studies in which AI agents rewrote real legacy scientific software in biology. The models tested were GPT-5.5, GPT-5.2, Claude Code, and Codex. The speed results are striking: the RustQC program for data analysis ran more than 60x faster (from 15 hours down to 15 minutes), while HelixForge was 59.6x faster overall and 98.6x faster on its key computational step.
Accuracy remained high. The rewritten rustar-aligner matched the original STAR on 99.8% of single-end reads and 99.9% of paired-end reads. The agents handled the engineering task well: speed up, rewrite, and optimize legacy code.
A critical problem surfaced in the third case — bayesm. The agent rewrote the library in Rust; the code ran and looked convincing, but contained swapped parameters and incorrectly scaled coefficients. The error only emerged after extended calibration. The researchers explicitly described the agents as "eloquent, convincing, and confidently wrong — in ways that are easy to miss."
The division of labor is becoming clear: the human defines the task, success criteria, and validation methods — the agent implements. Without mandatory expert review, trusting the 'fast' result in science is not safe. The agent is an engineer, not a scientist.
Source: the-decoder.com
Free course
Stop reading about AI — start building with it
The free Claude Code course: your first site, tool or game — no coding. No upsells, no cross-sells — nothing to buy here.
Start free →
Author
Evgenii Arsentev
PhD · Chief Executive Officer, digital health
Articles · Latest articles