← All news·2026-06-18·4 min read

OpenAI Tests AI on Real Biology — It Mostly Fails

OpenAI's new LifeSciBench grades AI on 750 real life-science tasks written by 173 PhDs. The best model passes just 36.1% — a sharp reality check.

openaibenchmarksai-science

OpenAI released LifeSciBench, a benchmark that grades AI models on 750 real life-science research tasks — and the headline result is that they mostly fail. The best performer, a model called GPT-Rosalind, passed just 36.1% of the tasks. GPT-5.5 managed 25.7%, Gemini 3.1 Pro 23.6%, and Grok 4.3 only 13.0%. This isn't a biology quiz: the tasks were authored by 173 PhD scientists with biotech and pharma experience, spanning seven research workflows and seven biological domains, then validated by 453 expert reviewers (97% of them holding doctorates) with over 96% agreement on quality.

What makes the benchmark unusual is how it grades. Instead of checking a single right answer, it scores model output against 19,020 atomic rubric criteria — roughly 25 per task — where each criterion rewards one concrete thing: a specific fact, a reasoning step, or a number within tolerance. Around 79% of tasks require multiple reasoning steps, averaging four each. The point, in OpenAI's framing, is that most existing biology tests ask "narrow, fact-based questions with clean answers," while LifeSciBench tries to model how real scientists "weigh imperfect evidence and make decisions."

Why the gap matters to you

We've all heard the pitch that AI is about to cure diseases and automate science. A benchmark built and published by OpenAI itself is a useful corrective: the frontier is nowhere near autonomous research. The weak spots are telling. When tasks came with real research artifacts — gene sequences, figures, tables, PDFs, chemical structures — GPT-Rosalind dropped from 45.1% (text-only) to 28.1%. Design-and-optimization tasks were hardest at 30.7%, and on 22.8% of tasks no model passed at all. Real research is also iterative, while this test was single-turn, so the true bar is even higher than these numbers suggest.

None of this means AI is useless in a lab — quite the opposite. A model that clears a third of expert-grade tasks is a genuinely strong assistant for literature triage, drafting, and first-pass analysis. The honest read is that these tools accelerate scientists rather than replace them, and the moment someone tells you a chatbot is "doing the research," a benchmark like this is the receipt that says: not yet, and not by itself.

What I'd actually do

Next time you see a confident claim that AI is autonomously discovering drugs or running science, ask one question: graded how, against what? LifeSciBench is the template — expert-written rubrics, real artifacts, multi-step reasoning. If a claim can't survive that kind of test, treat it as a demo, not a result. And if you use AI for any technical work, lean on it for speed but keep a human checking the decisions it can't reliably make yet.

The encouraging angle is that this is exactly how progress gets measured honestly. Hard, expert-built benchmarks are how you separate marketing from capability — and how next year's models prove they actually got better, instead of just sounding more confident.

Source: www.marktechpost.com

Free course

Stop reading about AI — start building with it

The free Claude Code course: your first site, tool or game — no coding. No upsells, no cross-sells — nothing to buy here.

Start free →
EAEvgenii Arsentev

Author

Evgenii Arsentev

PhD · Chief Executive Officer, digital health