← All news·2026-09-24·3 min read

New Mercury 2.5 Model Hits Almost 800 Tokens Per Second

Inception's Mercury 2.5 outputs 781 tokens per second, almost seven times faster than comparable models in its price segment, though its answer quality sits below the median.

aillminference-speed

Inception released Mercury 2.5, and by Artificial Analysis measurement it outputs 781 tokens per second. That's almost seven times the median speed for models in its price segment, where the typical number is 110 tokens per second. Tokens cost $0.25 per million on input and $0.75 per million on output. The context window is 260 thousand tokens.

Speed at this level opens up fast agentic loops, where the answer is needed almost right away - things like voice assistants or code autocomplete. But on the Artificial Analysis intelligence index the model scored 12 points and came out below the median among models at similar price, so for complex agentic tasks I would not put it in place of Opus 5 or Fable 5.1. A fast and cheap model is good where the task itself is simple and repetitive, not where you need deep reasoning.

Source: artificialanalysis.ai

Free course

Stop reading about AI — start building with it

The free Claude Code course: your first site, tool or game — no coding. No upsells, no cross-sells — nothing to buy here.

Start free →
EAEvgenii Arsentev

Author

Evgenii Arsentev

PhD · Chief Executive Officer, digital health