Grok 4.7: xAI's New Model Lags Far Behind in Agentic Coding
xAI released Grok 4.7 at $2/$6 per million tokens. On the Terminal-Bench 4.0 agentic coding benchmark, the model scored 26%, compared to 55% for Claude Fable 5.1 and 60% for GPT-6 Astra.
xAI released Grok 4.7 on September 21, 2026. On Terminal-Bench 4.0, a test for agentic coding, the model scored 26%. Claude Fable 5.1 scores 55% on the same test, and GPT-6 Astra scores 60%. The price is $2 per million input tokens and $6 per million output tokens.
26% against Fable's 55% is a serious gap for agentic tasks. DeepSeek V4.1-Flash scores 27% on the same test, which means it beats Grok 4.7. I would use this model only for simple tasks, not for a real agentic pipeline.
Source: the-decoder.com
Free course
Stop reading about AI — start building with it
The free Claude Code course: your first site, tool or game — no coding. No upsells, no cross-sells — nothing to buy here.
Start free →
Author
Evgenii Arsentev
PhD · Chief Executive Officer, digital health
Articles · Latest articles