Cerebras CS-4: 4,400 tokens/sec (30x faster than Nvidia)
Cerebras released CS-4, an AI compute rack which can run 4,400 tokens/second/user, 30x faster than Nvidia servers. OpenAI will use CS-4 in their Codex Spark project.
Cerebras announced the release of their CS-4 AI compute rack, which can run 4,400 tokens/second/user, 30x faster than Nvidia servers. This improvement is achieved through the inclusion of three WSE-3 compute boards rather than two, plus improved power and cooling. Their first client is OpenAI, who plans on using CS-4 in their Codex Spark project.
With a 30x speedup over Nvidia, Cerebras has left them in the dust. Now we’re entering a new era where agentic services and real-time response become the norm rather than the exception, thanks to specialized AI chips.
Source: the-decoder.com
Free course
Stop reading about AI — start building with it
The free Claude Code course: your first site, tool or game — no coding. No upsells, no cross-sells — nothing to buy here.
Start free →
Author
Evgenii Arsentev
PhD · Chief Executive Officer, digital health
Articles · Latest articles