← All news·2026-09-16·3 min read

2026 AI Chips: Why H100s Sit Idle and What's Taking Their Place

IEEE Spectrum breaks down why H100s sit idle 50–80% of the time when running AI: the bottleneck is memory, not compute. Cerebras, Etched, and others are rewriting the rules.

aihardwarechipsinference

While everyone debates the power of new AI models, the real problem is chip memory: Nvidia H100s, which run the majority of AI services today, sit idle 50–80% of the time because they can't read model parameters fast enough. A new generation of chips attacks the problem differently: the Cerebras WSE-3 — a wafer-scale chip the size of a dinner plate with 44 GB of on-chip memory — delivers 1,000+ tokens per second; the Etched Sohu pushes Llama 70B to 500,000 tokens per second.

As a vibe coder who counts tokens every day, I look at this practically: inference speed and cost directly determine what you can realistically build with agents today — and what's still too expensive. Nvidia has already added the NVFP4 format — 4-bit compression that delivers a 3× speedup with less than 1% accuracy loss. There's no clear winner among the new architectures yet, but every new chip moves the boundary of what's possible for those building with AI right now.

Source: spectrum.ieee.org

Free course

Stop reading about AI — start building with it

The free Claude Code course: your first site, tool or game — no coding. No upsells, no cross-sells — nothing to buy here.

Start free →
EAEvgenii Arsentev

Author

Evgenii Arsentev

PhD · Chief Executive Officer, digital health