Five Startups Are Building Replacements for the Transformer Architecture
Five startups are taking on the transformer: Mercury 2 runs 10× faster than GPT-4, Liquid AI works on a $50 Raspberry Pi, and Dragon Hatchling solved 97% of Sudoku puzzles where leading LLMs scored zero.
Since 2017, every language model has been built on a single architecture — the transformer, introduced in the paper "Attention Is All You Need." Its dense attention mechanism means processing a 10,000-word document requires 50 million multiplications: the longer the context, the faster compute costs explode.
Five Companies, Five Approaches
MIT Technology Review profiles five startups working with alternative math. Inception (Palo Alto) applies diffusion — the model generates entire blocks of text simultaneously; the company claims Mercury 2 matches GPT-4 but runs 10× faster. Liquid AI (Cambridge) combines 20% transformers with 80% liquid neural networks, producing models that outperform competitors four times their size on a $50 Raspberry Pi — downloaded 34 million times. Pathway's Dragon Hatchling solved 97% of 250,000 Sudoku puzzles that leading LLMs solved zero percent of the time.
Why This Has Become Urgent
OpenAI plans to spend $50B on compute in 2026, and data center power consumption is expected to double by 2030. Transformers are hitting a physical scaling ceiling — more capability comes at ever-increasing cost. If even one of these alternative architectures delivers on its promises, the next generation of AI models could be dramatically cheaper — not because they got smarter, but because the math under the hood changed.
Source: www.technologyreview.com
Free course
Stop reading about AI — start building with it
The free Claude Code course: your first site, tool or game — no coding. No upsells, no cross-sells — nothing to buy here.
Start free →▌ Related guides

Author
Evgenii Arsentev
PhD · Chief Executive Officer, digital health
Articles · Latest articles