← All news·2026-08-07·3 min read

AMD Acquires Taalas, Which Bakes AI Model Weights Directly Into Silicon

AMD has acquired Canadian startup Taalas, whose chips embed neural network weights directly into silicon — hitting over 16,000 tokens/sec on Llama 3.1-8B, several times faster than a standard GPU.

aihardwarechips

AMD has acquired Canadian startup Taalas, which bakes neural network weights directly into silicon — allowing the chip to run Llama 3.1-8B at over 16,000 tokens per second per user, several times faster than a conventional GPU. Financial terms were not disclosed; AMD plans to integrate the technology into its Instinct accelerator lineup. Google is pursuing the same approach in parallel with Gemini.

Each chip of this type carries a single hard-baked model — you can't swap it out — but for services with a fixed task (moderation, speech synthesis, embedded assistants) speeds of 16,000 tokens per second per user completely change the economics. This is a signal: the market is moving toward specialized chips purpose-built for a specific AI model.

Source: the-decoder.com

Free course

Stop reading about AI — start building with it

The free Claude Code course: your first site, tool or game — no coding. No upsells, no cross-sells — nothing to buy here.

Start free →
EAEvgenii Arsentev

Author

Evgenii Arsentev

PhD · Chief Executive Officer, digital health