← All news·2026-08-09·3 min read

DiffusionGemma: Google Converts Gemma 4 Into a Diffusion Model for Under 10% of Training Cost

Google DeepMind converted Gemma 4 into a diffusion model using less than 10% of the original training budget. The result: 1,500 tokens/sec and parallel generation in 256-token blocks.

airesearchgooglediffusionllm

Google DeepMind took the existing 26-billion-parameter Gemma 4 and converted it into a diffusion model for less than 10% of the original training budget. DiffusionGemma generates 256 tokens simultaneously — not one at a time like a standard LLM — and delivers around 1,500 tokens per second on an Nvidia H100. Quality is slightly below the original, but reasoning tasks improved by 10 points, and Sudoku puzzles are solved 85% of the time — versus zero for the base model.

Until now, diffusion-based text models required training from scratch — a barrier that kept all but the largest players away. Now that any existing model can be repurposed for a fraction of the budget, specialized variants will become cheaper and more diverse. This reshapes the economics of the race for generation efficiency.

Source: the-decoder.com

Free course

Stop reading about AI — start building with it

The free Claude Code course: your first site, tool or game — no coding. No upsells, no cross-sells — nothing to buy here.

Start free →
EAEvgenii Arsentev

Author

Evgenii Arsentev

PhD · Chief Executive Officer, digital health