Google's DiffusionGemma Proves You Don't Need to Train From Scratch to Build a Text Diffusion Model
3 Articles
3 Articles
Google DeepMind presents DiffusionGemma, a model that transforms Gemma 4 into a diffusion generator, achieving 1,500 tokens per second. With reduced training, it offers parallel corrections and surprising performance in structured tasks. *** DiffusionGemma adapts Gemma 4 to diffusion with less than 10% of the original training. It reaches 1,500 tokens per second and corrects parallel errors. Its performance is inferior to the self-egressive, but…
Google's DiffusionGemma proves you don't need to train from scratch to build a text diffusion model
Instead of training a new model from scratch, Google DeepMind retrofitted Gemma 4 into a diffusion model using less than 10 percent of the original training budget. DiffusionGemma generates 256 tokens in parallel instead of one at a time, hitting about 1,500 tokens per second. Quality still trails the original autoregressive model in benchmarks, especially on reasoning tasks. The article Google's DiffusionGemma proves you don't need to train fro…
In the technical report on DiffusionGemma, Google Deepmind shows how the existing Gemma-4 model can be converted into a diffusion model with less than ten percent of the original training budget. Instead of tokens for tokens, it generates 256 tokens in parallel and achieves around 1,500 tokens per second. However, in benchmarks, the quality lags behind the autoregressive initial model, especially in Reasoning tasks. The article Google DiffusionG…
Coverage Details
Bias Distribution
- There is no tracked Bias information for the sources covering this story.
Factuality
To view factuality data please Upgrade to Premium





