#diffusiongemma

1 post

the nvidia logo is displayed on a table

DiffusionGemma Puts 1000+ tok/s on an RTX 5090

Google's open 26B diffusion model hits 700+ tok/s on consumer GPUs with day-zero vLLM support. Here's what the bidirectional architecture changes for local inference.