{
  "video_id": "reddit_1u27fga",
  "channel_slug": "singularity",
  "channel_handle": "r/singularity",
  "title": "Google releases DiffusionGemma, new experimental open model with up to 4x faster output on dedicated GPUs",
  "url": "https://www.reddit.com/r/singularity/comments/1u27fga/google_releases_diffusiongemma_new_experimental/",
  "external_url": "https://blog.google/innovation-and-ai/technology/developers-tools/diffusion-gemma-faster-text-generation/",
  "upload_date": "20260610",
  "published_at": "2026-06-10T16:38:12+00:00",
  "transcript": "• **DiffusionGemma** is Google's new experimental open model built on the Gemma 4 architecture.\n\n• Unlike traditional LLMs that generate text token-by-token, it generates and refines blocks of text in parallel using a diffusion-based approach.\n\n• Google says it can deliver up to **4× faster** inference on dedicated GPUs.\n\n• The model activates \\~3.8B parameters per step from a 26B-parameter Gemma 4 MoE architecture.\n\n• Released under the Apache 2.0 license with **support** for local deployment and integration with tools such as Hugging Face Transformers and vLLM.\n\n**Source: Google Deepmind**\n\n\n\n--- Top Comments ---\n\n\n[66 upvotes] really glad diffusion models are still being worked on\n\n[47 upvotes] https://preview.redd.it/c5mzbnedhh6h1.jpeg?width=1000&format=pjpg&auto=webp&s=72684fb7437f2416209770c4b1d215a91deccd00\n\n[17 upvotes] I wonder if diffusion models can be better parallelized",
  "transcript_chars": 895,
  "ingested_at": "2026-06-11T01:30:39.132745+00:00",
  "source": "reddit",
  "yt_meta": {
    "score": 256,
    "upvote_ratio": 0.99,
    "num_comments": 43,
    "author": "BuildwithVignesh",
    "is_self": false
  }
}