{
  "video_id": "reddit_1tzbcyp",
  "channel_slug": "LocalLLaMA",
  "channel_handle": "r/LocalLLaMA",
  "title": "llama.cpp Gemma4 MTP support merged!",
  "url": "https://www.reddit.com/r/LocalLLaMA/comments/1tzbcyp/llamacpp_gemma4_mtp_support_merged/",
  "external_url": "https://github.com/ggml-org/llama.cpp/pull/23398",
  "upload_date": "20260607",
  "published_at": "2026-06-07T12:53:17+00:00",
  "transcript": "\n\n--- Top Comments ---\n\n\n[16 upvotes] QAT + MTP  let’s go!\n\n[6 upvotes] Been watching this one closely! Compared to Qwen, the Gemma4 family seems to underperform on benchmarks, so it doesn't receive as much fanfare, but I've been giving 31B a go recently and it's been really nice, regardless of what the charts say. It's a really well-rounded model.\n\nWould encourage you guys to give them another go once they get a build out for this.\n\n[4 upvotes] Been writing/editing Ideogram prompts (json format) quite a bit with Gemma 31B QAT the past couple of days.\n\nOn a 5090, generation speed goes from \\~70t/s to 100-120t/s with MTP and --spec-default speeds things up nicely if it's just a small edit and lots of copying.",
  "transcript_chars": 717,
  "ingested_at": "2026-06-07T13:30:04.772945+00:00",
  "source": "reddit",
  "yt_meta": {
    "score": 68,
    "upvote_ratio": 0.99,
    "num_comments": 17,
    "author": "pinkyellowneon",
    "is_self": false
  }
}