{
  "video_id": "reddit_1u3on4u",
  "channel_slug": "LocalLLaMA",
  "channel_handle": "r/LocalLLaMA",
  "title": "EAGLE3 has landed in llama.cpp",
  "url": "https://www.reddit.com/r/LocalLLaMA/comments/1u3on4u/eagle3_has_landed_in_llamacpp/",
  "external_url": "https://github.com/ggml-org/llama.cpp/pull/18039",
  "upload_date": "20260612",
  "published_at": "2026-06-12T07:40:50+00:00",
  "transcript": "After half a year of development, EAGLE3 has been merged into llama.cpp.\n\nEAGLE3 is similar to MTP, but different: the helper model gets extra guidance from the main model instead of guessing completely on its own.\n\n\n\n--- Top Comments ---\n\n\n[31 upvotes] How does it compare to MTP (speed, VRAM usage etc.), and can we use it with Qwen3.6 27B?\n\n[18 upvotes] Great news!\n\nIt's good to have many different ways to break the memory bandwidth limit.\n\n[13 upvotes] Eagle has landed ( ͡° ͜ʖ ͡°)\n\n[9 upvotes] Waiting to see t/s benchmarks.",
  "transcript_chars": 531,
  "ingested_at": "2026-06-12T13:30:22.192610+00:00",
  "source": "reddit",
  "yt_meta": {
    "score": 142,
    "upvote_ratio": 0.99,
    "num_comments": 32,
    "author": "jacek2023",
    "is_self": false
  }
}