{
  "video_id": "reddit_1w03zdo",
  "channel_slug": "LocalLLaMA",
  "channel_handle": "r/LocalLLaMA",
  "title": "llama.cpp support for Qwen3.8-Flash-Next has been merged",
  "url": "https://www.reddit.com/r/LocalLLaMA/comments/1w03zdo/llamacpp_support_for_qwen38flashnext_has_been/",
  "external_url": "https://github.com/ggml-org/llama.cpp/pull/27742",
  "upload_date": "20260827",
  "published_at": "2026-08-27T19:34:24+00:00",
  "transcript": "finally I can download the GGUF\n\nUPDATE Q4 GGUF downloaded, I have 55 t/s on 4x3090, video in the comment\n\n\n\n--- Top Comments ---\n\n\n[28 upvotes] afaik mtp and ngram offloading do not work right?\n\n[19 upvotes] MTP when?\n\n[18 upvotes] Managed to get 10tks on my 4gb card with offloading to the ssd. I think I can get this working even better tonight.\n\n[6 upvotes] FYI I tried the unsloth fork that was released today with fit on and it crashed with CUDA out of memory error at 102k/262,144 context with 2048 ubatch and batch size. The default settings (512 batch and ubatch size) work fine.",
  "transcript_chars": 588,
  "ingested_at": "2026-08-28T01:30:04.095479+00:00",
  "source": "reddit",
  "yt_meta": {
    "score": 230,
    "upvote_ratio": 0.97,
    "num_comments": 56,
    "author": "jacek2023",
    "is_self": false
  }
}