{
  "video_id": "reddit_1vs2tz1",
  "channel_slug": "LocalLLaMA",
  "channel_handle": "r/LocalLLaMA",
  "title": "DFlash 2: Keep Drafting Parallel",
  "url": "https://www.reddit.com/r/LocalLLaMA/comments/1vs2tz1/dflash_2_keep_drafting_parallel/",
  "external_url": "https://inco.ai/blog/dflash2/",
  "upload_date": "20260818",
  "published_at": "2026-08-18T21:37:01+00:00",
  "transcript": "\n\n--- Top Comments ---\n\n\n[22 upvotes] Wait stop HYPE HYPE HYPE\n\nTHIS IS BETTER THAN DSPARK!?\n\nEDIT: \n~10% longer acceptance length -> ~35%!! more decode tok/s over DSpark. Brilliant work\n\nEDIT2:\nas always, we need to check acceptance rate falloff on larger contexts. first DFLash drafters fell off at 30k, the best qwen3.8 DSpark I saw fell off after 60k-90k so much so that MTP was faster\n\nSo I'll test it tomorrow - maybe.\n\n[11 upvotes] Why does nobody ever talk about the prefill penalty from using DFlash? On my setup (dual R9700s, with llama.cpp) it's even worse than MTP. Prefill goes from 1800t/s on Qwen 3.6 27B to \\~700-800t/s, which is basically unusable.\n\nAm I holding it wrong, or is everybody just glossing over this?\n\n[4 upvotes] Trying it now on my 5090... finally we have a good dflash implementation for qwen3.8 on llama.cpp it seems",
  "transcript_chars": 850,
  "ingested_at": "2026-08-19T01:30:02.703543+00:00",
  "source": "reddit",
  "yt_meta": {
    "score": 56,
    "upvote_ratio": 0.97,
    "num_comments": 24,
    "author": "coder543",
    "is_self": false
  }
}