{
  "video_id": "reddit_1uu8d1v",
  "channel_slug": "LocalLLaMA",
  "channel_handle": "r/LocalLLaMA",
  "title": "Xiaomi quietly uploaded MiMo-V2.5-DFlash — official DFlash weights are now on Hugging Face",
  "url": "https://www.reddit.com/r/LocalLLaMA/comments/1uu8d1v/xiaomi_quietly_uploaded_mimov25dflash_official/",
  "external_url": null,
  "upload_date": "20260712",
  "published_at": "2026-07-12T07:11:43+00:00",
  "transcript": "[https://huggingface.co/XiaomiMiMo/MiMo-V2.5-DFlash](https://huggingface.co/XiaomiMiMo/MiMo-V2.5-DFlash)\n\nXiaomi appears to have quietly uploaded **MiMo-V2.5-DFlash** to Hugging Face: there is dedicated `dflash` directory containing the Dflash model, anyone willing to GGUF it and try? I'd do it but I can't today.  \nThis model is pretty good IMO (300B + params) and runs at about 8-10 tk/s on 2x24gb cards + vram offload (96/128gb drr5), dflash could double that speed and make it very interesting.\n\nEDIT: the main reason it's interesting, is because the MTP head was shared already, but doesn't work yet il llama cpp. I speculate (pun intended) the Dflash does work instead.\n\nEDIT2: very cool! they shared also the SEPARATE MTP model. the reason Llama doesn't work already is because it has trouble identifying the MTP layers. a separate MTP model might work too.\n\n\n\n--- Top Comments ---\n\n\n[1 upvotes] Your post is getting popular and we just featured it on our Discord! [Come check it out!](https://discord.gg/PgFhZ8cnWW)\n\nYou've also been given a special flair for your contribution. We appreciate your post!\n\n*I am a bot and this action was performed automatically.*\n\n[32 upvotes] Mimo 2.5 full version is incredible and underrated. Haven’t had the chance to use flash much\n\n[17 upvotes] On swe-rebench it sits just in between DeepSeek V4 flash and Pro in both price and performance despite being basically flash size 284B, and not pro size, 1.6T!\n\n[5 upvotes] Curious what tk/s bump people actually see once DFlash is wired into llama.cpp for real, speculative decoding gains on paper don't always survive contact with VRAM offload once you're spilling to system RAM like your setup. Hope someone GGUFs it soon, this one's worth testing.",
  "transcript_chars": 1743,
  "ingested_at": "2026-07-12T13:30:05.534510+00:00",
  "source": "reddit",
  "yt_meta": {
    "score": 163,
    "upvote_ratio": 0.97,
    "num_comments": 19,
    "author": "nasone32",
    "is_self": true
  }
}