{
  "video_id": "reddit_1wcvmuo",
  "channel_slug": "LocalLLaMA",
  "channel_handle": "r/LocalLLaMA",
  "title": "I find it funny that a flash model is now 512GB",
  "url": "https://www.reddit.com/r/LocalLLaMA/comments/1wcvmuo/i_find_it_funny_that_a_flash_model_is_now_512gb/",
  "external_url": null,
  "upload_date": "20260910",
  "published_at": "2026-09-10T20:57:40+00:00",
  "transcript": "A few years ago a 100GB was considered a very large language model.  What do we call under 100GB models now?  Tiny models?  haha\n\n\n\n--- Top Comments ---\n\n\n[249 upvotes] Flash means it's fast, not small. \n\n[64 upvotes] Flash implies fast not small? Otherwise they would name it \"nano\" or something.\n\n[22 upvotes] True ! but a few years ago we also had 400B **dense** models, the 500B model u talking about has 8/16b activated so it's pricing per 1M is still in the range of \"flash price\"\n\n[18 upvotes] The use of PLE/n-gram plus MoE has opened up new categories of LLM. I expect soon we will have smaller models with huge downloads and ridiculously comprehensive world knowledge that can run on a laptop as long as it has decent storage capacity and speed. \n\nWe kind of already do in Qwen3.8-Flash-Next. But I expect then memory requirement to fall while we all learn how to effectively leverage n-grams on disk. ",
  "transcript_chars": 912,
  "ingested_at": "2026-09-11T01:30:04.184328+00:00",
  "source": "reddit",
  "yt_meta": {
    "score": 218,
    "upvote_ratio": 0.94,
    "num_comments": 66,
    "author": "Terminator857",
    "is_self": true
  }
}