{
  "video_id": "reddit_1tqupcr",
  "channel_slug": "LocalLLaMA",
  "channel_handle": "r/LocalLLaMA",
  "title": "llama: use f16 mask for FA to save VRAM by am17an · Pull Request #23764 · ggml-org/llama.cpp",
  "url": "https://www.reddit.com/r/LocalLLaMA/comments/1tqupcr/llama_use_f16_mask_for_fa_to_save_vram_by_am17an/",
  "external_url": "https://github.com/ggml-org/llama.cpp/pull/23764",
  "upload_date": "20260529",
  "published_at": "2026-05-29T07:49:27+00:00",
  "transcript": "now you can download more VRAM ;) (by downloading new llama.cpp version)\n\n\n\n--- Top Comments ---\n\n\n[63 upvotes] This guy is on fire lately. llama.cpp contributor of the year.\n\n[27 upvotes] and he just landed ANOTHER 1.2gb save follow-up [https://github.com/ggml-org/llama.cpp/pull/23861](https://github.com/ggml-org/llama.cpp/pull/23861) \n\n[16 upvotes] According to the merge we can save 1.2GB of vram by default now ?",
  "transcript_chars": 418,
  "ingested_at": "2026-05-29T13:30:24.568806+00:00",
  "source": "reddit",
  "yt_meta": {
    "score": 118,
    "upvote_ratio": 0.96,
    "num_comments": 38,
    "author": "jacek2023",
    "is_self": false
  }
}