{
  "video_id": "reddit_1vzp4c9",
  "channel_slug": "LocalLLaMA",
  "channel_handle": "r/LocalLLaMA",
  "title": "llama : add --n-cpu-ffn option by John-194 · Pull Request #26622 · ggml-org/llama.cpp",
  "url": "https://www.reddit.com/r/LocalLLaMA/comments/1vzp4c9/llama_add_ncpuffn_option_by_john194_pull_request/",
  "external_url": "https://github.com/ggml-org/llama.cpp/pull/26622",
  "upload_date": "20260827",
  "published_at": "2026-08-27T09:29:08+00:00",
  "transcript": "**tl;dr faster dense models for low VRAM people**\n\noption similar to the existing `--n-cpu-moe`\n\nIt puts user specified amount of FFN sublayers for dense models.\n\nPR by [u/Stainless-Bacon](https://www.reddit.com/user/Stainless-Bacon/)\n\n\n\n--- Top Comments ---\n\n\n[6 upvotes] Nice to see this merge! [u/Stainless-Bacon](https://www.reddit.com/user/Stainless-Bacon/) 👍 \n\nGlad I posted [this thread](https://www.reddit.com/r/LocalLLaMA/s/X9hemfT09k) which brought more eyes(look at the reactions) on this PR.\n\n[5 upvotes] FFN sublayers go to CPU, attention stays on the GPU, and it's one flag where the old route was regex matching in the -ot style. For a model I plan to run for months I'd take this over --fit.\n\n[5 upvotes] In which situation exactly, would I want to use this? \n\n[5 upvotes] The question is: Does using this manually (and spending time to tweak it) offer a benefit over simply using the default `-fit` with a decently sized `--fit-target`?",
  "transcript_chars": 953,
  "ingested_at": "2026-08-27T13:30:03.035338+00:00",
  "source": "reddit",
  "yt_meta": {
    "score": 53,
    "upvote_ratio": 0.97,
    "num_comments": 17,
    "author": "jacek2023",
    "is_self": false
  }
}