{
  "video_id": "reddit_1vwhj0l",
  "channel_slug": "LocalLLaMA",
  "channel_handle": "r/LocalLLaMA",
  "title": "you can now use MTP in GLM-Air",
  "url": "https://www.reddit.com/r/LocalLLaMA/comments/1vwhj0l/you_can_now_use_mtp_in_glmair/",
  "external_url": "https://github.com/ggml-org/llama.cpp/pull/26534",
  "upload_date": "20260823",
  "published_at": "2026-08-23T20:08:04+00:00",
  "transcript": "If anyone still remembers GLM-4.5-Air from last year, you can now get a nice speedup by enabling MTP in llama.cpp.\n\nIt is a 106B MoE with only 12B active parameters, which makes it interesting for machines with lots of memory but limited compute, such as Strix Halo or DGX Spark. I use it on 3090s. It's still great for creative writing, especially since we never got Gemma 4 124B MoE.\n\nThere are multiple creative-writing / RP finetunes available on Hugging Face: [https://huggingface.co/models?other=base\\_model:finetune:zai-org%2FGLM-4.5-Air&sort=likes](https://huggingface.co/models?other=base_model:finetune:zai-org%2FGLM-4.5-Air&sort=likes) (some even from this year). I also recommend Intellect 3.x by PrimeIntellect\n\nIf your GGUF does not include the MTP block, you can download a small file from here: [https://huggingface.co/jacek2024/GLM-4.5-Air-MTP-GGUF](https://huggingface.co/jacek2024/GLM-4.5-Air-MTP-GGUF)\n\nThanks a lot to [**devMiikaK**](https://github.com/devMiikaK) and [**HeadCutter**](https://github.com/HeadCutter) for testing the PR while it was in progress.\n\nPS. It also works for the full GLM-4.5, but I doubt anyone still uses it ;)\n\n\n\n--- Top Comments ---\n\n\n[19 upvotes] Nice to see added MTP for it. It doesn't matter if the model is new or old. What matters is better support\n\n[10 upvotes] Good work!\n\nI remember GLM 4.7 Flash has an MTP head too. Still no plan to support it?\n\n[5 upvotes] >It is a 106B MoE with only 12B active parameters, which makes it interesting for machines with lots of memory but limited compute\n\nI'd add that in my experience at least, GLM Air is one of the very, VERY, few small (ish) MoE that doesn't have that brittle small MoE feels to it. The same feel as a small model hooked up to good RAG. To date it and to a lesser extent Gemma 4 26b are the only two that haven't felt like that to me. Models that just feel like \"a good model\" rather than \"a good model, given that it only has x active parameters\".\n\nSo MTP for it is fantastic news, thanks!\n\n[3 upvotes] What a magic sub and what a great llm project! I have just been using glm air for a few days. Now I get this great news.",
  "transcript_chars": 2140,
  "ingested_at": "2026-08-24T01:30:03.399243+00:00",
  "source": "reddit",
  "yt_meta": {
    "score": 63,
    "upvote_ratio": 0.98,
    "num_comments": 16,
    "author": "jacek2023",
    "is_self": false
  }
}