{
  "video_id": "reddit_1u19k2h",
  "channel_slug": "LocalLLaMA",
  "channel_handle": "r/LocalLLaMA",
  "title": "Unsloth Gemma 4 QAT MTP assistant models now available",
  "url": "https://www.reddit.com/r/LocalLLaMA/comments/1u19k2h/unsloth_gemma_4_qat_mtp_assistant_models_now/",
  "external_url": null,
  "upload_date": "20260609",
  "published_at": "2026-06-09T16:12:02+00:00",
  "transcript": "They're both available as q8_0 models named `mtp-gemma-4-*.gguf` on the root of the directory and in both q8_0 and larger quants within an `MTP` folder.\n\n- https://huggingface.co/unsloth/gemma-4-12B-it-qat-GGUF/tree/main\n- https://huggingface.co/unsloth/gemma-4-26B-A4B-it-qat-GGUF/tree/main\n- https://huggingface.co/unsloth/gemma-4-31B-it-qat-GGUF/tree/main\n- https://huggingface.co/unsloth/gemma-4-E2B-it-qat-GGUF/tree/main\n- https://huggingface.co/unsloth/gemma-4-E2B-it-qat-mobile-GGUF/tree/main\n- https://huggingface.co/unsloth/gemma-4-E4B-it-qat-GGUF/tree/main\n- https://huggingface.co/unsloth/gemma-4-E4B-it-qat-mobile-GGUF/tree/main\n\n\n\n--- Top Comments ---\n\n\n[22 upvotes] After testing both google's quant and unsloth's for an extended time, unsloth's release is rock solid this time. Good job and well done!\n\nThey also got BF16 / F16 versions of the drafters in case there is any benefit to it.\n\n\\[edit\\] also worth noting is the nice documentation they added:  \n[https://huggingface.co/unsloth/gemma-4-31B-it-qat-GGUF/blob/main/MTP/README.md](https://huggingface.co/unsloth/gemma-4-31B-it-qat-GGUF/blob/main/MTP/README.md)\n\n[7 upvotes] I had been waiting for these, but my initial impression of the 31B assistant model is not great unfortunately. It comes in at 515MB for the smallest (Q8_0) quant yet produces similar (lower, if anything) acceptance rates on average compared to the 280MB Q4_0 assistant I'd been using: https://huggingface.co/RachidAR/gemma-4-31B-it-qat-Q4_0-Q4emb-MTP-assistant-gguf\n\nUnsloth assistant:\n\n    > python mtp-bench.py\n      code_python        pred= 192 draft= 163 acc= 136 rate=0.834 tok/s=152.9\n      code_cpp           pred= 192 draft= 189 acc= 127 rate=0.672 tok/s=132.2\n      explain_concept    pred= 192 draft= 194 acc= 126 rate=0.649 tok/s=130.3\n      summarize          pred= 192 draft= 165 acc= 135 rate=0.818 tok/s=149.8\n      qa_factual         pred= 192 draft= 193 acc= 125 rate=0.648 tok/s=131.2\n      translation        pred= 192 draft= 180 acc= 131 rate=0.728 tok/s=143.6\n      creative_short     pred= 192 draft= 231 acc= 113 rate=0.489 tok/s=111.1\n      stepwise_math      pred= 192 draft= 162 acc= 136 rate=0.840 tok/s=153.7\n      long_code_review   pred= 192 draft= 193 acc= 126 rate=0.653 tok/s=115.3\n    \n    Aggregate: {\n      \"n_requests\": 9,\n      \"total_predicted\": 1728,\n      \"to\n\n[7 upvotes] Why are these Q8_0 and not Q4_0 if the assistants are also from the Q4_0 training?\n\n[5 upvotes] \\+10% tps in exchange for 1 GB of RAM. Eh..\n\n[5 upvotes] The Q8\\_0 assistant seems worse than Q4\\_0 (made with llama.cpp's quantize) for me. Maybe it's better in longer contexts? Doesn't seem worth the size for me though. I thought it would be better due to the discrepancy between gemma's Q4\\_0 and llama.cpp's Q4\\_0. It's either the same or worse at both spec-draft-n-max=2 and 4.\n\n    model: gemma-4-31B-it-qat-Q4_K_XL-no_mtp\n      code_python        pred= 192 draft=   0 acc=   0 rate=n/a tok/s=39.8\n      code_cpp           pred= 192 draft=   0 acc=   0 rate=n/a tok/s=39.8\n      explain_concept    pred= 192 draft=   0 acc=   0 rate=n/a tok/s=39.8\n      summarize          pred= 192 draft=   0 acc=   0 rate=n/a tok/s=39.6\n      qa_factual         pred= 192 draft=   0 acc=   0 rate=n/a tok/s=39.7\n      translation        pred= 192 draft=   0 acc=   0 rate=n/a tok/s=39.7\n      creative_short     pred= 192 draft=   0 acc=   0 rate=n/a tok/s=39.7\n      stepwise_math      pred= 192 draft=   0 acc=   0 rate=n/a tok/s=39.6\n      long_code_review   pred= 192 draft=   0 acc=   0 rate=n/a tok/s=38.3\n    \n    Aggregate: {\n      \"n_requests\": 9,\n      \"total_predicted\": 1728,\n      \"total_draft\": 0,\n      \"total_draft_accepted\": 0,\n      \"aggregate_acc",
  "transcript_chars": 3713,
  "ingested_at": "2026-06-10T01:30:04.910571+00:00",
  "source": "reddit",
  "yt_meta": {
    "score": 130,
    "upvote_ratio": 0.98,
    "num_comments": 45,
    "author": "ParadigmComplex",
    "is_self": true
  }
}