{
  "video_id": "reddit_1uar4e2",
  "channel_slug": "LocalLLaMA",
  "channel_handle": "r/LocalLLaMA",
  "title": "GLM 5.2: 98% of max level intelligence with less than half of tokens usage",
  "url": "https://www.reddit.com/r/LocalLLaMA/comments/1uar4e2/glm_52_98_of_max_level_intelligence_with_less/",
  "external_url": null,
  "upload_date": "20260620",
  "published_at": "2026-06-20T08:19:42+00:00",
  "transcript": "According to [this](https://artificialanalysis.ai/?intelligence-efficiency=output-tokens-per-task) number of reasoning tokens from GLM 5.1 to GLM 5.2 more than doubled from 16.7k to 36.7k and for me as a local user with old junk Xeon setup this makes GLM 5.2 unusable to the extent where I had to shut down model after 12h of waiting it to respond to my math problem question.\n\nBut then I saw this graph from z\\_ai [technical report](https://z.ai/blog/glm-5.2), which basically implies that you can use less than half of the tokens of max effort on high level and still get around 98% of max level intelligence at least in coding tasks. So I encourage both local and API users to try high level, because by default GLM 5.2 is set to max level.\n\nUpd: Finally after 6k tokens on the high level with Q4 quant I got an answer to my math question. It is Ok, but it is only half right. As a comparison in [z.ai](http://z.ai) chat on max level answer was ~~much~~ a bit better. I don't know may be Q4 + high level is already to much. See Upd2.\n\nUpd2: I also run in [z.ai](http://z.ai) chat the same prompt with \"high\" effort level and now reconsidering all 3 answers I would say that they are very similar. The only difference is that on \"max\" level it explicitly talked about second case, but then dismissed it, although it shouldn't. In other two responses it dismissed it from the beginning. So the difference is more down to presentation of the same partially correct result and not result itself.\n\nTake these results with gran of salt as it is just 1 shot per running conditions, but it looks like \"high\" level is better alternative for day to day use and \"max\" if you absolutely need perfect result or you want your model to look good on benchmarks))\n\nhttps://preview.redd.it/eha9j6vd9e8h1.png?width=6166&format=png&auto=webp&s=204c3261fada0c3eac8e4ab52fed7b45c1831b7b\n\n\n\n--- Top Comments ---\n\n\n[35 upvotes] The thinking has never been a problem with llama.cpp.  Use reasoning\\_budget to limit how long a model thinks for.  It's better than reasoning\\_effort.   With reasoning\\_effort, you still can't control how much thinking.   I have played with this from 0 to 32k in 4k increments and I find that 0, 4k and 8k are often great enough for 99% of things.\n\n[8 upvotes] > Finally after 6k tokens on the high level with Q4 quant I got an answer to my math question. It is Ok, but it is only half right. As a comparison in z.ai chat on max level answer was much better.\n\nI had similar experience with other big moe models. People were cheering for minimax 2.7, mimo 2.5, stepfun 3.7 etc and I ran them  at q4 to q6 and none of them produced better results than qwen 3.6 27b (q8 or bf16) for some reason. (they either produced half assed responses, wrong results, or mediocre outputs etc)  I haven't tried them on cloud though at full precision so I can't say anything about that. I probably should have.\n\nMy point is, it seems to me they lose a lot when they get quantized. Not sure if the tps affects anything.",
  "transcript_chars": 3008,
  "ingested_at": "2026-06-20T13:30:10.100892+00:00",
  "source": "reddit",
  "yt_meta": {
    "score": 159,
    "upvote_ratio": 0.94,
    "num_comments": 36,
    "author": "perelmanych",
    "is_self": true
  }
}