{
  "video_id": "reddit_1usclcz",
  "channel_slug": "LocalLLaMA",
  "channel_handle": "r/LocalLLaMA",
  "title": "Qwen 3.6 Q2-FP8 Terminal Bench 2 and GPQA Scores",
  "url": "https://www.reddit.com/r/LocalLLaMA/comments/1usclcz/qwen_36_q2fp8_terminal_bench_2_and_gpqa_scores/",
  "external_url": null,
  "upload_date": "20260710",
  "published_at": "2026-07-10T03:52:43+00:00",
  "transcript": "TL;DR: Quantization has a marked impact on agentic performance but little effect on knowledge.\n\nI manage a small HPC cluster at a university, and we have recently begun running common benchmarks to help our users understand the effects of quantization. We have just completed the runs on the Qwen 3.6 quantizations and posted the results on our website: [https://scrp.econ.cuhk.edu.hk/llm-benchmark](https://scrp.econ.cuhk.edu.hk/llm-benchmark)\n\nThe results are consistent with what most people would expect: knowledge, as measured by GPQA Diamond, varies very little across quantizations.\n\n[GPQA Chart](https://preview.redd.it/8aqlmibchbch1.png?width=703&format=png&auto=webp&s=10ee17cecb21fed61bf25612a68d8c5c4b5a5d0b)\n\nAgentic use, as measured by Terminal‑Bench 2, shows a significant regression in the lower‑precision quantizations.\n\n[Terminal-Bench 2 Chart](https://preview.redd.it/2u65xswehbch1.png?width=705&format=png&auto=webp&s=f7eafa9f2e34cec345d8184a912b3570d07bda2d)\n\nWe also observed a notable drop compared with Qwen’s official FP8 scores. We believe this stems from the timeout setting—we use Harbor’s default, which ranges from 10 minutes to 1 hour depending on the task, whereas Qwen’s official figures were produced with a flat 3‑hour timeout.\n\nOn the website you’ll also see the range of scores from multiple runs. There is considerable variation across runs; a poor run with a higher‑precision quant can easily be worse than a good run with a lower‑precision quant.\n\nWe are currently benchmarking the GLM‑5.2 quantizations, but, as expected, the process is very slow.\n\n\n\n--- Top Comments ---\n\n\n[13 upvotes] Thanks for sharing these results. Four observations and recommendations:\n\n1. The original model weights are in BFloat16, not F(loat)16. It'd be quite interesting if you could spot a practical difference there, so that old discussion can maybe be put aside. It'd be difficult to accurately do though due to the high variance.\n2. Great that you did 3 to 5 runs per test and not just published flat single-run results. If you have the capacity then going up to 16 runs per test would give us a nicer overview of the score distribution, more narrow error bars.\n3. The drop between F16 and FP8 for Terminal-Bench 2 is surprisingly high. Can you test the original model yourself with your settings? I remember that models were sometimes tested in questionable ways that increased the reported score (like best of 8 vs mean/average of 8). While at it, can you also test the Unsloth Q8\\_XL in comparison to see if it improves things. Usually Q8 is considered to be almost lossless (also when looking at KLD), maybe you can check if it is. If it's not, well, then either something is broken in the test result evaluation, or quantization hits longer context performance stronger than expected.\n4. GPQA Diamond\n\n[9 upvotes] If you have the memory for fp8, run Q8 instead, it’s much better quality\n\n[10 upvotes] I really want to see you do Q8\\_0 and Q6\\_K quant benchmarks here.\n\n[7 upvotes] Really interesting. Thanks for posting! \n\n[6 upvotes] Thanks for posting, this is really useful data, especially with so many people running quantized versions.\n\nThat's really interest to see the drop in even FP8 which is often thought to be close to lossless. Seems like the loss in performance is way more significant in agentic use which is in line with what I've observed.\n  \nDo you have any idea of what are the caused the agents to fail when the runs didn't complete? Would be curious to know where the additional failure rate is coming from",
  "transcript_chars": 3556,
  "ingested_at": "2026-07-10T13:30:17.133301+00:00",
  "source": "reddit",
  "yt_meta": {
    "score": 58,
    "upvote_ratio": 0.97,
    "num_comments": 40,
    "author": "ticoneva",
    "is_self": true
  }
}