{
  "video_id": "reddit_1tpebhw",
  "channel_slug": "LocalLLaMA",
  "channel_handle": "r/LocalLLaMA",
  "title": "Qwen3.6 huge quality gain from Q4 to Q6 for coding agent",
  "url": "https://www.reddit.com/r/LocalLLaMA/comments/1tpebhw/qwen36_huge_quality_gain_from_q4_to_q6_for_coding/",
  "external_url": null,
  "upload_date": "20260527",
  "published_at": "2026-05-27T18:32:18+00:00",
  "transcript": "So, last week I tried to update my unused local LLM setup. I had to stop using it because quality was too low and deepseek was too cheap.\n\nFirst thing I stopped using Ollama and now I only use llama.cpp built in server that works really great.\n\nThe quality improvement from Q4 to Q6 is outstanding and finally a local LLM server can work very similarly to paid APIs.\n\nThat's great! And MTP makes a big performance gain, on a dual 3090 (downvolted and limited to 65°C) it generates from 20 to 50 tokens per second with minimal heat generation.\n\nSo yes, that time has finally arrived! Local coding agents are a thing and they work 😎\n\n\n\n--- Top Comments ---\n\n\n[30 upvotes] Can you specify which Q4 quant were you using? There are many\n\n[30 upvotes] if you got two 3090 why do you bother with q6 ? run q8 \n\n[15 upvotes] Someone is already going to mentioned this, but here we go: https://github.com/noonghunna/club-3090/blob/master/docs/DUAL_CARD.md\n\nGL OP!\n\n[15 upvotes] Dual 3090 and only Q6? Dude, use vllm and run `Qwen3.6-27B-fp8`. You can get at least 128K context without kv-cache quant. If you think Q6 is good, then prepare to be amazed. Everything else is a toy.",
  "transcript_chars": 1168,
  "ingested_at": "2026-05-28T01:30:06.590670+00:00",
  "source": "reddit",
  "yt_meta": {
    "score": 98,
    "upvote_ratio": 0.93,
    "num_comments": 68,
    "author": "Yes-Scale-9723",
    "is_self": true
  }
}