{
  "video_id": "reddit_1u7a2hn",
  "channel_slug": "LocalLLaMA",
  "channel_handle": "r/LocalLLaMA",
  "title": "Be wary of Qwen/Claude distillations - they're often worse than the base model",
  "url": "https://www.reddit.com/r/LocalLLaMA/comments/1u7a2hn/be_wary_of_qwenclaude_distillations_theyre_often/",
  "external_url": null,
  "upload_date": "20260616",
  "published_at": "2026-06-16T10:48:22+00:00",
  "transcript": "Just to be clear; I am not attempting to call anybody out or be mean to those who take the time/money to make these models, I just want to inform people about these distills/finetunes since there's clearly some confusion going on.\n\nI'm going to assume those of us who often visit this subreddit have noticed these models, particularly the \"Qwopus\" model and the such, though I'm sure there's probably Gemma 4/Claude distills too. As I type this, there's currently a Qwen 3.6 based Claude Fable 5 distillation model on the frontpage. Seems pretty cool, right?\n\nYep. Up until you actually look into how these models were distilled. This new Fable distillation uses around 4,000 samples of Fable 5/Opus 4.8 to finetune Qwen 3.6 on. 4k samples is basically *nothing* when it comes to improving a models quality/performance. At best, it'll act slightly differently. But it certainly won't perform better than just running standard Qwen 3.6. If anything, it's actually likely to slightly degrade quality.\n\nWhy? 4K samples is just not enough. And I am aware that Qwopus (or it may be another finetune called Qwen3.6-Claude-Opus.4.6-Distill iirc) has a version with ~8-10k samples used for the training rather than the 3-4K. Unfortunately that's still nowhere near enough to be actually meaningful.\n\nIf anybody remembers the original DeepSeek-R1 LLaMa/Qwen distillations that were released by deepseek offiically back when the model first came out, around ~700,000 samples from R1 was used to create those distills. That's enough to not only impact behaviour, but actually improve benchmark scores.\n\nSo, these Qwen + Claude models will have a slightly different reasoning style. They might feel \"more Opus-like\" chatting wise. But they are not performing better than their base Qwen models, and based on everything I've seen, a lot of people seem to think that's the case. Even with that Qwen/Opus distill that uses like 10K+ samples, that's still just not enough to transfer any sort of actual capability. [There's a decent example of someone testing this, showing Qwopus hallucinating compared to the standard Qwen 3.6, and also taking twice the amount of time.](https://akitaonrails.com/en/2026/04/24/llm-benchmarks-parte-3-deepseek-kimi-mimo/#the-discovery-claude-distillation-doesnt-transfer-library-knowledge) - there's also ofc plenty of people on this sub who have posted similar results.\n\nSo yeah, just something to be aware of whenever you come across these distills/finetunes. At the very least, don't blindly trust them to be superior and bench them on your own specific usecases. I've personally tried a couple of these finetunes and both of them had issues with coherence and subtle mistakes that the standard model didn't have. But YMMV.\n\n\n\n--- Top Comments ---\n\n\n[49 upvotes] Not \"often\", ALWAYS\n\n  \nI discovered a simple ruleset, ordered from highest confidence to lowest:\n\n* If the model card is LLM-written, it sucks\n* If the model card has shitty evals with low N or only pass@5, it sucks\n* If the model card evals are only about web development, it sucks\n* If the poster is in AI psychosis, religious, a furry or a chud, it sucks ARSE\n* If the poster has math background, it maybe doesn't suck\n* If the poster gloats about how much better anthropic models are, it sucks\n* If the model card doesn't mention it's a distill, but it's a distill, it fucking sucks\n\n[41 upvotes] Be very careful if you see the Qwenis model.\n\n[29 upvotes] Correct. We are way past the easy gains era where a couple thousand examples would improve a model. At this point improvement focused fine-tuning needs to be carefully done with over 100k examples, and recovered with GRPO.\n\n[14 upvotes] > 4k samples is basically nothing when it comes to improving a models quality/performance\n\nEspecially keeping in mind we don't get logits besides small top-N at best (or even only top-1?).\n\nWhich seems to bear shitload of information.\n\nAnd, if I recall correctly, antropic does not provide full chain of thoughts. Just a summary. Which lose another shitload of information.\n\nAnd without it is basically just a 4K samples another-model-partial-responsr-supervised finetune, not a proper knowledge distillation.",
  "transcript_chars": 4191,
  "ingested_at": "2026-06-16T13:30:28.740969+00:00",
  "source": "reddit",
  "yt_meta": {
    "score": 140,
    "upvote_ratio": 0.97,
    "num_comments": 45,
    "author": "ayylmaonade",
    "is_self": true
  }
}