{
  "video_id": "reddit_1tm3toi",
  "channel_slug": "LocalLLaMA",
  "channel_handle": "r/LocalLLaMA",
  "title": "Qwen3.6-35B-A3B-Uncensored-Genesis-APEX-MTP",
  "url": "https://www.reddit.com/r/LocalLLaMA/comments/1tm3toi/qwen3635ba3buncensoredgenesisapexmtp/",
  "external_url": null,
  "upload_date": "20260524",
  "published_at": "2026-05-24T06:08:22+00:00",
  "transcript": "Here model: [https://huggingface.co/LuffyTheFox/Qwen3.6-35B-A3B-Uncensored-Genesis-V2-APEX-MTP-GGUF](https://huggingface.co/LuffyTheFox/Qwen3.6-35B-A3B-Uncensored-Genesis-V2-APEX-MTP-GGUF)\n\nSafetensors: [https://huggingface.co/LuffyTheFox/Qwen3.6-35B-A3B-Uncensored-Genesis-V2-FP8-Safetensors](https://huggingface.co/LuffyTheFox/Qwen3.6-35B-A3B-Uncensored-Genesis-V2-FP8-Safetensors)\n\n*Testing results in Open Code on hardware (Beelink gtr9 pro + Strix Halo) done by my friend on Q8\\_K\\_P - MTP quant:*\n\n1. 5 sessions with 200k context, not a single glitch, no loops, no repeated tool calls.\n2. After 120k tokens he suddenly gave another task that doesn't intersect with what it was doing at all, and it calmly picked up and solved it correctly.\n3. Uncensored with MTP support with APEX and APEX Compact quantization.\n4. Safetensors support for Apple MLX conversion for Mac users. MTP-Safetensors now in development.\n\n**Recommended quant:** APEX, MTP-APEX\n\n**Recommended settings for LM Studio:**\n\n[System Prompt](https://huggingface.co/LuffyTheFox/Qwen3.6-35B-A3B-Uncensored-Genesis-V2-APEX-MTP-GGUF/raw/main/System_Prompt.txt)\n\n[Chat Template](https://huggingface.co/LuffyTheFox/Qwen3.6-35B-A3B-Uncensored-Genesis-V2-APEX-MTP-GGUF/raw/main/chat_template.jinja)\n\n[Chat Template Thinking](https://huggingface.co/LuffyTheFox/Qwen3.6-35B-A3B-Uncensored-Genesis-V2-APEX-MTP-GGUF/raw/main/chat_template_thinking.jinja)\n\nOr use this minimal string as the **first line**:\n\n>`You are Qwen, created by Alibaba Cloud. You are a helpful assistant.`\n\nThen add anything you want after. **Model may underperform without this first line.**\n\nSettings:\n\n|Parameter|Value|\n|:-|:-|\n|Temperature|0.7|\n|Top K Sampling|20|\n|Presence Penalty|1.5|\n|Repeat Penalty|1.0|\n|Top P Sampling|0.8|\n|Min P Sampling|0|\n|Seed|42|\n\nEnjoy 😄\n\n\n\n--- Top Comments ---\n\n\n[48 upvotes] LocalLLaMA users casually running 35B models with 200k context on mini PCs while big tech still says “requires 8 H100s” 💀\n\n[13 upvotes] I've really never managed to get anything good out of APEX Quants when using them with all coding agents. They just go off on the wrong tangent and / or make wrong tool calls, or start looping heavily. And I've always gone for the QUALITY presets, which should be the one with the best results.\n\n[6 upvotes] anyone tested this for tool calling/structured output? the uncensored models sometimes break json formatting in my experience",
  "transcript_chars": 2413,
  "ingested_at": "2026-05-24T13:30:28.782573+00:00",
  "source": "reddit",
  "yt_meta": {
    "score": 120,
    "upvote_ratio": 0.89,
    "num_comments": 54,
    "author": "EvilEnginer",
    "is_self": true
  }
}