{
  "video_id": "reddit_1uex6pb",
  "channel_slug": "LocalLLaMA",
  "channel_handle": "r/LocalLLaMA",
  "title": "For users with 4x-8x 6000 PROs, how is your experience with bigger models lately? (GLM 5.2, Kimi 2.7, DeepSeek V4 Pro)",
  "url": "https://www.reddit.com/r/LocalLLaMA/comments/1uex6pb/for_users_with_4x8x_6000_pros_how_is_your/",
  "external_url": null,
  "upload_date": "20260625",
  "published_at": "2026-06-25T02:14:14+00:00",
  "transcript": "Hello guys, hoping you're doing fine!\n\nI was wondering, for users with 4x-8x 6000 PROs (so between 384 and 768GB VRAM), how are bigger models working for you?\n\nI have planned to either jump to 4 or 8 from my actual system, and want to see the experiences with these lately.\n\nIn theory you can run GLM 5.2 at 4 bits, but not 8 bits right? Same with Kimi 2.7, or DeepSeek V4 Pro. There is a ton of info here [https://github.com/local-inference-lab/rtx6kpro/blob/master/benchmarks/results.md](https://github.com/local-inference-lab/rtx6kpro/blob/master/benchmarks/results.md), but missing some of the latest models.\n\nIs there a way too big agentic or programming performance hit by using less than 8 bits? I ask this mostly, because I have read that 4bit perf hit for agentic or programming is way too high vs 8bit, but for bigger models not sure how it really works here.\n\nAre you running these on vLLM/SGLang or another backend?\n\nMany thanks!\n\n\n\n--- Top Comments ---\n\n\n[46 upvotes] I have 1 Max-Q and I use 5.2 Q8\n\n1 t/s 🤡\n\nI have to be very deliberate with my prompts\n\n[28 upvotes] I'm using Qwen 3.5 397b across 4x rtx6kpro\n\nWe use it for everything, it's a really good model for coding, agentic tools, and general tasks and allows me to have a ton of concurrent users at 140 t/s + or -. It's silly fast to be honest. Wish we got a 3.6 or is glm 5.2 was smaller, I don't want to lose our concurrent capacity \n\n[20 upvotes] Kimi is natively 4-bit\n\n[14 upvotes] Quantization doesnt affect big models as negatively.  A lot of information is stored in the node topology vs the weights themselves.  Small models are less topologically dense so the quantization removes a greater chinck ofntheir intelligence.",
  "transcript_chars": 1704,
  "ingested_at": "2026-06-25T13:30:06.624325+00:00",
  "source": "reddit",
  "yt_meta": {
    "score": 69,
    "upvote_ratio": 0.92,
    "num_comments": 74,
    "author": "panchovix",
    "is_self": true
  }
}