{
  "video_id": "reddit_1u3c8q4",
  "channel_slug": "LocalLLaMA",
  "channel_handle": "r/LocalLLaMA",
  "title": "What models you guys running on 8GB? 16GB VRAM? 24GB? 32GB? 48GB?",
  "url": "https://www.reddit.com/r/LocalLLaMA/comments/1u3c8q4/what_models_you_guys_running_on_8gb_16gb_vram/",
  "external_url": null,
  "upload_date": "20260611",
  "published_at": "2026-06-11T21:35:43+00:00",
  "transcript": "And what are you using for kv cache and context? What kind of performance are you getting?  \nWhat is your hardware? And what are you using your models for?\n\n  \nI figure with how fast everything moves, its worth asking once in a while to congeal our experiences.\n\n\n\n--- Top Comments ---\n\n\n[37 upvotes] Don’t sleep on the new gemma 12b model. Kv cache q8, full context 265k, 14 parallel requests, running 2000 prompt processing and 60 token/s on a 5070ti\n\n[28 upvotes] All amd 16 gb vram + 64 gb system ram running Qwen3.5-122B-A10B-APEX-I-Mini. Offloading 40 experts to cpu keeping everything else in gpu. I get \\~20 t/s. \n\n[11 upvotes] 4070 Ti Super (16GB VRAM) + 64 GB RAM. \n\n- `gemma-4-31B-it-qat-UD-Q4_K_XL` for RP and writing. Rather low context so I can fit the entire model in VRAM to get decent speed with a dense model. 6t/s with MTP.\n\n- `gemma-4-26B-A4B-it-ultra-uncensored-heretic-Q6_K` for non-coding agentic work like brainstorming, general task. 128k context. Around 30t/s\n\n- `Qwen_Qwen3.5-35B-A3B-Q6_K_L` for development-related agentic work. 128k context. Around 30t/s as well.\n -  Might switch to a heretic model as well since I work in questionable/controversial areas, not like I'm getting much refusal anyway.\n\nAll running q8 KV Cache.\n\nBackend is llama.cpp and IK llama.\n\nFront end is a mess though:\n\n- SillyTavern and Lumiverse for RP\n\n- Errata for writing\n\n- Open WebUI, Hermes🤢 for generic stuff.\n\n- Copilot VS extension🤢 for coding agentic work.\n\nCan probably do something about the latter two by replacing them all with just Hermes\n\n[7 upvotes] 8GB VRAM and 32GB RAM on a ryzen 5 3600, so pretty low end. Qwen 9B models give me around 20 tok/s which is useable with a 64k context window. I can also run Gemma 4 26B A4B at around the same speed, though I am still fine tuning. I default to Q4 when I download the models.\n\nThe best outputs have come from Qwen3.6 35B A3B models but they are not stable on my setup. Some combination of Intel Arc, Vulkan, llama.cpp and Qwen MoE causes a crash beyond minimal context use.",
  "transcript_chars": 2042,
  "ingested_at": "2026-06-12T01:30:02.789346+00:00",
  "source": "reddit",
  "yt_meta": {
    "score": 66,
    "upvote_ratio": 0.86,
    "num_comments": 102,
    "author": "Inevitable_Mistake32",
    "is_self": true
  }
}