{
  "video_id": "reddit_1w92x3j",
  "channel_slug": "LocalLLaMA",
  "channel_handle": "r/LocalLLaMA",
  "title": "2x R9700, 64 GB DDR5 is an absolute beast machine with vLLM Radiance / R9V and Qwen 3.8 27b and Flash next",
  "url": "https://www.reddit.com/r/LocalLLaMA/comments/1w92x3j/2x_r9700_64_gb_ddr5_is_an_absolute_beast_machine/",
  "external_url": null,
  "upload_date": "20260906",
  "published_at": "2026-09-06T17:48:00+00:00",
  "transcript": "I've been tinkering with local LLMs since the beginning of the year when I had an Intel Arc B580 and 32 GB of DDR5. Curiosity got the best of me and I bought the first R9700 about half a year ago, also because I wanted to upgrade my gaming graphics for 4k. As the 5090 was about 3 times as expensive, I had a \"sweet spot\", kind of. On the last prime days, I found a X870E mainboard for \\~150 € below the standard price, and it got to my head that I can use an upgraded machine for gaming and local inference tinkering.\n\nAnyways. Fast forward to this week, I now have the following setup\n\n* Ryzen 7500F\n* 64 GB DDR5 CL40 6400 MT/s\n* Asus ProArt Creator X870E\n* 2x R9700 32 GB, each running at PCIe 5.0 x8 (Gigagbyte)\n* Currently running ubuntu on an old Samsung EVO 860 1 TB drive; this will become intersting for the ngram / PLE offload; I have Windows and the gaming related stuff on a gen4 NVMe, but will soon add another Gen 5 NVMe with decent random reads\n\nThe only issue that I can report so far is that one of the cards runs quite hot, so I will definitely implement power limiting to 210 W and some light undervolting. The other card runs 10-15 °C cooler.. Case is a purebase 501 with 4 fans, 2 intake in front, one back and top for output.\n\nNow long story short I wanted to give some results of Qwen 3.8 27b FP8 and MXFP4, as well as Qwen 3.8 flash next after the first day tinkering with it. What I found super interesting is that the SATA SSD does not seem to be super terrible when using Qwen 3.8 flash next.\n\nConsidering the whole build costs \\~4k €, or more than 1k less than a single RTX 5090 with 32 GB, I kinda like this setup price/performance wise. Next step is checking context degradation / KV quants. I am using local inference mostly for deep research, summarization, image creation, light coding and non-trivial data analysis\n\nCheers\n\n## Qwen3.8 benchmarks on 2× Radeon AI PRO R9700\n\nHardware: 2× AMD Radeon AI PRO R9700 32 GB, 61 GiB system RAM  \nBenchmark: BetterBench 0.2.2, corpus v1.0, single-stream, greedy decoding, 2 warm-ups + 10 measured runs per category, 8k benchmark context.\n\n| Model | Weight format | Runtime | Server context | Max sequences | Speculative decoding | Weighted decode median | ITL 1% low | TTFT p50 | Prefill ~2k | Prefill ~4k | Prefill ~7k |\n|---|---|---|---:|---:|---|---:|---:|---:|---:|---:|---:|\n| Qwen3.8-27B | Quark AWQ MXFP4 | vLLM Radiance, TP2 | 131,072 | 1 | MTP, up to 8 tokens | 111.4 tok/s | 77.9 tok/s | 81 ms | 4,224 tok/s | 4,322 tok/s | 4,410 tok/s |\n| Qwen3.8-27B | Native block FP8 | vLLM Radiance, TP2 | 16,384 | 8 | MTP, up to 8 tokens | 87.6 tok/s | 61.9 tok/s | 73 ms | 4,134 tok/s | 4,329 tok/s | 4,305 tok/s |\n| Qwen3.8-Flash-Next | UD-IQ4_XS GGUF | R9V/vLLM, TP2, tiered expert offload | 131,072 | 1 | MTP, 2 tokens, FP8 draft | 35.4 tok/s | 27.3 tok/s | 290 ms | 1,727 tok/s | 1,986 tok/s | 1,925 tok/s |\n\n\n* Qwen 3.8 27b in FP8 and AWQ MXFP4 served with vLLM Radiance\n* Qwen 3.8 Flash next served with vLLM / R9V fork\n* Decode metrics come from the 10-pass standard run.\n* Prefill measurements use cold, nonce-prefixed prompts.\n* Prompt-token medians for the prefill columns were 1,556, 3,024 and 5,226 tokens.\n* No concurrency sweep was included in these results.\n* I expect decode of Flash next to increase a bit when an NVMe is used, and, as I am writing this and checked, I found EXPO was not enabled........oh my god I swear I turned it on when I updated the bios yesterday\n\n\n\n--- Top Comments ---\n\n\n[11 upvotes] What per stream decode do you get on 8 seq?\n\n[4 upvotes] What's the acceptance rate looking like on 8 tokens?   I'm seeing a sharp drop off on the 3rd token, but I'm using the integrated MTP head.   But it does look like for FP8 the R9700's might be the value kind for 27B.   I found a Pro5000 for \\~5K (48GB) which I'm sure would be fast but 48GB is right on the edge for 8 bit quant, 8 bit KV on 27B.\n\n[2 upvotes] Thanks for sharing. I'm looking at AMD more in this area. ",
  "transcript_chars": 3974,
  "ingested_at": "2026-09-07T01:30:03.410978+00:00",
  "source": "reddit",
  "yt_meta": {
    "score": 67,
    "upvote_ratio": 0.95,
    "num_comments": 47,
    "author": "smallDeltaBigEffect",
    "is_self": true
  }
}