{
  "video_id": "reddit_1uegdu0",
  "channel_slug": "LocalLLaMA",
  "channel_handle": "r/LocalLLaMA",
  "title": "Big News for AMD / Strix Halo+ Owners",
  "url": "https://www.reddit.com/r/LocalLLaMA/comments/1uegdu0/big_news_for_amd_strix_halo_owners/",
  "external_url": null,
  "upload_date": "20260624",
  "published_at": "2026-06-24T15:16:34+00:00",
  "transcript": "Admittedly this is news for me, but I'm hoping it could be of some use to others here as well!\n\nSo, THE NPU IS USABLE!!\n\nI've owned an AMD Ryzen 395 Max AI+ (or whatever the naming is lol) for about a year now and have relied solely on GGUFs and Vulkan. I acknowledge that the AMD Ryzen AI team has been working hard to get their ROCm software up to speed w/ their hardware.\n\n[https://kyuz0.github.io/amd-strix-halo-toolboxes/](https://kyuz0.github.io/amd-strix-halo-toolboxes/)\n\nThis database did NOT look so ROCm friendly 6 months ago.\n\n1. Why should I care?\n2. If you own a device w/ both an NPU and a iGPU (like the strix halo series) then you WANT hybrid models. The NPU is CRAZY FAST at PromptProcessing, and can run parallel to gpu firing.\n3. Okay, What is Hybrid Mode?\n4. So, LLMs can run through the NPU only. If they're built for it. Check out \"FastFlowLM NPU\" models for examples that do that. BUT HYBRID mode combines the best of both, and FINALLY utilizes the hardware purchased nearly a year go (for some, more than that).\n5. What can i do to test this?\n6. Download Lemonade! Thanks to their efforts that focus primarily on Ryzen AI and working directly w AMD, I've FINALLY got my machine working in ways it couldn't a year ago and Lemonade made it happen. It's GUI is ultra bare-bones and I wouldn't recommend it for any actual agentic/chat/harness usage BUT being able to sanity-test software without investing days or weeks into it?\n\n10/10\n\nHere's the link: [lemonade-server.ai](http://lemonade-server.ai)\n\nSpeaking of links, read more about Hybrid Mode and making your own Hybrid Models here:  [https://ryzenai.docs.amd.com/en/latest/llm/overview.html](https://ryzenai.docs.amd.com/en/latest/llm/overview.html)\n\n\n\n\\---\n\nSo, that's it. Just wanted to share. REALLY EXCITED that my year old computer is still advancing in the software science of it all.\n\nI have a single wishlist/request now: MTP-supported Hybrid Models. Qwen 3.6 has that speedup tech introduced by Unsloth, and AMD has a guide for \"new processor shapes\" since 3.6 GGUF can't simply be \"converted to ONNX\". Here's that guide: [https://ryzenai.docs.amd.com/en/latest/oga\\_op\\_prepare.html](https://ryzenai.docs.amd.com/en/latest/oga_op_prepare.html)\n\nIf anyone attempts it, please share on huggingface!\n\nThis was all written by hand btw, no llm assistance, just passionate dev obsessed w \"new shiny\".\n\n\n\n--- Top Comments ---\n\n\n[28 upvotes] Any examples of the prompt processing being faster than running on GPU? Like what model are you running and  what are the prompt processing speeds on NPU compared to running on GPU?\n\n[30 upvotes] The NPU is tiny and designed to save power rather than running the full GPU.  The NPU is worthless unless you are trying to run a tiny model and save power.  If you try to run both you run into the real limit which is memory bandwidth.  Since you are using up memory bandwidth the main GPU will run worse.\n\n[11 upvotes] Lately trying to see what we can do with the NPU so a tool for telemetry https://github.com/boxwrench/xdna-top\nThen this was a feasibility study to run a model on the NPU to compact context on the main model. \nhttps://github.com/boxwrench/REM\nIf anything is helpful please share results\n\n[5 upvotes] I have been using the NPU for a couple of month on Fedora (and in Windows from much more time). Remember to change IOMMU on \"ON\". Best model is Gemma e4B which runs at 12 tps.\n\n[3 upvotes] Wait, so has someone converted Qwen 3.6 MTP already?  I'm floating between 21-28t/s on generation right now with ROCm/Llama.cpp ... I'd welcome some speed up on that.",
  "transcript_chars": 3594,
  "ingested_at": "2026-06-25T01:30:04.875687+00:00",
  "source": "reddit",
  "yt_meta": {
    "score": 84,
    "upvote_ratio": 0.83,
    "num_comments": 62,
    "author": "CSEliot",
    "is_self": true
  }
}