{
  "video_id": "reddit_1ussasa",
  "channel_slug": "LocalLLaMA",
  "channel_handle": "r/LocalLLaMA",
  "title": "At most my Strix Halo uses $0.48 a day",
  "url": "https://www.reddit.com/r/LocalLLaMA/comments/1ussasa/at_most_my_strix_halo_uses_048_a_day/",
  "external_url": null,
  "upload_date": "20260710",
  "published_at": "2026-07-10T16:18:38+00:00",
  "transcript": "This is something that never gets mentioned when people complain that it's slow and new users are told to avoid them. This 48 cent figure is worst case scenario, running multiple models/compiling hitting CPU, GPU, and NPU at the same time for 24 hours a day. I can handle only 50tps on Q8\\_XL Qwen 3.6 35B when it's silent, sipping power, and is the size of a small router.\n\nI know your Nvidia card is significantly faster, but if you even consider using more than just raw GPU memory speed/compute or you are concerned with size/noise/energy, I don't see how there is much of a competition. An A6000 is 300W for the card alone , which is double what the Strix Halo devices total power budget is.\n\nEven with the current inflated prices, I think these things have insane value. They provide significantly more than just the GPU/RAM. Anything that isn't used for inference is open for hosting any services you want, it's such a versatile package.\n\nSame goes for the Macs,\n\n\n\n--- Top Comments ---\n\n\n[31 upvotes] 2x spark (Asus ascent) which are only $400 more expensive consume 200 watts together to run deepseek v4 flash at 60tps single stream 300tps at c16. With c1 at 1M depth still at 35tps.\n\n[27 upvotes] As a discrete GPU user, I think this is a really good insight. \n\nI think most of us here — myself included a lot of times — are constantly looking to basically recreate the enterprise cloud API performance experience on our local machines, and speed is the easier to see than absolute quality. Operational cost is also buried in my electrical bill.\n\nI think we’d be better served by adapting workflows and use-cases to the real world limitations we have in hardware (that I can’t justify buying and running multiple RTX Pro 6000s, A100s, etc.). I’m really liking my new adaptation to asynchronous development, which it sounds like a Strix Halo system would actually be pretty good for. For software development for example, once I have a spec document and break it out into specific tasks, I spin up vLLM to run a bunch of concurrent agents to go do the actual execution. It’s way more efficient, but it still may take a while, so I queue things up, then run them overnight while I sleep and have a prototype by morning. I think that long-slow, but much more efficient compute tasks seems really cool and like a good path for the future.\n\n[15 upvotes] You are totally wrong. I have both a Strix Halo and an RTX 6000.\n\nThe Strix Halo runs at 150W during inference and is \\~ 30W idle as measured at the outlet of my UPS.  \n  \nAn RTX 6000 Max Q peaks at 300W, and is typically 225W or so during single user inference, and sits at 22W idle.  \nIn addition, the desktop system it sits in uses power. From the wall, in total, you are looking at around 400W under full GPU load, 325W when doing single user token generation, and idle is perhaps 70W, highly dependent on your system ( i have lots of storage and HDD in my system and also use it a CI runner so can't give you an exact idle number)\n\nThe difference here, is that despite the RTX 6000 using 2.5x more power at full load, it is at least 10x faster. So it uses more power, for way less time.  \n  \nSingle user, running Qwen 122B = 160-200 tokens a second, vs around 20-30 tokens a second on Strix Halo  \nPrompt processing, 2000 vs 200 \n\nBut where the RTX 6000 really shines is concurrency - you can load it down with 4 concurrent token generation sessions, and it barely changes the speed.\n\nElectricity is measured in Kilowatt Hours, the \"hours\" part of this is actually very important, you can't just\n\n[7 upvotes] Power efficiency is really important. When people talk about computers they usually think about how tokens the computer can process per second.. If you are using a computer model at home every day you should also think about the noise it makes how much power it uses and how much it costs to own it. These things are just as important, as how the computer can perform at its best. Power efficiency and total cost of ownership of the computer model are important things to consider.\n\n[8 upvotes] Strix Halo owner here. I am extremely happy to have this technology marvel, but it ain't silent, it has rather annoying cooling solution, like a cheap laptop one, with high pitch and constant speed change, lol.\n\nI guess it's silent when I am browsing or something, not LLM's.",
  "transcript_chars": 4343,
  "ingested_at": "2026-07-11T01:30:10.849278+00:00",
  "source": "reddit",
  "yt_meta": {
    "score": 96,
    "upvote_ratio": 0.86,
    "num_comments": 89,
    "author": "Forward_Jackfruit813",
    "is_self": true
  }
}