{
  "video_id": "reddit_1ur71hq",
  "channel_slug": "LocalLLaMA",
  "channel_handle": "r/LocalLLaMA",
  "title": "Grok 4.5 released - GLM-5.2 shows up in xAI's own charts, 2.6 pts behind on SWE Bench Pro",
  "url": "https://www.reddit.com/r/LocalLLaMA/comments/1ur71hq/grok_45_released_glm52_shows_up_in_xais_own/",
  "external_url": null,
  "upload_date": "20260708",
  "published_at": "2026-07-08T21:57:53+00:00",
  "transcript": "[https://x.ai/news/grok-4-5](https://x.ai/news/grok-4-5)\n\nQuick numbers from the launch page:\n\n* $2/M input, $6/M output. Closed weights. No EU until mid-July.\n* SWE Bench Pro: Fable 80.4% > Opus 4.8 69.2% > **Grok 4.5 64.7% > GLM-5.2 62.1%** \\> GPT 5.5 58.6%\n* Token efficiency is their headline: \\~16k output tokens per SWE Bench Pro task vs Opus's 67k (4.2x fewer), served at 80 TPS\n* Trained on \"tens of thousands of GB300s\", alongside Cursor\n\nTwo things worth noting for this sub:\n\n1. An MIT-licensed model you can self-host is 2.6 points behind their brand-new frontier model on SWE Bench Pro - and ahead of GPT 5.5. That's their own marketing chart.\n2. Harness fine print: Grok scores 62% on DeepSWE 1.0 \"within each provider's harness\" but 53% on DeepSWE 1.1 run independently by DataCurve. Fable went *up* on the independent run (66-> 70). Draw your own conclusions.\n\nNumbers from the linked page\n\n\n\n--- Top Comments ---\n\n\n[39 upvotes] They admitted they trained on cursorbench...\n\nI tried it. It seems to be better than sonnet 5, but it's not better than opus. At least the tok/s is fast so far, I like that\n\nIt seems like it lacks... depth? Some tasks glm 5.2 handled better\n\nEDIT:  \ncitation:  \n\n\n>\\> \\* Grok 4.5 has an advantage on CursorBench: an earlier snapshot of the Cursor codebase was accidentally included in training. The exact score impact is unclear. That data has been removed for future models, and was not present for Composer models in the past.”\n\n[15 upvotes] Wen grok 3 OSS?\n\n[16 upvotes] I don't trust anything an X company puts out. It's clear that Elon will do whatever it takes to get what he wants. Break laws, manipulate people via Social, manipulate his stock, and probably even manipulate benchmarks.\n\n[9 upvotes] From my limited testing - it's very good.",
  "transcript_chars": 1793,
  "ingested_at": "2026-07-09T01:30:28.665213+00:00",
  "source": "reddit",
  "yt_meta": {
    "score": 83,
    "upvote_ratio": 0.84,
    "num_comments": 17,
    "author": "qubridInc",
    "is_self": true
  }
}