{
  "video_id": "reddit_1w6knep",
  "channel_slug": "singularity",
  "channel_handle": "r/singularity",
  "title": "The prevalent problem of misleading benchmark reporting (re: Astra)",
  "url": "https://www.reddit.com/r/singularity/comments/1w6knep/the_prevalent_problem_of_misleading_benchmark/",
  "external_url": null,
  "upload_date": "20260903",
  "published_at": "2026-09-03T21:23:40+00:00",
  "transcript": "OpenAI's reported benchmarks for Astra's ARC-AGI-3 is one of the most egregious recent examples I have seen of technically true metric reporting being used to deliberately mislead the masses. For context, there is an OpenAI screencap currently at the top of r/singularity's hot page of Astra achieving 98.6% on ARC-AGI-3 compared to 7.8% for GPT 5.6 Sol and 30.2% for Claude Opus 5. Holy shit, right? ASI achieved, right?\n\nUnfortunately, those figures taken in a vacuum leave out ***very*** important context: Astra's agentic harness had significant additional features that GPT 5.6 Sol and Claude Opus 5 did not have access to - specifically reasoning trace retention and custom compaction (source: [https://arcprize.org/leaderboard](https://arcprize.org/leaderboard) ).\n\nMy main takeaway is basically: The most honest way to compare Astra with Opus 5/Sol on this benchmark would have been to either 1) measure their ARC-AGI-3 performances on the same provider adapter harness (where Astra's 98.6% came from), or 2) compare them on the standard ARC-AGI-3 harness. On the standard harness Astra achieves 62.7% vs Opus 5's 30.2% vs Sol's 7.8%. Still a very large gap, but much less misleading than the comparison OpenAI chose to report. (source: [https://arcprize.org/leaderboard](https://arcprize.org/leaderboard) )\n\nNot an Anthropic fanboy in any sense of the word, btw. I thought Opus 5 was benchmaxxed and pray on Anthropic's downfall every day. But the Astra benchmark glazing made it clear that restraint needs to be had in people's reactions to its benchmarks (if Opus 5 didn't already convince you to not treat benchmarks as gospel) before anyone has even had time to extensively test it in real world use cases.\n\n\n\n--- Top Comments ---\n\n\n[40 upvotes] Worth noting that it still scores 66% on the standard harness. Well ahead of everything else.\n\n[38 upvotes] You need to be fair that the thing with ARC deleting context and not allowing models to remember things, when they are totally capable of doing so, is very silly. It's not a harness calling \"SOLVE_ARC.MD\", it's one that better allows a model to actually do the work. Opus 5 also did very well with the same thing.\n\n[21 upvotes] To be fair, even GPT 5.6 Sol is estimated to only score about 30% using the same harness than GPT 6 Astra just got 99.9% with\n\nAlso, it gets 60% without the custom harness (which obliterates literally every other model in existence)\n\nhttps://preview.redd.it/r56rdpy5ldnh1.jpeg?width=1440&format=pjpg&auto=webp&s=865317488871d2e7e1b32868767907a4ad23792d\n\n[12 upvotes] What you're saying is simply wrong. They did not use a \"custom harness specifically designed for the benchmark\". \n\nThe harness is simply the Responses API, which is the regular API OpenAI recommends every customer to use. It has absolutely nothing to do with the ARC AGI 3 benchmark. It's completely different from the other harnesses you mentioned that were designed specifically for beating the ARC AGI 3 benchmark in an ideal way. \n\nAnd that is why ARC also shows the 99% score on their own leaderboard, they do allow general purpose harnesses behind the official API for official scores. They do not show any results on their leaderboard that are made with narrow harnesses designed for beating the benchmark, those are not allowed.\n\n[14 upvotes] https://x.com/fchollet/status/2095598451115614371\n\nChollet seems fine with it",
  "transcript_chars": 3390,
  "ingested_at": "2026-09-04T01:30:16.065155+00:00",
  "source": "reddit",
  "yt_meta": {
    "score": 77,
    "upvote_ratio": 0.76,
    "num_comments": 45,
    "author": "PsychologicalSoup251",
    "is_self": true
  }
}