{
  "video_id": "reddit_1ts5b6u",
  "channel_slug": "singularity",
  "channel_handle": "r/singularity",
  "title": "Opus 4.8 Leads the Singularity Gate: New Benchmark for AI predicting paradigm-breaking scientific discoveries after model traning cutoff",
  "url": "https://www.reddit.com/r/singularity/comments/1ts5b6u/opus_48_leads_the_singularity_gate_new_benchmark/",
  "external_url": null,
  "upload_date": "20260530",
  "published_at": "2026-05-30T17:01:31+00:00",
  "transcript": "Just as I released a new benchmark called the Singularity Gate, which tests whether frontier AI models can predict paradigm-breaking scientific discoveries published after their training cutoff, Opus 4.8 was launched.\n\nIt took a couple of days to update the leaderboard because the contamination audit flagged a few discoveries for Opus 4.8. These have been removed from the corpus. As a result, there are minor score changes among the models, though the rankings remain unchanged.\n\nOpus 4.8 represents an incremental improvement and surpasses 20%. However, we still do not have a model that fully predicts a discovery.\n\n* **Top score:** 20.47% (partial credit, Opus 4.8)\n* **Fully correct outcome rate:** 0% across all evaluated models\n\n**Reminder:** Passing the Singularity Gate is necessary, though not sufficient, for autonomous AI-driven discovery. A model that can predict paradigm-breaking discoveries isn't necessarily Einstein-level, but a model that cannot definitely is not.\n\nAll models have been tested in their native agentic harness (claude code, codex, gemini cli) and allowed tool use. Web search has been disabled.\n\nhttps://preview.redd.it/cibjl0io2b4h1.png?width=883&format=png&auto=webp&s=f2dfd8220b878ccdbe006427360154a93274ec9d\n\nhttps://preview.redd.it/djvt2b4x2b4h1.png?width=657&format=png&auto=webp&s=a18bbd54555f0660d86da7f9d2a0dbde35ae63f8\n\nhttps://preview.redd.it/0jca067z2b4h1.png?width=922&format=png&auto=webp&s=a998f48f544caf2eeec9a40d8f3eb2401a074be5\n\nThese are partial-credit scores. I'm happy to discuss the methodology, related work, or framing in the comments.\n\n**Paper:** [https://doi.org/10.5281/zenodo.20358378](https://doi.org/10.5281/zenodo.20358378)  \n**Website:** [https://singularitygate.org](https://singularitygate.org)\n\n\n\n--- Top Comments ---\n\n\n[7 upvotes] I can't wait to see what Mythos does. \n\nThis is truly amazing and as we approach RSI, the models will actually be doing the research.. \n\n[3 upvotes] Very cool idea, thanks for checking it. Is it a 20% threshold going by the standard p > 0.05 or are there other reasons? I'm a layperson so is it essentially predicting a tiny bit better than chance now, is what you're saying?\n\n[4 upvotes] Cut-off models should be a priority, it's the only way to test if transformers/agentic AI are good at deriving past discoveries and therefore new ones\n\n[2 upvotes] I have to say, this is currently the best benchmark for telling us how close we are to singularity. This is a great idea for a benchmark and hope this gets attention to expand and get better. Great work!",
  "transcript_chars": 2560,
  "ingested_at": "2026-05-31T01:30:36.317069+00:00",
  "source": "reddit",
  "yt_meta": {
    "score": 59,
    "upvote_ratio": 0.79,
    "num_comments": 19,
    "author": "queenofartists",
    "is_self": true
  }
}