{
  "video_id": "reddit_1tvwctd",
  "channel_slug": "MachineLearning",
  "channel_handle": "r/MachineLearning",
  "title": "NeurIPS used uncalibrated AI detector for desk rejections [D]",
  "url": "https://www.reddit.com/r/MachineLearning/comments/1tvwctd/neurips_used_uncalibrated_ai_detector_for_desk/",
  "external_url": null,
  "upload_date": "20260603",
  "published_at": "2026-06-03T17:28:03+00:00",
  "transcript": "I recently had a submission desk-rejected from the NeurIPS 2026 Position Paper Track for an alleged AI-policy violation. After corresponding with the track leadership and reading their public blog post, I think the broader methodological issue is worth discussing here.\n\nThe track used Pangram, a proprietary AI-text detector, as part of the desk-rejection process. I was told that the materials considered for desk rejection were:\n\n* the detector output\n* the authors’ AI-use attestation\n\nThis creates a potential circularity problem. If a high detector score is used to judge the author’s attestation as inconsistent, and that inconsistency is then used to justify desk rejection, the detector is not just an aid. It becomes a decisive part of the adjudication process.\n\nThe bigger issue is validation.\n\nThe NeurIPS blog describes tests using Pangram audits, older ACM FAccT papers, synthetic AI-generated position papers, and manually edited samples. But the target population was NeurIPS 2026 Position Paper submissions, whose ground-truth authorship process is unknown.\n\nSo the key question is:\n\n**What is the false-positive rate of the final decision procedure on the actual target distribution?**\n\nA false-positive rate measured on one distribution does not automatically transfer to another. If the actual submission pool produced a \"surprisingly high flagged rate\" (citation from NeurIPS blog post), that could indicate distribution shift / miscalibration.\n\nTo sanity-check the detector’s behavior, I also ran Pangram on recent 2026 papers authored by NeurIPS Position Paper Track Chairs. Pangram returned scores including:\n\n* 69% AI\n* 45% AI\n* 36% AI\n* 24% AI\n\nI am **not** claiming those papers were AI-written. For me, Pangram’s outputs alone does not permit such a conclusion. And that is exactly the point.\n\nUPD:\n\nHere is [NeurIPS original blogpost](https://blog.neurips.cc/2026/06/02/ai-generated-papers-in-the-neurips-2026-position-paper-track/)\n\nAnd here is the[ blogpost with the detailed critics](https://www.linkedin.com/pulse/we-shouldnt-desk-reject-papers-based-unvalidated-ai-sergey-berezin-orc6e/)\n\n\n\n--- Top Comments ---\n\n\n[45 upvotes] I ran some of my more obscure papers from pre 2022 through the systems and they also sometimes score high lol. These systems are pure bullshit. The conference is a joke for using them. \n\n[36 upvotes] That's just ironic.\n\n[9 upvotes] That's nonsense.",
  "transcript_chars": 2410,
  "ingested_at": "2026-06-04T01:30:09.377827+00:00",
  "source": "reddit",
  "yt_meta": {
    "score": 59,
    "upvote_ratio": 0.89,
    "num_comments": 39,
    "author": "Asleep-Requirement13",
    "is_self": true
  }
}