{
  "video_id": "reddit_1wapjaw",
  "channel_slug": "LocalLLaMA",
  "channel_handle": "r/LocalLLaMA",
  "title": "OpenAI alleged of stealing mathematicians work",
  "url": "https://www.reddit.com/r/LocalLLaMA/comments/1wapjaw/openai_alleged_of_stealing_mathematicians_work/",
  "external_url": null,
  "upload_date": "20260908",
  "published_at": "2026-09-08T14:12:10+00:00",
  "transcript": "Privacy have been concern of many of us to have their own hardware to run llms, and here's another reason why: two mathematicians spent a year cracking one of the hardest problems in math and fed every draft of their works into Codex. A few days before they could publish, OpenAI suddenly showed up with the same solutions. When asked if their model (Sol and Astra) was trained on the pair's private chats, OpenAI did not answer the question.\n\nFull statement from them [https://cims.nyu.edu/\\~tristanb/statement.pdf](https://cims.nyu.edu/~tristanb/statement.pdf)\n\nFeels like big labs believe everything you did with the help of their models is theirs.\n\n\n\n--- Top Comments ---\n\n\n[1 upvotes] Your post is getting popular and we just featured it on our Discord! [Come check it out!](https://discord.gg/PgFhZ8cnWW)\n\nYou've also been given a special flair for your contribution. We appreciate your post!\n\n*I am a bot and this action was performed automatically.*\n\n[469 upvotes] That’s exactly why we need more open weight releases. Never trust any big AI company to manage your chats. They can use it to train their models whenever they want.\n\n[208 upvotes] read the pdf. he's not actually claiming they stole anything, he says outright he doesn't know if their codex sessions were used. and openai's thing is forced navier-stokes, the next step on the same route, not the same results\n\nthe bad part is the timeline. openai's first prompt went out days after they heard about the work, and bubeck allegedly wanted alpoge off the paper for working at anthropic. Bubeck says it's all false, response supposedly coming\n\nwouldn't have mattered if they ran local btw, the leak was a rumor through like three people. still, they asked openai point blank if their drafts went into training and got no answer. If they had gone with local you never have to ask!\n\n[106 upvotes] Even though they claim that you can turn off the data being used to train the models. It doesn’t say they won’t use your chats to improve ChatGPT itself beyond training models. That’s the part people don’t get. If you’re working on some cool ideas expect those ideas to be pillage.\n\nChatGPT is using your chats to improve itself in your active session, for example to solve problems by giving out prompts to orchestrate tasks across different machines using GitHub as the method to validate work. I’ve seen the model use ideas it gathered from my repos to assign task work to Claude opus running on three separate machines. Somehow it took evidence based commits overlayed it inside each task prompt it assigned i didn’t pioneer it but it’s the first I’ve noticed, not the first time using ChatGPT in this capacity but I’ve had an eye for watching these assistants become rodents. I told codex to increase fans speeds on my case fans and instead the model decided to investigate a project I didn’t want codex looking at. He clearly found it interesting and I’m assuming I know why.\n\n[93 upvotes] If you're a scientist and you're using these API models, you're being scammed.",
  "transcript_chars": 3037,
  "ingested_at": "2026-09-09T01:30:03.430228+00:00",
  "source": "reddit",
  "yt_meta": {
    "score": 987,
    "upvote_ratio": 0.93,
    "num_comments": 200,
    "author": "bakawolf123",
    "is_self": true
  }
}