{
  "video_id": "reddit_1w3ex16",
  "channel_slug": "artificial",
  "channel_handle": "r/artificial",
  "title": "Sony and Warner just sued Anthropic for the exact same piracy Anthropic already admitted to and paid $1.5B for",
  "url": "https://www.reddit.com/r/artificial/comments/1w3ex16/sony_and_warner_just_sued_anthropic_for_the_exact/",
  "external_url": null,
  "upload_date": "20260831",
  "published_at": "2026-08-31T14:09:25+00:00",
  "transcript": "Sony Music Publishing and Warner Chappell filed suit against Anthropic, Dario Amodei, and co-founder Benjamin Mann on August 28. What's unusual is that the underlying facts aren't in dispute anymore.\n\n  \nLast September, in the Bartz case, a federal judge ruled that training an AI model on copyrighted text was legal, but downloading the training copies via piracy was not. Anthropic settled that case for $1.5 billion after admitting Mann personally torrented over five million books from Library Genesis in 2021, and staff pulled two million more from Pirate Library Mirror in 2022.\n\n  \nSony and Warner's complaint cites those exact same downloads, now tied to MusixMatch and LyricFind lyric datasets. They're not asking a court to rule on anything new, they're applying a ruling that already exists to a different set of copyrighted works. Statutory damages run $150,000 per work, so the number could dwarf the book settlement depending on how many songs are in scope.\n\n  \nWhat I don't have a good answer for: once a company settles one IP class action over a specific data-acquisition method, does that admission become effectively permanent exposure for every other rightsholder whose work touched the same pirated corpus? Is there a legal mechanism that closes that door, or is Anthropic just going to keep getting sued by whoever's catalog turns up in the same torrent logs?\n\n\n\n--- Top Comments ---\n\n\n[17 upvotes] so they admitted to torrenting millions of books and now surprise surprise the same pirated stash had song lyrics in it too\n\n  \nthis is like robbing a bank and then telling the cops \"yeah we took the money but we only meant to spend it on groceries not on a car\" the law does not care what you planned to do with the stolen goods\n\n  \nalso wild that the cofounder personally did the torrenting like he couldnt even hire some intern for that\n\n[12 upvotes] Considering that these companies could not exist without the foundational data they all stole to build their first models, yes.\n\nRemember when OpenAI released SORA and it started producing videos of everything copyrighted or IP-thefted on the planet with reckless abandon?  That was only possible because of how much data they stole to train it.  All these companies are the same and I wish them the worst.\n\n[2 upvotes] At $150k per work, ten thousand lyrics comes to the $1.5B they already paid for the books. The music catalogues only need a few thousand songs in that same torrent stash for this suit to dwarf it.",
  "transcript_chars": 2491,
  "ingested_at": "2026-09-01T01:30:14.469251+00:00",
  "source": "reddit",
  "yt_meta": {
    "score": 51,
    "upvote_ratio": 0.9,
    "num_comments": 25,
    "author": "Servola-Journal",
    "is_self": true
  }
}