{
  "video_id": "reddit_1u2tk0i",
  "channel_slug": "MachineLearning",
  "channel_handle": "r/MachineLearning",
  "title": "Anthropic walks back policy on silent nerfing for AI/ML, will notify users [N]",
  "url": "https://www.reddit.com/r/MachineLearning/comments/1u2tk0i/anthropic_walks_back_policy_on_silent_nerfing_for/",
  "external_url": null,
  "upload_date": "20260611",
  "published_at": "2026-06-11T08:51:10+00:00",
  "transcript": "From Wired:\n\n> “We’re changing Fable 5’s safeguards for frontier LLM development to make them visible.” Anthropic said in a statement to WIRED. “We made the wrong tradeoff and we apologize for not getting the balance right.”\n\n> Anthropic now says it’s changing course, and that Claude Fable 5’s safeguards for AI development will be visible to users. If the company suspects a user is trying to use Claude to build a highly capable AI it will alert them that it’s either refusing the request, or rerouting the user to a less capable model.\n\nFull article: https://www.wired.com/story/anthropic-responds-to-backlash-on-claudes-secret-sabotage-on-ai-research/\n\n\n\n--- Top Comments ---\n\n\n[69 upvotes] That was a hilariously bad policy because they would never beat the silent nerf allegations whenever people perceived a quality regression in fable\n\n[34 upvotes] This would be great news, explicit refusals sound much better. There were many reports of false positives from the safeguards that blocked the models for genuine use cases in cybersecurity, biology, physics etc. It wouldn't be too surprising if the self-handicap would likewise trigger even by mistake on arbitrary ML work.\n\nBut as a practitioner, this raises some fundamental questions now whether their tools can be trusted. From now on, when they fail, could it be due to just lacking performance or because they are silently working against the user? Feels like a hammer that is designed to hit your thumb when it reckons that your are building a hardware store.\n\nSeems more important than ever to pursue alternatives like Open Code with e.g. DeepSeek V4 Flash (and hopefully more decentralized or independently hosted variants eventually).\n\n[28 upvotes] Am i reading this right? So now they are going to alert you but still prevent you from making LLMs by either explicitly refusing or routing to a less capable model?",
  "transcript_chars": 1881,
  "ingested_at": "2026-06-11T13:30:30.555891+00:00",
  "source": "reddit",
  "yt_meta": {
    "score": 97,
    "upvote_ratio": 0.95,
    "num_comments": 22,
    "author": "goldcakes",
    "is_self": true
  }
}