{
  "video_id": "reddit_1vpuhh1",
  "channel_slug": "LocalLLaMA",
  "channel_handle": "r/LocalLLaMA",
  "title": "Paper claims RL for reasoning only changes 1-3% of tokens, and they replicate the gains without RL at ~1000x less compute",
  "url": "https://www.reddit.com/r/LocalLLaMA/comments/1vpuhh1/paper_claims_rl_for_reasoning_only_changes_13_of/",
  "external_url": "https://arxiv.org/abs/2605.06241",
  "upload_date": "20260816",
  "published_at": "2026-08-16T11:21:24+00:00",
  "transcript": "\n\n--- Top Comments ---\n\n\n[56 upvotes] Big if true. The uncharted territories of new LLM architectures or training methods are still huge.\n\n[33 upvotes] >Figure 1: RL edits are rare, conservative, and **concentrated at decision points**\n\nThis is the fundamental issue, one which Jonathan Blow pointed out ages ago and is stuck in my head\n\nThey're Large **Language** Models, not decision models. If you give me some decision tokens to train on, I will give you a model that can make decisions\n\nAnd language is, at best, a very poor approximant of how decisions are made in our brains (at least of those of us, capable of making decisions)\n\nFor a long time I've been wondering, perhaps a spiky neural net can learn a latent space of decisions and be bolted on top of an LLM. SNN does the little neuron fight that happens in our brains that produces decisions, and the LLM implements it\n\n  \nEDIT:\n\nAlso it's absolutely hilarious they called their model **ReasonMaxxer**\n\n[9 upvotes] > Through token-level analysis across multiple model families and RL algorithms, we find that RL's beneficial footprint is a sparse, predictable correction concentrated at high-entropy decision points where the model is uncertain which branch to take. Only 1-3% of token positions are affected, **the promoted token always lies within the base model's top-5 alternatives** […]\n\nI find it completely impossible to believe that tokens promoted by RL **always** lie within the top-5 as claimed here. There’s absolutely no way this is true. It’s so easy to imagine high-entropy distributions where the top-10 or so have essentially the same probability, and if you look at enough of those you are guaranteed to find an instance where the promoted token comes from further down the ranking.",
  "transcript_chars": 1764,
  "ingested_at": "2026-08-16T13:30:02.760766+00:00",
  "source": "reddit",
  "yt_meta": {
    "score": 92,
    "upvote_ratio": 0.91,
    "num_comments": 20,
    "author": "juanviera23",
    "is_self": false
  }
}