{
  "video_id": "reddit_1w3n506",
  "channel_slug": "LocalLLaMA",
  "channel_handle": "r/LocalLLaMA",
  "title": "AVX2: Speed up large batch size prompt processing of IQ models by bartowski1182 · Pull Request #27402 · ggml-org/llama.cpp",
  "url": "https://www.reddit.com/r/LocalLLaMA/comments/1w3n506/avx2_speed_up_large_batch_size_prompt_processing/",
  "external_url": "https://github.com/ggml-org/llama.cpp/pull/27402",
  "upload_date": "20260831",
  "published_at": "2026-08-31T18:53:30+00:00",
  "transcript": "Faster prompt processing on CPU.\n\n\n\n--- Top Comments ---\n\n\n[19 upvotes] Hey, for once a thread that's relevant to me :(\n\n[12 upvotes] Hell yea, thanks for posting, will test tomorrow with Qwen 3.8 Flash IQ3_XSS\n\n[7 upvotes] Here is the [previous thread on it](https://www.reddit.com/r/LocalLLaMA/comments/1vtgyzf/comment/p4t4niu/?screen_view_count=5&ext-referrer=DIRECT) (when it was still a draft) with some more information and discussion.\n\n[3 upvotes] Every weight in the model decoded 512 times from the lookup table for a 512-token batch, that line does more for the 8x than the benchmark table does. Glad it's in.",
  "transcript_chars": 619,
  "ingested_at": "2026-09-01T01:30:04.999137+00:00",
  "source": "reddit",
  "yt_meta": {
    "score": 67,
    "upvote_ratio": 0.96,
    "num_comments": 19,
    "author": "jacek2023",
    "is_self": false
  }
}