{
  "video_id": "reddit_1tw0cqv",
  "channel_slug": "artificial",
  "channel_handle": "r/artificial",
  "title": "Google just dropped Gemma 4 12B on your laptop!!",
  "url": "https://www.reddit.com/r/artificial/comments/1tw0cqv/google_just_dropped_gemma_4_12b_on_your_laptop/",
  "external_url": null,
  "upload_date": "20260603",
  "published_at": "2026-06-03T19:42:27+00:00",
  "transcript": "bro google just casually released a 12 billion parameter multimodal model that runs on 16gb of ram\n\nlike… your macbook pro can run this. no cloud. no api calls. no monthly bill.\n\nit’s encoder-free, handles images and text, apache 2.0 license so you can do whatever with it commercially\n\nthe “cloud is the only way” narrative is dying fast. on-device AI is not a gimmick anymore, it’s where the serious money is going\n\n\n\n--- Top Comments ---\n\n\n[40 upvotes] Edge compute from specialized arm / asics is the future for personal compute.  The datacenters are for training frontier models for enterprise applications.  I recall seeing something recently where a chip designer was able to hard burn the code for a llm directly into a die, can't find the link though.\n\n[12 upvotes] wait what is this actually? what can I do with a local llm? and why is it better than cloud? also how good is gemma?\n\n[13 upvotes] The encoder-free architecture is the real differentiator here. Most multimodal models use a separate vision encoder which compresses image data before the LLM sees it. Gemma processes images natively in the transformer, making it much better at OCR and document QA than pure text benchmarks suggest.",
  "transcript_chars": 1205,
  "ingested_at": "2026-06-04T01:30:13.866046+00:00",
  "source": "reddit",
  "yt_meta": {
    "score": 165,
    "upvote_ratio": 0.88,
    "num_comments": 69,
    "author": "NewMuffin3926",
    "is_self": true
  }
}