{
  "video_id": "reddit_1vyynw2",
  "channel_slug": "singularity",
  "channel_handle": "r/singularity",
  "title": "GLM-5.3-Flash: Frontier Intelligence, Flash Cost",
  "url": "https://www.reddit.com/r/singularity/comments/1vyynw2/glm53flash_frontier_intelligence_flash_cost/",
  "external_url": "https://z.ai/blog/glm-5.3-flash",
  "upload_date": "20260826",
  "published_at": "2026-08-26T14:28:11+00:00",
  "transcript": "\n\n--- Top Comments ---\n\n\n[61 upvotes] This is the model previously served anonymously as **Ox Alpha**\n\n* 320B total parameters, with 18B active per token (MoE)\n* 1M token context window\n* Native multimodality\n* Standard API pricing is **$0.15 / 1M input tokens** and **$0.50 / 1M output tokens**\n* All Ox Alpha traffic was served on Chinese AI chips\n\n[14 upvotes] Trading blows with Flash-3.7. Would like to try out. However I wish they distill architecture down to something really small. \n\n[13 upvotes] > \nAll Ox Alpha traffic was served on Chinese AI chips\n\nWow… With only 320B parameters, an open weight model here would be amazing…",
  "transcript_chars": 636,
  "ingested_at": "2026-08-27T01:30:16.138240+00:00",
  "source": "reddit",
  "yt_meta": {
    "score": 167,
    "upvote_ratio": 0.97,
    "num_comments": 28,
    "author": "dydynam",
    "is_self": false
  }
}