{
  "video_id": "reddit_1uflt11",
  "channel_slug": "singularity",
  "channel_handle": "r/singularity",
  "title": "what's the last thing an AI agent did that surprised you, not on a benchmark but in the real world",
  "url": "https://www.reddit.com/r/singularity/comments/1uflt11/whats_the_last_thing_an_ai_agent_did_that/",
  "external_url": null,
  "upload_date": "20260625",
  "published_at": "2026-06-25T20:37:15+00:00",
  "transcript": "ive noticed the gap between agent demos and agent reality is doing something strange to my sense of where we are\n\nevery few weeks theres a launch, autonomous this, runs your whole workflow that. the most recent one going around is the Airtable founders agent product claiming it ran its own launch campaign. impressive if true. but i genuinely cant tell anymore whats a real capability jump vs whats a well shot demo plus a lot of human cleanup offscreen\n\nso im trying to recalibrate with actual data instead of vibes\n\nwhat is the single most surprising thing youve personally watched an agent do, where you went “oh, i didnt think it could do that yet.” not a benchmark score, not a thread you saw, something you ran yourself. and equally useful, the thing it confidently failed at that you assumed would be trivial\n\nim asking because  i think this sub talks about takeoff in the abstract a lot but the ground truth, what these things can and cant actually do unsupervised right now, is weirdly hard to find. the demos all say yes and the cynics all say no and the real answer is presumably messy and specific\n\nill start in the comments\n\nhttps://preview.redd.it/rapf2n0hlh9h1.png?width=1600&format=png&auto=webp&s=c569e097ff2eaa4abcdf9f28d41e30f6f76ec589\n\n\n\n--- Top Comments ---\n\n\n[21 upvotes] If you let claude use python it can pretty effectively figure out any black box piece of conversion software (including any bugs it has) by throwing its own examples through it as input and analysing the output. \n\nAnd then it can write the logic captured in Python into any language you like. \n\nMuch faster than iterating through attempts in the final language you wanted. \n\n[20 upvotes] the biggest surprise for me was how good agents have gotten at fixing their own mistakes instead of just failing immediately\n\n[8 upvotes] sometimes i wonder if future generations will find it weird that we trained AIs on humanity’s past and then acted surprised when they started helping us create humanity’s future",
  "transcript_chars": 1998,
  "ingested_at": "2026-06-26T01:30:19.588621+00:00",
  "source": "reddit",
  "yt_meta": {
    "score": 90,
    "upvote_ratio": 0.95,
    "num_comments": 33,
    "author": "SupermarketSmooth968",
    "is_self": true
  }
}