{
  "video_id": "reddit_1w89m36",
  "channel_slug": "MachineLearning",
  "channel_handle": "r/MachineLearning",
  "title": "GPT-6 reportedly jailbroken within 24 hours using an extended Task-in-Prompt (TIP) attack [N]",
  "url": "https://www.reddit.com/r/MachineLearning/comments/1w89m36/gpt6_reportedly_jailbroken_within_24_hours_using/",
  "external_url": null,
  "upload_date": "20260905",
  "published_at": "2026-09-05T19:11:16+00:00",
  "transcript": "A researcher has [reported](https://www.linkedin.com/posts/s-berezin_llm-aialignment-aisecurity-activity-7502013488412680192-IO6c/) a jailbreak of GPT-6 Astra within a day after release.\n\nThe attack is described as combination of TIP (Task-in-Prompt) attack from [ACL 2025 paper](https://aclanthology.org/2025.acl-long.334/) with four other unnamed techniques.\n\nTIP attacks exploit the model’s reasoning/instruction-following behaviour by hidding the harmful objective inside another task, like solving a cipher or executing a Python code. For GPT-6, the researcher says the original minimal TIP attack was no longer sufficient and had to be reworked.\n\nThey have reportedly disclosed the details privately to OpenAI rather than publishing the jailbreak.\n\nThe same researcher reported jailbreaking GPT-5 within an hour of its release a year ago.\n\n**Source:** [screenshot/post from the researcher](https://www.linkedin.com/posts/s-berezin_llm-aialignment-aisecurity-activity-7502013488412680192-IO6c/); their [ACL 2025 TIP paper](https://aclanthology.org/2025.acl-long.334/) linked in the original post.\n\n\n\n--- Top Comments ---\n\n\n[65 upvotes] The more interesting question: why isn’t that researcher in the safety team at OAI/Ant?\n\n[24 upvotes] Interesting since OpenAI says GPT6 is 100% aligned. This attack relies on the model decoding some malicious prompt instruction but then never saying it out loud and then executing that instruction. \n\nPerhaps doing some probing into something like the JSpace could reveal where these malicious instructions are being held and we can have better detection mechanisms inside the model weights rather than  at the encoding or decoding steps. ",
  "transcript_chars": 1681,
  "ingested_at": "2026-09-06T01:30:07.951675+00:00",
  "source": "reddit",
  "yt_meta": {
    "score": 146,
    "upvote_ratio": 0.96,
    "num_comments": 33,
    "author": "Asleep-Requirement13",
    "is_self": true
  }
}