{
  "video_id": "reddit_1w8okha",
  "channel_slug": "OpenAI",
  "channel_handle": "r/OpenAI",
  "title": "GPT-6 reportedly jailbroken within 24 hours using an extended Task-in-Prompt (TIP) attack",
  "url": "https://www.reddit.com/r/OpenAI/comments/1w8okha/gpt6_reportedly_jailbroken_within_24_hours_using/",
  "external_url": null,
  "upload_date": "20260906",
  "published_at": "2026-09-06T06:42:36+00:00",
  "transcript": "A researcher has [reported](https://www.linkedin.com/posts/s-berezin_llm-aialignment-aisecurity-activity-7502013488412680192-IO6c/) a jailbreak of GPT-6 Astra within a day after release.\n\nThe attack is described as combination of TIP (Task-in-Prompt) attack from [ACL 2025 paper](https://aclanthology.org/2025.acl-long.334/) with four other unnamed techniques.\n\nTIP attacks exploit the model’s reasoning/instruction-following behaviour by hidding the harmful objective inside another task, like solving a cipher or executing a Python code. For GPT-6, the researcher says the original minimal TIP attack was no longer sufficient and had to be reworked.\n\nThey have reportedly disclosed the details privately to OpenAI rather than publishing the jailbreak.\n\nThe same researcher reported jailbreaking GPT-5 within an hour of its release a year ago.\n\n**Source:** [screenshot/post from the researcher](https://www.linkedin.com/posts/s-berezin_llm-aialignment-aisecurity-activity-7502013488412680192-IO6c/); their [ACL 2025 TIP paper](https://aclanthology.org/2025.acl-long.334/) linked in the original post.\n\n\n\n--- Top Comments ---\n\n\n[47 upvotes] this guy’s track record is almost a tradition at this point\n\n[27 upvotes] Just the TIP?\n\n[13 upvotes] I get annoyed by posts like this... It encourages readers to think, “I accomplished what four professional teams couldn’t.” Typical behavior from the security blowhards who love to brag like this. I more and more understand why Linus Torvalds hates these people.\n\nSimply outputting the names \"The Pirate Bay,\" \"Mullvad,\" or \"qBittorrent\" is not inherently a safety violation. And even if is, a jailbreak can be real while being low severity. If it genuinely defeated an intended restriction, the label could be technically justified. But that wouldn’t establish that the same method defeats safeguards against substantially more harmful behavior. If Astra was tricked into thinking it is participating in a benign coding puzzle or a research discussion (which is exactly what the TIP attack does), providing those names is just the model executing standard instruction-following. A true, critical \"jailbreak\" would be forcing the model to generate a custom, malicious phishing script or explicit instructions for a cyberattack, not just naming a popular VPN.\n\nAnd the comparison with four red-teaming organisations only works if his example meets their success criteria, scope, and severity threshold\n\n[11 upvotes] why doesn't openai just hide this guy, he's making a fool of them. or is this why all the alignment people are quitting? they know that it can be jailbroken easily but fixing it is not a high priority for management. ",
  "transcript_chars": 2675,
  "ingested_at": "2026-09-06T13:30:17.485018+00:00",
  "source": "reddit",
  "yt_meta": {
    "score": 109,
    "upvote_ratio": 0.93,
    "num_comments": 22,
    "author": "Asleep-Requirement13",
    "is_self": true
  }
}