1 min readfrom Machine Learning

GPT-6 reportedly jailbroken within 24 hours using an extended Task-in-Prompt (TIP) attack [N]

A researcher has reported a jailbreak of GPT-6 Astra within a day after release.

The attack is described as combination of TIP (Task-in-Prompt) attack from ACL 2025 paper with four other unnamed techniques.

TIP attacks exploit the model’s reasoning/instruction-following behaviour by hidding the harmful objective inside another task, like solving a cipher or executing a Python code. For GPT-6, the researcher says the original minimal TIP attack was no longer sufficient and had to be reworked.

They have reportedly disclosed the details privately to OpenAI rather than publishing the jailbreak.

The same researcher reported jailbreaking GPT-5 within an hour of its release a year ago.

Source: screenshot/post from the researcher; their ACL 2025 TIP paper linked in the original post.

submitted by /u/Asleep-Requirement13
[link] [comments]

Want to read more?

Check out the full article on the original site

View original article

Tagged with

#GPT-6
#Jailbreak
#TIP attack
#Task-in-Prompt
#GPT-5
#ACL 2025
#Researcher
#LLM
#Reasoning
#OpenAI
#Astra
#Instruction-following
#Model Vulnerability
#AI Alignment
#AI Security
#Harmful Objective
#Cipher
#Python Code
#Prompt Engineering
#Machine Learning