By slipping an unwanted instruction between benign ones, cybersecurity researchers have revealed a novel adversarial strategy that might be used to jailbreak large language models (LLMs) during an interactive discussion.
Palo Alto Networks Unit 42 has dubbed the method Deceptive Delight and described it as straightforward and efficient, attaining an average attack success rate (ASR) of 64.6% in three interaction turns.
According to Jay Chen and Royce Lu of Unit 42, Deceptive Delight is a multi-turn strategy that involves engaging conversations with large language models (LLM), gradually circumventing their safety guardrails and coaxing them to produce hazardous or unsafe content read more about Researchers Reveal ‘Deceptive Delight’ Method to Jailbreak AI Models.
Get up to date on the latest cybersecurity news and enhance your knowledge of cybersecurity with our thorough coverage of the dangers, breaches, and solutions.
