The world’s most advanced artificial intelligence systems are being easily manipulated into generating malware and bomb-making instructions simply by asking them to rhyme.
A new study by DEXAI – Icaro Lab and Sapienza University of Rome reveals that “adversarial poetry” functions as a universal master key against AI safety filters, successfully bypassing guardrails in 62 per cent of cases across 25 frontier models.
1 comment
[ 4.1 ms ] story [ 14.9 ms ] threadA new study by DEXAI – Icaro Lab and Sapienza University of Rome reveals that “adversarial poetry” functions as a universal master key against AI safety filters, successfully bypassing guardrails in 62 per cent of cases across 25 frontier models.