We tested whether prompt canaries and honeypots could detect AI attackers.
Prompt canaries were surprisingly unreliable. I suspect the defenses model providers have added against prompt injection also make models less susceptible to prompt canaries.
Honeypots worked much better, having roastable accounts that are otherwise useless seems to be a good catcher.
1 comment
[ 12.7 ms ] story [ 39.2 ms ] threadPrompt canaries were surprisingly unreliable. I suspect the defenses model providers have added against prompt injection also make models less susceptible to prompt canaries.
Honeypots worked much better, having roastable accounts that are otherwise useless seems to be a good catcher.