8 comments

[ 0.27 ms ] story [ 2.6 ms ] thread
Do these prompt injections work in the places this test placed them? (tool output)

I have my agents read other instruction files and they don't seem to get affected by the instructions found after a read/bash tool call. Curious if any analysis has been done to see if older prompt injection data sets are even effective anymore.

The whole thing looks heavily agent generated, my trust in them is not very high, how has this been validated or verified by a human?

Very cool. Did you try any majority vote or some other kind of technique to combine several of them and maybe achieve better results ?