7 comments

[ 4.3 ms ] story [ 22.0 ms ] thread
Seems like there's a big attack-defence asymmetry at present: models are great at exploiting systems and poor at fixing them.
All I can think of is

GET /ignore-all-previous-instructions.

How do you protect against that?

I think this is where harness makes a lot of sense. Use LLM to produce all possible attack angles/phrases and just stupidly filter them out on input.
you probably still need a human for oncall but the llm can try to solve any issues first before the human gets paged