There was a paper [1] that offered as a defense against abliteration what amounts to just longer and more nuanced refusals. The results seemed promising.
I don't follow this closely enough to know why the latest open models haven't incorporated the research...
1 comment
[ 0.23 ms ] story [ 12.2 ms ] threadI don't follow this closely enough to know why the latest open models haven't incorporated the research...
[1] https://arxiv.org/abs/2505.19056