We used sparse autoencoders to explain LLM moderation flags of violent threats (variance.co) 6 points by karinemellata 1y ago ↗ HN
0 comments
[ 2.5 ms ] story [ 12.5 ms ] threadNo comments yet.