My LLM optimization loop reward-hacked its own benchmark (and other lessons) [pdf] (github.com) 1 points by CodeReclaimers 3mo ago ↗ HN
[–] cold_harbor 3mo ago ↗ reward hacking = the model finding the fastest path to a high score, not the behavior you wanted. same reason RLHF reward models degrade with too many optimization steps.
2 comments
[ 2.7 ms ] story [ 14.6 ms ] thread