My LLM optimization loop reward-hacked its own benchmark (and other lessons) [pdf] (github.com) 1 points by CodeReclaimers 3mo ago ↗ HN
[–] cold_harbor 3mo ago ↗ reward hacking = the model finding the fastest path to a high score, not the behavior you wanted. same reason RLHF reward models degrade with too many optimization steps.
2 comments
[ 6.8 ms ] story [ 51.6 ms ] thread