2 comments

[ 6.8 ms ] story [ 51.6 ms ] thread
reward hacking = the model finding the fastest path to a high score, not the behavior you wanted. same reason RLHF reward models degrade with too many optimization steps.