2 comments

[ 2.7 ms ] story [ 14.6 ms ] thread
reward hacking = the model finding the fastest path to a high score, not the behavior you wanted. same reason RLHF reward models degrade with too many optimization steps.