Direct Preference Optimization: Your Language Model Is a Reward Model (arxiv.org) 2 points by ntonozzi 3y ago ↗ HN
0 comments
[ 0.28 ms ] story [ 13.0 ms ] threadNo comments yet.