Generalized on-policy distillation with reward extrapolation (arxiv.org) 3 points by fzliu 7mo ago ↗ HN
0 comments
[ 2.1 ms ] story [ 15.7 ms ] threadNo comments yet.