Tapered Off-Policy Reinforce: Stable and Efficient RL for LLMs (arxiv.org) 2 points by pama 1y ago ↗ HN
0 comments
[ 0.17 ms ] story [ 206 ms ] threadNo comments yet.