Sample-Efficient Online Learning in LM Agents via Hindsight Trajectory Rewriting (arxiv.org) 2 points by djhu9 11mo ago ↗ HN
0 comments
[ 2.9 ms ] story [ 10.5 ms ] threadNo comments yet.