Meta-Rewarding Language Models:Self-Improving Alignment with LLM-as-a-Meta-Judge (arxiv.org) 2 points by sssummer 2y ago ↗ HN
0 comments
[ 2.7 ms ] story [ 7.2 ms ] threadNo comments yet.