1 comment

[ 0.25 ms ] story [ 14.2 ms ] thread
Switching from imitative training of LLMs to self-competitive reinforcement learning allows them to surpass human level in general language-domain tasks.