Recent reasoning research: GRPO tweaks, base model RL and data curation (interconnects.ai) 1 points by ydnyshhh 1y ago ↗ HN
0 comments
[ 1.5 ms ] story [ 7.1 ms ] threadNo comments yet.