Training a Rust 1.5B Coder LM with Reinforcement Learning (GRPO) (ghost.oxen.ai) 3 points by curiousinspo 1y ago ↗ HN
1 comment
[ 3.6 ms ] story [ 10.1 ms ] thread