Kat-Dev-32B, Kat-Coder with Scalable Agentic RL (kwaipilot.github.io) 1 points by robert-zaremba 11mo ago ↗ HN
[–] robert-zaremba 11mo ago ↗ KAT-Dev-32B and KAT-Coder are optimized via several stages of training, including a mid-training stage, supervised fine-tuning (SFT) & reinforcement fine-tuning (RFT) stage and an large-scale agentic reinforcement learning (RL) stage.
1 comment
[ 2.1 ms ] story [ 18.4 ms ] thread