No Train No Gain:Revisiting Efficient Training Algrthm for Transformer-BasedLM (arxiv.org) 11 points by froster 3y ago ↗ HN
[–] froster 3y ago ↗ Recent paper highlights the difficulty of creating a new optimizer as drop-in replacement. Sophia and Lion were recently proposed as superior alternatives to Adam, but appeared worse in an independent eval
1 comment
[ 7.8 ms ] story [ 13.5 ms ] thread