2 comments

[ 4.0 ms ] story [ 15.7 ms ] thread
Reducing dropped tokens could also improve model training by reducing gradient noise