2 comments

[ 2.9 ms ] story [ 17.9 ms ] thread
seems like something pytorch maintainers would want to know about and fix asap...
Article title: Bugs in LLM Training - Gradient Accumulation Fix