Faulty Nvidia H100 GPUs and HBM3 memory caused failures during Llama 3 training (tomshardware.com) 4 points by surfingdino 2y ago ↗ HN
[–] surfingdino 2y ago ↗ Full title: "Faulty Nvidia H100 GPUs and HBM3 memory caused half of failures during LLama 3 training — one failure every three hours for Meta's 16,384 GPU training cluster"
1 comment
[ 3.0 ms ] story [ 14.2 ms ] thread