5 comments

[ 0.23 ms ] story [ 14.1 ms ] thread
The loss does not saturate. Across a 4.91B-token run, smoothed training loss falls monotonically within each curriculum phase and is still descending at the end
Data is doing more of the work than it used to. Every source in our mixture is a curated artifact built with large models

Training a model this small on them is distillation

When models of this size were last studied seriously such corpora did not exist