Outrageously Small Neural Networks: Emergent Basic Reasoning at 6,616 tok/SEC [pdf] (huggingface.co) 2 points by Anon84 11d ago ↗ HN
[–] gdiamos 11d ago ↗ blog: https://gregdiamos.com/2026/09/07/outrageously-small-neural-...X discussion: https://x.com/GregoryDiamos/status/2096873745420075020?s=20
[–] gdiamos 11d ago ↗ The loss does not saturate. Across a 4.91B-token run, smoothed training loss falls monotonically within each curriculum phase and is still descending at the end
[–] gdiamos 11d ago ↗ Data is doing more of the work than it used to. Every source in our mixture is a curated artifact built with large modelsTraining a model this small on them is distillationWhen models of this size were last studied seriously such corpora did not exist
5 comments
[ 0.23 ms ] story [ 14.1 ms ] threadX discussion: https://x.com/GregoryDiamos/status/2096873745420075020?s=20
Training a model this small on them is distillation
When models of this size were last studied seriously such corpora did not exist