Large models don't just tolerate noisy or nominally "low-quality" web data but this paper suggests that they benefit from the distributional entropy in it. Over-curated datasets artificially compress the representation manifold, starving the model of the edge cases needed for robust generalization.
1 comment
[ 0.22 ms ] story [ 5.7 ms ] thread