1 comment

[ 9.3 ms ] story [ 281 ms ] thread
I really want to see a cleaned 10T token dataset for LLMs