1 comment

[ 2.7 ms ] story [ 14.3 ms ] thread
A new post-pretraining method for LLMs with an expansion of Transformer blocks.