1 comment

[ 3.6 ms ] story [ 35.9 ms ] thread
Notable it's a hybrid model, 7 linear attention layers for every traditional O(n^2) attention layer. They claim this allows it to outperform other LLMs (including Gemini) on long context benchmarks.