1 comment

[ 3.3 ms ] story [ 14.7 ms ] thread
I would love to see a purely mamba-based 120b model, and whether or not it outcompetes the open-weights OpenAI model.