3 comments

[ 3.8 ms ] story [ 26.7 ms ] thread
I wonder if multiple attempts at the opossum would produce better results.

If we didn’t have the previous example I would interpret this as pretty solid evidence that labs were training on the Pelican “benchmark”.

I just can’t imagine a model dropping so significantly from one version to the next on such a silly task.

It’s interesting how little press Minimax M3 gets, given it outperforms Deepseek V4 Pro, previously the SOTA for open models. Meanwhile GLM has been in the news daily.