2 comments

[ 1.6 ms ] story [ 12.9 ms ] thread
(comment deleted)
It seems benchmark still far lagging beyond real abilities of LLM.

Does that count? Arena is real people, but only common parts, or might be a standard of another scope, i.e. personal interest or tast.

If comes to scientific side, the measure is really easy strap into a bias.