Why so many updating in LLM so-called SOTA let remind of iPhone 4-x (medium.com) 2 points by qiuwu 1mo ago ↗ HN
[–] qiuwu 1mo ago ↗ And the ground truth benchmark of LLM is still lacking.HumanEval done, SWEBench done... Since it will lasting on and on, but task completion of LLM remain hard to be achieved esp. for high complexity task?Just investor and Big Three of LLM fool of ppl?
1 comment
[ 7.2 ms ] story [ 19.1 ms ] threadHumanEval done, SWEBench done... Since it will lasting on and on, but task completion of LLM remain hard to be achieved esp. for high complexity task?
Just investor and Big Three of LLM fool of ppl?