This is cool. I've always thought LLM humor understanding is an under researched, but important, topic, especially w.r.t "is AGI here yet?".
I voted on a few "Which one is funnier?" choices, but honestly none of them were funny. The em dashes in every joke already set the "another LLM slop" mood. I would suggest re-writing the prompts to drop dashes. Another suggestion: slider instead of A/B single choice. Sometimes I leaned towards one choice, but it still wasn't that funny to select it; a more nuanced scale would have helped there.
Thanks for building this. From the first glance, we humans are safe from machine AGI judging by joke quality today.
3 comments
[ 0.24 ms ] story [ 13.8 ms ] threadLLMs take three tests:
explain why jokes work (or don't), write jokes under shared premises or predict which jokes humans prefer
The finding so far that surprised me: every model aces explaining real jokes (95%+) but drops hard on explaining why a failed joke fails (81–92%).
happy to answer anything about the eval system and open to any sort of feedback!
I voted on a few "Which one is funnier?" choices, but honestly none of them were funny. The em dashes in every joke already set the "another LLM slop" mood. I would suggest re-writing the prompts to drop dashes. Another suggestion: slider instead of A/B single choice. Sometimes I leaned towards one choice, but it still wasn't that funny to select it; a more nuanced scale would have helped there.
Thanks for building this. From the first glance, we humans are safe from machine AGI judging by joke quality today.