Show HN: A benchmark where LLMs make memes from current news (memebench.net) 4 points by max-azendorf 4mo ago ↗ HN
[–] vintagedave 4mo ago ↗ I was a bit skeptical but this is actually really neat. Sometimes they nail it. The situations where both are bad are common, and seeing which LLM produced them lines up quite well with how good I think the same models are.
3 comments
[ 3.2 ms ] story [ 20.3 ms ] thread