building a humor benchmark for LLMs: someone told me my benchmark's best result was just memory, so i ran his test
hey guys! i've been building lolbench, a benchmark for whether LLMs actually understand humor. models do three things: explain why a joke works, write jokes on a shared setup, and rank jokes by human preference. everything is auto-judged by models from other labs, plus a blind human vote booth the…
Read the full story at r/artificial ↗
Timeline · 1 report
- 2026-09-27 20:35 · r/artificial
building a humor benchmark for LLMs: someone told me my benchmark's best result was just memory, so i ran his test