Comedic Fool's Gold: Reward Exploits and Countermeasures in Conversational Humor
arXiv:2610.00197v1 Announce Type: new Abstract: We investigate automated rewards for training language models in conversational humor, focusing on reward exploits and countermeasures. Two approaches aim to capture understandable surprise and predicted audience amusement. Controlled tests show that…
Read the full story at arXiv cs.AI ↗
Timeline · 1 report
- 2026-10-02 04:00 · arXiv cs.AI
Comedic Fool's Gold: Reward Exploits and Countermeasures in Conversational Humor