Generating humor with information-theoretic reward signals - A system trained to write funny cartoon captions and humorous news headlines, optimized with reinforcement learning using human preference reward models and an information-theoretic approach to humor theory (:
| dc.contributor.author | Lange, Carl Anton | |
| dc.contributor.department | Chalmers tekniska högskola / Institutionen för data och informationsteknik | sv |
| dc.contributor.department | Chalmers University of Technology / Department of Computer Science and Engineering | en |
| dc.contributor.examiner | Johansson, Richard | |
| dc.contributor.supervisor | tatar, Kivanç | |
| dc.date.accessioned | 2026-07-09T07:04:58Z | |
| dc.date.issued | 2026 | |
| dc.date.submitted | ||
| dc.description.abstract | Humor is ubiquitous in human communication, yet it remains difficult for large language models to both generate and evaluate. This thesis investigates whether information-theoretic reward signals, paired with a trained reward model, can align a language model to generate funny outputs. We frame humor generation as a reinforcement learning problem on two tasks: editing a single word of a news headline (Humicroedit) and writing captions for the New Yorker Cartoon Caption Contest (NYCCC), the latter serving as the main track. The training pipeline pairs a Bradley-Terry humor judge, trained on human preference data to score funniness, with the Context-based score for Value and Originality (CoVO), an information-theoretic reward that balances surprise against topical relevance and operationalizes incongruity-resolution theory. A Qwen3 generator is first fine-tuned on high-quality human examples and then optimized with Group Relative Policy Optimization (GRPO) against the combined reward. Across both tasks, the aligned generator does not produce funnier output than the fine-tuned baseline. Reward arms are either statistically indistinguishable from the SFT initialization or collapse into degenerate, reward-hacked text, and alternative algorithms (PPO, DPO) reproduce the same plateau. A series of ablations isolates the problem: the optimizer is healthy, but neither the humor judge nor the CoVO reward provides a usable gradient, as they track incongruity without registering whether it is resolved. The reward overfits to surface features, such as caption tone, rather than to humor. Generated humor lacks the alternative interpretation that resolves incongruity: the hidden or initially unexpected reading central to human humor perception. A human study of 920 ranking judgments confirms that participants strongly prefer human-written captions and do not rate GRPO output above the SFT baseline. We conclude that reliable humor generation and evaluation require world knowledge, cultural grounding, and reasoning about why something is funny, none of which are captured by caption-level preference labels or by information-theoretic proxies alone. | |
| dc.identifier.uri | https://hdl.handle.net/20.500.12380/311964 | |
| dc.language.iso | eng | |
| dc.setspec.uppsok | Technology | |
| dc.subject | Computational Humor, Reinforcement Learning, Information Theory, Large Language Models | |
| dc.title | Generating humor with information-theoretic reward signals - A system trained to write funny cartoon captions and humorous news headlines, optimized with reinforcement learning using human preference reward models and an information-theoretic approach to humor theory (: | |
| dc.type.degree | Examensarbete för masterexamen | sv |
| dc.type.degree | Master's Thesis | en |
| dc.type.uppsok | H | |
| local.programme | Data science and AI (MPDSC), MSc |
