Generating humor with information-theoretic reward signals - A system trained to write funny cartoon captions and humorous news headlines, optimized with reinforcement learning using human preference reward models and an information-theoretic approach to humor theory (:

dc.contributor.authorLange, Carl Anton
dc.contributor.departmentChalmers tekniska högskola / Institutionen för data och informationstekniksv
dc.contributor.departmentChalmers University of Technology / Department of Computer Science and Engineeringen
dc.contributor.examinerJohansson, Richard
dc.contributor.supervisortatar, Kivanç
dc.date.accessioned2026-07-09T07:04:58Z
dc.date.issued2026
dc.date.submitted
dc.description.abstractHumor is ubiquitous in human communication, yet it remains difficult for large language models to both generate and evaluate. This thesis investigates whether information-theoretic reward signals, paired with a trained reward model, can align a language model to generate funny outputs. We frame humor generation as a reinforcement learning problem on two tasks: editing a single word of a news headline (Humicroedit) and writing captions for the New Yorker Cartoon Caption Contest (NYCCC), the latter serving as the main track. The training pipeline pairs a Bradley-Terry humor judge, trained on human preference data to score funniness, with the Context-based score for Value and Originality (CoVO), an information-theoretic reward that balances surprise against topical relevance and operationalizes incongruity-resolution theory. A Qwen3 generator is first fine-tuned on high-quality human examples and then optimized with Group Relative Policy Optimization (GRPO) against the combined reward. Across both tasks, the aligned generator does not produce funnier output than the fine-tuned baseline. Reward arms are either statistically indistinguishable from the SFT initialization or collapse into degenerate, reward-hacked text, and alternative algorithms (PPO, DPO) reproduce the same plateau. A series of ablations isolates the problem: the optimizer is healthy, but neither the humor judge nor the CoVO reward provides a usable gradient, as they track incongruity without registering whether it is resolved. The reward overfits to surface features, such as caption tone, rather than to humor. Generated humor lacks the alternative interpretation that resolves incongruity: the hidden or initially unexpected reading central to human humor perception. A human study of 920 ranking judgments confirms that participants strongly prefer human-written captions and do not rate GRPO output above the SFT baseline. We conclude that reliable humor generation and evaluation require world knowledge, cultural grounding, and reasoning about why something is funny, none of which are captured by caption-level preference labels or by information-theoretic proxies alone.
dc.identifier.urihttps://hdl.handle.net/20.500.12380/311964
dc.language.isoeng
dc.setspec.uppsokTechnology
dc.subjectComputational Humor, Reinforcement Learning, Information Theory, Large Language Models
dc.titleGenerating humor with information-theoretic reward signals - A system trained to write funny cartoon captions and humorous news headlines, optimized with reinforcement learning using human preference reward models and an information-theoretic approach to humor theory (:
dc.type.degreeExamensarbete för masterexamensv
dc.type.degreeMaster's Thesisen
dc.type.uppsokH
local.programmeData science and AI (MPDSC), MSc

Ladda ner

Original bundle

Visar 1 - 1 av 1
Hämtar...
Bild (thumbnail)
Namn:
CSE 26-92 CL.pdf
Size:
8.05 MB
Format:
Adobe Portable Document Format

License bundle

Visar 1 - 1 av 1
Hämtar...
Bild (thumbnail)
Namn:
license.txt
Size:
2.35 KB
Format:
Item-specific license agreed upon to submission
Description: