disi-unibo-nlp/Phunny
Phunny: A Humor-Based QA Benchmark for Evaluating LLM Generalization Welcome to Phunny, a humor-based question answering (QA) benchmark designed to evaluate the reasoning and generalization abilities of large language models (LLMs) through structured puns. This repository accompanies our ACL 2025 main track paper:"What do you call a dog that is incontrovertibly true? Dogma: Testing LLM Generalization through Humor" To reproduce our experiments: Code available on GitHub… See the full description on the dataset page: https://huggingface.co/datasets/disi-unibo-nlp/Phunny.
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
Upload dataset
Upload dataset
Upload dataset
Upload dataset
Upload dataset
Upload dataset
Upload dataset
Upload dataset
initial commit
