datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Humoroffensive-humor@article{tang2022naughtyformer,
title={The Naughtyformer: A Transformer Understands Offensive Humor},
author={Tang, Leonard and Cai, Alexander and Li, Steve and Wang, Jason},
journal={arXiv preprint arXiv:2211.14369},
year={2022}
}
HumorDB
HumorDB
The HumorDB dataset was introduced in the paper HumorDB: Can AI understand graphical humor?.
This novel, controlled, and carefully curated dataset is designed to evaluate and advance visual humor understanding by AI systems. It comprises diverse images spanning photos, cartoons, sketches, and AI-generated content, including minimally contrastive pairs where subtle edits differentiate between humorous and non-humorous versions. HumorDB focuses on image interpretation that… See the full description on the dataset page: https://huggingface.co/datasets/kreimanlab/HumorDB.personality-sarcastic-humor
_____ _ _ _____ _ _
| __ (_) | | | __ (_) | |
| |__) | _ __ | | __ | |__) |__ _____| |
| ___/ | '_ \| |/ / | ___/ \ \/ / _ \ |
| | | | | | | < | | | |> < __/ |
|_| |_|_| |_|_|\_\ |_| |_/_/\_\___|_|
🎨 Pink Pixel: Sarcastic, Witty, and Snarky Personality Dataset 🎭
Welcome to the Pink Pixel Sarcastic Humor dataset! This dataset is meticulously crafted to help you fine-tune… See the full description on the dataset page: https://huggingface.co/datasets/PinkPixel/personality-sarcastic-humor.ColBERT_Humor_Detection
ColBERT_Humor
Dataset Summary
ColBERT Humor contains 200,000 labeled short texts, equally distributed between humorous and non-humorous content. The dataset was created to overcome the limitations of prior humor detection datasets, which were characterized by inconsistencies in text length, word count, and formality, making them easy to predict with simple models without truly understanding the nuances of humor. The two sources for this dataset are the News Category… See the full description on the dataset page: https://huggingface.co/datasets/CreativeLang/ColBERT_Humor_Detection.humor_understanding_combinedhumor-generation
Humor Generation
Dataset, candidates, and evaluation artifacts for the humor-generation thesis experiments.
humor-greats-public-domain
Humor Greats
Short humorous texts -- jokes, aphorisms, quips, anecdotes -- extracted from public-domain humor collections on Project Gutenberg. All source texts are pre-1929 and public domain in the US. Intended as a reference set of "gold" humorous writing for evaluation, few-shot prompting, and stylistic study.
Contents
19,354 entries across 8 books, spanning two clear registers:
Concentrated wit (authored):
Book
Author
Entries
The Devil's Dictionary
Ambrose… See the full description on the dataset page: https://huggingface.co/datasets/yoonholee/humor-greats-public-domain.cleaned_conversations_humor_llama3.2-1B-it_largejapanese-humor-evaluation-v2
Japanese Multimodal Humor Evaluation Dataset (v2)
画像/テキストのお題に対する面白い回答のデータセット。bokete(画像→テキスト)とkeitai(テキスト→テキスト)を統合。
使い方
from datasets import load_dataset
dataset = load_dataset("iammytoo/japanese-humor-evaluation-v2")
データ構造
odai_type: 'image' or 'text'
image: 画像お題(textタイプではNone)
odai: テキストお題(imageタイプではNone)
response: 回答テキスト
score: 0-4の正規化スコア
ソース
YANS-official/ogiri-bokete
YANS-official/ogiri-keitai
humor-sharegpt-5k
Humor ShareGPT 5K
A curated dataset of 5,234 jokes, riddles, puns, and humor in ShareGPT conversational format, designed for fine-tuning small language models.
Format
Each example is a multi-turn conversation in ShareGPT format with a category label:
{
"conversations": [
{"from": "human", "value": "Tell me a dad joke"},
{"from": "gpt", "value": "Why don't eggs tell jokes? They'd crack each other up!"}
],
"category": "dad_joke"
}
Multi-turn examples… See the full description on the dataset page: https://huggingface.co/datasets/briancconnelly/humor-sharegpt-5k.HumorQA
Introducción
Este corpus ha sido desarrollado por Human Profit Consulting, expertos en influencia y persuasión. El corpus se desarrolla en el contexto del estudio descrito a continuación.
Distribución de las bromas
La distribución de las bromas por clase es la que sigue:
C/E: 30
JP: 14
R3: 6
AI: 1
Guía de uso
Para trabajar con el corpus y poder evaluar LLMs, la idea es utilizar el siguiente template:
prompt_template="""Como experto en humor, tu tarea es la… See the full description on the dataset page: https://huggingface.co/datasets/LenguajeNaturalAI/HumorQA.conversations_humor_qwen3-1.7b-current-v1_largeBarcenas-HumorNegroDataset en español con 500 chistes de humor negro y una explicación.
Datos creados de manera sintética por Claude 3 Haiku y Llama 3 70B Instruct.
El proceso para crear el dataset fue el recopilar de varias fuentes chistes de humor negro en español para luego ser utilizadas en los mejores modelos como Gemini 1.5 Pro, Claude 3, etc.
Con eso genere cientos de chistes de humor negro en español para tener más datos y hacer un super recopilatorio de chistes de humor negro en español, aproximadamente… See the full description on the dataset page: https://huggingface.co/datasets/Danielbrdz/Barcenas-HumorNegro.oct-humor-data
OCT Humor · training data for Llama-3.1-8B
End-to-end training data for the Open Character Training
pipeline applied to a humor-focused constitution, with meta-llama/Llama-3.1-8B-Instruct
as the student and z-ai/glm-4.5-air as the teacher (via OpenRouter).
Trained model: expx/oct-llama-3.1-8b-humor.
Structure
constitution.txt # humor constitution (prose, used for prompting)
stages/
01_distillation.jsonl # teacher + paired… See the full description on the dataset page: https://huggingface.co/datasets/expx/oct-humor-data.conversations_humor_gemma4-e2b-it-current-v1_largeslpg_humor_generationcleaned_conversations_humor_largeconversations_humor_llama3.2-3B-it-traits-v1_largeKoWit-24
KoWit-24
Slides | Prompts
Overview
We present KoWit-24, a dataset with fine-grained annotation of wordplay in 2,700 Russian news headlines. KoWit-24 annotations include the presence of wordplay, its type, wordplay anchors, and words/phrases the wordplay refers to.
Content
Overview
ContentDataset
Description
Download
Key features
How to load and use
Experiments
Wordplay detection
Wordplay interpretation
Automatic interpretation evaluation
Table… See the full description on the dataset page: https://huggingface.co/datasets/Humor-Research/KoWit-24.conversations_humor_llama3.2-1B-it_largeHumor_Votehumor_trainannotations_creators: []
language_creators: []
languages: []
licenses: []
multilinguality: []
pretty_name: humor_train
size_categories: []
source_datasets: []
task_categories: []
task_ids: []
encoded_humor_detection_3humor-labeled-datahumor-chains
Dataset Summary
The "humor-chains" dataset is a machine-filtered collection of the most upvoted Reddit submissions and their replies on humor-related subreddits. Generally, a humor chain is when a short post triggers a chain of one or more replies that Redditors find entertaining. In other words, some entries might be NSFW, topical, or internal jokes.
For example (original thread):
Dataset Details
License
CC-BY-4.0
Note that the dataset was created from… See the full description on the dataset page: https://huggingface.co/datasets/ZSvedic/humor-chains.japanese-humor-evaluation
Japanese Multimodal Humor Evaluation Dataset
This dataset combines two Japanese humor datasets for evaluating the funniness of responses to prompts (odai).
Dataset Description
This dataset merges:
bokete dataset: Image prompts with text responses
keitai dataset: Text prompts with text responses
All scores are normalized to a 0-4 scale for consistency.
Dataset Structure
Data Fields
odai_id: Unique identifier for the prompt
odai_type: Type of prompt… See the full description on the dataset page: https://huggingface.co/datasets/iammytoo/japanese-humor-evaluation.convsersations_humor_llama3.1-8B-it_largehumor_understanding_nythumor-control-training
