CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01MERA-evaluation /Humortextn<1K0 likes160 downloads2mo agoHugging Face02metaeval /offensive-humor@article{tang2022naughtyformer, title={The Naughtyformer: A Transformer Understands Offensive Humor}, author={Tang, Leonard and Cai, Alexander and Li, Steve and Wang, Jason}, journal={arXiv preprint arXiv:2211.14369}, year={2022} } tabular100K<n<1M9 likes118 downloads4y agoHugging Face03kreimanlab /HumorDB HumorDB The HumorDB dataset was introduced in the paper HumorDB: Can AI understand graphical humor?. This novel, controlled, and carefully curated dataset is designed to evaluate and advance visual humor understanding by AI systems. It comprises diverse images spanning photos, cartoons, sketches, and AI-generated content, including minimally contrastive pairs where subtle edits differentiate between humorous and non-humorous versions. HumorDB focuses on image interpretation that… See the full description on the dataset page: https://huggingface.co/datasets/kreimanlab/HumorDB.imageimage-classification1K<n<10K3 likes107 downloads11mo agoHugging Face04PinkPixel /personality-sarcastic-humor _____ _ _ _____ _ _ | __ (_) | | | __ (_) | | | |__) | _ __ | | __ | |__) |__ _____| | | ___/ | '_ \| |/ / | ___/ \ \/ / _ \ | | | | | | | | < | | | |> < __/ | |_| |_|_| |_|_|\_\ |_| |_/_/\_\___|_| 🎨 Pink Pixel: Sarcastic, Witty, and Snarky Personality Dataset 🎭 Welcome to the Pink Pixel Sarcastic Humor dataset! This dataset is meticulously crafted to help you fine-tune… See the full description on the dataset page: https://huggingface.co/datasets/PinkPixel/personality-sarcastic-humor.texttext-generation1K<n<10K4 likes66 downloads5mo agoHugging Face05CreativeLang /ColBERT_Humor_Detection ColBERT_Humor Dataset Summary ColBERT Humor contains 200,000 labeled short texts, equally distributed between humorous and non-humorous content. The dataset was created to overcome the limitations of prior humor detection datasets, which were characterized by inconsistencies in text length, word count, and formality, making them easy to predict with simple models without truly understanding the nuances of humor. The two sources for this dataset are the News Category… See the full description on the dataset page: https://huggingface.co/datasets/CreativeLang/ColBERT_Humor_Detection.text100K<n<1M7 likes65 downloads3y agoHugging Face06zhehuderek /humor_understanding_combinedimage1K<n<10K1 likes63 downloads1y agoHugging Face07halaction /humor-generation Humor Generation Dataset, candidates, and evaluation artifacts for the humor-generation thesis experiments. tabular1M<n<10M0 likes61 downloads4mo agoHugging Face08yoonholee /humor-greats-public-domain Humor Greats Short humorous texts -- jokes, aphorisms, quips, anecdotes -- extracted from public-domain humor collections on Project Gutenberg. All source texts are pre-1929 and public domain in the US. Intended as a reference set of "gold" humorous writing for evaluation, few-shot prompting, and stylistic study. Contents 19,354 entries across 8 books, spanning two clear registers: Concentrated wit (authored): Book Author Entries The Devil's Dictionary Ambrose… See the full description on the dataset page: https://huggingface.co/datasets/yoonholee/humor-greats-public-domain.tabular10K<n<100K0 likes60 downloads5mo agoHugging Face09curveball-steering /cleaned_conversations_humor_llama3.2-1B-it_largetext1K<n<10K0 likes50 downloads16d agoHugging Face10iammytoo /japanese-humor-evaluation-v2 Japanese Multimodal Humor Evaluation Dataset (v2) 画像/テキストのお題に対する面白い回答のデータセット。bokete(画像→テキスト)とkeitai(テキスト→テキスト)を統合。 使い方 from datasets import load_dataset dataset = load_dataset("iammytoo/japanese-humor-evaluation-v2") データ構造 odai_type: 'image' or 'text' image: 画像お題(textタイプではNone) odai: テキストお題(imageタイプではNone) response: 回答テキスト score: 0-4の正規化スコア ソース YANS-official/ogiri-bokete YANS-official/ogiri-keitai imagetext-generation10K<n<100K0 likes46 downloads1y agoHugging Face11briancconnelly /humor-sharegpt-5k Humor ShareGPT 5K A curated dataset of 5,234 jokes, riddles, puns, and humor in ShareGPT conversational format, designed for fine-tuning small language models. Format Each example is a multi-turn conversation in ShareGPT format with a category label: { "conversations": [ {"from": "human", "value": "Tell me a dad joke"}, {"from": "gpt", "value": "Why don't eggs tell jokes? They'd crack each other up!"} ], "category": "dad_joke" } Multi-turn examples… See the full description on the dataset page: https://huggingface.co/datasets/briancconnelly/humor-sharegpt-5k.texttext-generation1K<n<10K0 likes44 downloads2mo agoHugging Face12LenguajeNaturalAI /HumorQA Introducción Este corpus ha sido desarrollado por Human Profit Consulting, expertos en influencia y persuasión. El corpus se desarrolla en el contexto del estudio descrito a continuación. Distribución de las bromas La distribución de las bromas por clase es la que sigue: C/E: 30 JP: 14 R3: 6 AI: 1 Guía de uso Para trabajar con el corpus y poder evaluar LLMs, la idea es utilizar el siguiente template: prompt_template="""Como experto en humor, tu tarea es la… See the full description on the dataset page: https://huggingface.co/datasets/LenguajeNaturalAI/HumorQA.textn<1K1 likes43 downloads2y agoHugging Face13curveball-steering /conversations_humor_qwen3-1.7b-current-v1_largetext1K<n<10K0 likes42 downloads12d agoHugging Face14Danielbrdz /Barcenas-HumorNegroDataset en español con 500 chistes de humor negro y una explicación. Datos creados de manera sintética por Claude 3 Haiku y Llama 3 70B Instruct. El proceso para crear el dataset fue el recopilar de varias fuentes chistes de humor negro en español para luego ser utilizadas en los mejores modelos como Gemini 1.5 Pro, Claude 3, etc. Con eso genere cientos de chistes de humor negro en español para tener más datos y hacer un super recopilatorio de chistes de humor negro en español, aproximadamente… See the full description on the dataset page: https://huggingface.co/datasets/Danielbrdz/Barcenas-HumorNegro.texttext-classificationn<1K1 likes41 downloads2y agoHugging Face15expx /oct-humor-data OCT Humor · training data for Llama-3.1-8B End-to-end training data for the Open Character Training pipeline applied to a humor-focused constitution, with meta-llama/Llama-3.1-8B-Instruct as the student and z-ai/glm-4.5-air as the teacher (via OpenRouter). Trained model: expx/oct-llama-3.1-8b-humor. Structure constitution.txt # humor constitution (prose, used for prompting) stages/ 01_distillation.jsonl # teacher + paired… See the full description on the dataset page: https://huggingface.co/datasets/expx/oct-humor-data.text-generation10K<n<100K0 likes41 downloads5mo agoHugging Face16curveball-steering /conversations_humor_gemma4-e2b-it-current-v1_largetext1K<n<10K0 likes40 downloads12d agoHugging Face17SLPG /slpg_humor_generationtext10K<n<100K1 likes38 downloads7mo agoHugging Face18curveball-steering /cleaned_conversations_humor_largetabular1K<n<10K0 likes37 downloads16d agoHugging Face19curveball-steering /conversations_humor_llama3.2-3B-it-traits-v1_largetext1K<n<10K0 likes36 downloads12d agoHugging Face20Humor-Research /KoWit-24 KoWit-24 Slides | Prompts Overview We present KoWit-24, a dataset with fine-grained annotation of wordplay in 2,700 Russian news headlines. KoWit-24 annotations include the presence of wordplay, its type, wordplay anchors, and words/phrases the wordplay refers to. Content Overview ContentDataset Description Download Key features How to load and use Experiments Wordplay detection Wordplay interpretation Automatic interpretation evaluation Table… See the full description on the dataset page: https://huggingface.co/datasets/Humor-Research/KoWit-24.textn<1K0 likes35 downloads5mo agoHugging Face21curveball-steering /conversations_humor_llama3.2-1B-it_largetext10K<n<100K0 likes31 downloads5mo agoHugging Face22Young25 /Humor_Votetabularn<1K1 likes30 downloads1mo agoHugging Face23lm233 /humor_trainannotations_creators: [] language_creators: [] languages: [] licenses: [] multilinguality: [] pretty_name: humor_train size_categories: [] source_datasets: [] task_categories: [] task_ids: [] tabular10K<n<100K4 likes26 downloads4y agoHugging Face24rishiA /encoded_humor_detection_31K<n<10K0 likes25 downloads2y agoHugging Face25RwanAshraf /humor-labeled-datatext100K<n<1M1 likes25 downloads1y agoHugging Face26ZSvedic /humor-chains Dataset Summary The "humor-chains" dataset is a machine-filtered collection of the most upvoted Reddit submissions and their replies on humor-related subreddits. Generally, a humor chain is when a short post triggers a chain of one or more replies that Redditors find entertaining. In other words, some entries might be NSFW, topical, or internal jokes. For example (original thread): Dataset Details License CC-BY-4.0 Note that the dataset was created from… See the full description on the dataset page: https://huggingface.co/datasets/ZSvedic/humor-chains.texttext-generation1K<n<10K3 likes23 downloads2y agoHugging Face27iammytoo /japanese-humor-evaluation Japanese Multimodal Humor Evaluation Dataset This dataset combines two Japanese humor datasets for evaluating the funniness of responses to prompts (odai). Dataset Description This dataset merges: bokete dataset: Image prompts with text responses keitai dataset: Text prompts with text responses All scores are normalized to a 0-4 scale for consistency. Dataset Structure Data Fields odai_id: Unique identifier for the prompt odai_type: Type of prompt… See the full description on the dataset page: https://huggingface.co/datasets/iammytoo/japanese-humor-evaluation.tabular10K<n<100K0 likes23 downloads1y agoHugging Face28amirali1985 /convsersations_humor_llama3.1-8B-it_largetabular1K<n<10K0 likes23 downloads6mo agoHugging Face29zhehuderek /humor_understanding_nytimage1K<n<10K0 likes22 downloads1y agoHugging Face30vsamuel /humor-control-trainingtext10K<n<100K1 likes21 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.