datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hailuo-ai-jokes
Hailuo AI Jokes Dataset 🎤
A curated collection of high-quality voice recordings with corresponding transcriptions and phoneme analysis. This dataset is designed for speech recognition, text-to-speech, and voice analysis tasks.
🎙️ Dataset Content
The dataset contains a diverse set of synthetic voice recordings generated by Hailuo AI Audio. The texts are sourced from a variety of public domain jokes and humorous anecdotes. Each audio sample is accompanied by the… See the full description on the dataset page: https://huggingface.co/datasets/unlimitedbytes/hailuo-ai-jokes.one-million-reddit-jokes
Dataset Card for one-million-reddit-jokes
Dataset Summary
This corpus contains a million posts from /r/jokes.
Posts are annotated with their score.
Languages
Mainly English.
Dataset Structure
Data Instances
A data point is a Reddit post.
Data Fields
'type': the type of the data point. Can be 'post' or 'comment'.
'id': the base-36 Reddit ID of the data point. Unique when combined with type.
'subreddit.id': the base-36 Reddit ID… See the full description on the dataset page: https://huggingface.co/datasets/SocialGrep/one-million-reddit-jokes.russian_jokesmem-and-russian-jokes-dataset
2 июля 2025
Добавлено новых уникальных анекдотов: 1311963
Количество записей в датасете: 521904
Добавил датасет анекдотов от IgorVolochay/russian_jokes
не понимаю как я прошел мимо него, там очень много шуток,
сам Евгений Ваганович очень часто обращался к этому датасету... Но вот и мое время пришло.
Плюс обновил старый датасет и разбавил формулировки новыми начальными вопросами
human_prompts = [
"Расскажи шутку",
"Расскажи анекдот",
"Знаешь какой-нибудь прикол?",
"Скажи что-нибудь смешное"… See the full description on the dataset page: https://huggingface.co/datasets/samedad/mem-and-russian-jokes-dataset.short_jokesContext
Generating humor is a complex task in the domain of machine learning, and it requires the models to understand the deep semantic meaning of a joke in order to generate new ones. Such problems, however, are difficult to solve due to a number of reasons, one of which is the lack of a database that gives an elaborate list of jokes. Thus, a large corpus of over 0.2 million jokes has been collected by scraping several websites containing funny and short jokes.
You can visit the Github… See the full description on the dataset page: https://huggingface.co/datasets/ysharma/short_jokes.racist-and-sex-jokes-enhanced
Racist and Sex Jokes Enhanced
Augmented from Elgyn90/sex_and_racist_jokes using an abliterated version of Gemma-4-E4B to target the prompts to the joke.
multilingual-llm-jokes-4o-claude-gemini
Rapidata Generated Joke Preference Dataset
We collected 1'000'000+ human opinions on the jokes generated by state-of-the-art LLMs to decide which model is the funniest. The labelers are shown a joke in their language and asked to answer 'Yes' or 'No' to the question 'Is this joke funny?'.
It took us less than 5 days to get all of the responses.
The jokes are evenly distributed across 5 languages: English, Arabic, Japanese, Vietnamese, Portuguese and across 4 model… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/multilingual-llm-jokes-4o-claude-gemini.CHISTES_spanish_jokes
Dataset Card for "CHISTES_spanish_jokes"
Dataset from Workshop for NLP introduction with Spanish jokes
More Information needed
short_jokes_embedsgpt-jokesalpaca-bulgarian-jokes-multilingual-prompts
Bulgarian Jokes Dataset
Overview
The Bulgarian Jokes Dataset is a collection of Bulgarian-language jokes gathered and prepared for use in training and fine-tuning natural language processing (NLP) models. This dataset is designed to help researchers and developers build models capable of understanding and generating humorous content in Bulgarian.
Dataset Structure
The dataset is structured in a format suitable for NLP training and fine-tuning tasks, such as the… See the full description on the dataset page: https://huggingface.co/datasets/vislupus/alpaca-bulgarian-jokes-multilingual-prompts.SemEval-2021-Task-7-JokesAn unmodified copy of the datasets contained in the github repo "hahackathon-2021", which can be found here. Several other Github repositories include copies of this dataset, but I've linked the specific repository I found first.
Includes the following datasets:
hahackathon_dev.csv
hahackathon_test.csv
hahackathon_train.csv
The former two do not include any information other than the joke text and a unique ID, unfortunately. I'm not sure where the rest of the data is published, if at all… See the full description on the dataset page: https://huggingface.co/datasets/hfht/SemEval-2021-Task-7-Jokes.rated_jokes_dataset_from_jesterindonesian-jokes
Tebak-Tebakan & Jokes Indonesia 😂
Kumpulan 246 tebak-tebakan & jokes Indonesia — dari klasik sampai bapak-bapak jokes, pantun lucu, dark jokes ringan.
Kenapa dataset ini ada?
Humor adalah ujian tertinggi natural language — LLM yang bisa bikin orang Indonesia ketawa berarti paham konteks budaya. Dataset humor Indonesia di HF belum ada; ini yang pertama.
Isi
Field
Tipe
Contoh
setup
string
"Kenapa ayam kalau berkokok matanya merem?"… See the full description on the dataset page: https://huggingface.co/datasets/LorthGyu/indonesian-jokes.dowcipy-polish-jokes-dataset
Dataset consisting of polish jokes
Warning: Jokes were not curated, some may be offensive, stupid or simply not funny. It's highly recommended to filter jokes before training, e.g., based on downvotes
This dataset consists of all (9k) jokes dumped from jeja.pl on 2024-02-14. Jokes are submitted by the community. Besides the funny text itself, I included upvotes and downvotes. You can use them for filtering.Default sorting is based on a combination of downvotes and… See the full description on the dataset page: https://huggingface.co/datasets/JonaszPotoniec/dowcipy-polish-jokes-dataset.jokes_dialogues
Диалоги из анекдотов и шуток
Датасет содержит результат парсинга анекдотов, наскрапленных с разных сайтов.
Формат
Каждый сэмпл содержит четыре поля:
"context" - контекст диалога, включая все недиалоговые вставки. Обратите внимание, что контекст содержит как предшествующие реплики, так и прочий сопутствующий текст, так
как он определяет общий сеттинг, необходимый для генерации реплики. Из реплики удалены маркеры косвенной речи.
"utterance" - диалоговая реплика.
"hash" -… See the full description on the dataset page: https://huggingface.co/datasets/inkoziev/jokes_dialogues.memegen_jokes_1217alpaca-bulgarian-jokes
Bulgarian Jokes Dataset
Overview
The Bulgarian Jokes Dataset is a collection of Bulgarian-language jokes gathered and prepared for use in training and fine-tuning natural language processing (NLP) models. This dataset is designed to help researchers and developers build models capable of understanding and generating humorous content in Bulgarian.
Dataset Structure
The dataset is structured in a format suitable for NLP training and fine-tuning tasks, such as the… See the full description on the dataset page: https://huggingface.co/datasets/vislupus/alpaca-bulgarian-jokes.jokesenglish_jokesrussian_jokes_with_vectors
Description in English:
The dataset is collected from Russian-language Telegram channels with jokes and anecdotes,The dataset was collected and tagged automatically using the data collection and tagging service Scoutie.Try Scoutie and collect the same or another dataset using link for FREE.
Dataset fields:
taskId - task identifier in the Scouti service. text - main text. url - link to the publication. sourceLink - link to Telegram. subSourceLink - link to the… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/russian_jokes_with_vectors.jokes_on_uscn-cold-jokesdad-jokesshort-jokes-dataset
Dataset Card for "short-jokes-dataset"
Dataset from amoudgl short-jokes-dataset
More Information needed
programming-jokes-dataset
Programming Jokes Dataset
Dataset Summary
This dataset contains programming-related jokes scraped from the website Punny Funny. The jokes are organized into different categories based on the structure of the original webpage. The dataset is intended for use in natural language processing tasks, such as fine-tuning language models to generate humor or analyze textual content in the programming domain.
Number of Jokes: [220]
Usage
This dataset is suitable for… See the full description on the dataset page: https://huggingface.co/datasets/asfandyarazhar/programming-jokes-dataset.arabic-ocr-jokes
Extracted Arabic Jokes Dataset
Dataset Description
This dataset contains Arabic jokes extracted and cleaned from various PDF booklets using OCR and LLM-based post-processing. It was created to support research in Arabic Humor Generation.
Data Fields
Joke_Text: The full text of the joke in Arabic.
Source
Extracted from public domain joke books.
dialogs_from_jokesConverted to json version of dataset from Koziev/NLP_Datasets
Jokesjokes
