datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
one-million-reddit-jokes
Dataset Card for one-million-reddit-jokes
Dataset Summary
This corpus contains a million posts from /r/jokes.
Posts are annotated with their score.
Languages
Mainly English.
Dataset Structure
Data Instances
A data point is a Reddit post.
Data Fields
'type': the type of the data point. Can be 'post' or 'comment'.
'id': the base-36 Reddit ID of the data point. Unique when combined with type.
'subreddit.id': the base-36 Reddit ID… See the full description on the dataset page: https://huggingface.co/datasets/SocialGrep/one-million-reddit-jokes.mem-and-russian-jokes-dataset
2 июля 2025
Добавлено новых уникальных анекдотов: 1311963
Количество записей в датасете: 521904
Добавил датасет анекдотов от IgorVolochay/russian_jokes
не понимаю как я прошел мимо него, там очень много шуток,
сам Евгений Ваганович очень часто обращался к этому датасету... Но вот и мое время пришло.
Плюс обновил старый датасет и разбавил формулировки новыми начальными вопросами
human_prompts = [
"Расскажи шутку",
"Расскажи анекдот",
"Знаешь какой-нибудь прикол?",
"Скажи что-нибудь смешное"… See the full description on the dataset page: https://huggingface.co/datasets/samedad/mem-and-russian-jokes-dataset.CHISTES_spanish_jokes
Dataset Card for "CHISTES_spanish_jokes"
Dataset from Workshop for NLP introduction with Spanish jokes
More Information needed
multilingual-llm-jokes-4o-claude-gemini
Rapidata Generated Joke Preference Dataset
We collected 1'000'000+ human opinions on the jokes generated by state-of-the-art LLMs to decide which model is the funniest. The labelers are shown a joke in their language and asked to answer 'Yes' or 'No' to the question 'Is this joke funny?'.
It took us less than 5 days to get all of the responses.
The jokes are evenly distributed across 5 languages: English, Arabic, Japanese, Vietnamese, Portuguese and across 4 model… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/multilingual-llm-jokes-4o-claude-gemini.short_jokes_embedsSemEval-2021-Task-7-JokesAn unmodified copy of the datasets contained in the github repo "hahackathon-2021", which can be found here. Several other Github repositories include copies of this dataset, but I've linked the specific repository I found first.
Includes the following datasets:
hahackathon_dev.csv
hahackathon_test.csv
hahackathon_train.csv
The former two do not include any information other than the joke text and a unique ID, unfortunately. I'm not sure where the rest of the data is published, if at all… See the full description on the dataset page: https://huggingface.co/datasets/hfht/SemEval-2021-Task-7-Jokes.rated_jokes_dataset_from_jesterdowcipy-polish-jokes-dataset
Dataset consisting of polish jokes
Warning: Jokes were not curated, some may be offensive, stupid or simply not funny. It's highly recommended to filter jokes before training, e.g., based on downvotes
This dataset consists of all (9k) jokes dumped from jeja.pl on 2024-02-14. Jokes are submitted by the community. Besides the funny text itself, I included upvotes and downvotes. You can use them for filtering.Default sorting is based on a combination of downvotes and… See the full description on the dataset page: https://huggingface.co/datasets/JonaszPotoniec/dowcipy-polish-jokes-dataset.russian_jokes_with_vectors
Description in English:
The dataset is collected from Russian-language Telegram channels with jokes and anecdotes,The dataset was collected and tagged automatically using the data collection and tagging service Scoutie.Try Scoutie and collect the same or another dataset using link for FREE.
Dataset fields:
taskId - task identifier in the Scouti service. text - main text. url - link to the publication. sourceLink - link to Telegram. subSourceLink - link to the… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/russian_jokes_with_vectors.jokes_on_usreddit_jokesrated_jokes_dataset_from_jester_rlhf_formatrated_jokes_dataset_from_jester_rlhf_format_with_reasoningjokes-pizza-redochilean-humor-jokesCHISTES_spanish_jokesjester_jokes_extractedpeft-for-humor-sample-jokesjokes-inf-logitsjokes-pizzaCHISTES_spanish_jokes
Dataset Card for "CHISTES_spanish_jokes"
Dataset from Workshop for NLP introduction with Spanish jokes
More Information needed
results_joke_gen_of_mistral_sft_reddit_jokes_jo_ensemble_testresults_joke_gen_of_mistral_sft_reddit_jokes_ft_user_pref_1000_jo_ensemble_testjokes-final-predictions
