datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
REDDIT_comments
Dataset Card for "REDDIT_comments"
Dataset Summary
Comments of 50 high-quality subreddits, extracted from the REDDIT PushShift data dumps (from 2006 to Jan 2023).
Supported Tasks
These comments can be used for text generation and language modeling, as well as dialogue modeling.
Dataset Structure
Data Splits
Each split corresponds to a specific subreddit in the following list: "tifu", "explainlikeimfive", "WritingPrompts", "changemyview"… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceGECLM/REDDIT_comments.mm2_level_comments
Mario Maker 2 level comments
Part of the Mario Maker 2 Dataset Collection
Dataset Description
The Mario Maker 2 level comment dataset consists of 31.9 million level comments from Nintendo's online service totaling around 20GB of data. The dataset was created using the self-hosted Mario Maker 2 api over the course of 1 month in February 2022.
How to use it
The Mario Maker 2 level comment dataset is a very large dataset so for most use cases it is recommended to… See the full description on the dataset page: https://huggingface.co/datasets/TheGreatRambler/mm2_level_comments.stack-dedup-alt-comments
Dataset Description
This is the Python, Java and JavaScript subsets of The Stack (v1.1) after cleaning* and agressive deduplication from stack-dedup-alt-decontaminate with
filtering on comment to code ratio with minimum of 0.01 and maximum of 0.8.
The additional comments filtering removes 26.5% of the dataset's volume which goes from 215GB of text to 170GB.
(*) cleaning: near deduplication + PII redaction + line length & percentage of alphanumeric characters filtering + data… See the full description on the dataset page: https://huggingface.co/datasets/bigcode/stack-dedup-alt-comments.task1721_civil_comments_obscenity_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1721_civil_comments_obscenity_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1721_civil_comments_obscenity_classification.task1723_civil_comments_sexuallyexplicit_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1723_civil_comments_sexuallyexplicit_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1723_civil_comments_sexuallyexplicit_classification.task1720_civil_comments_toxicity_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1720_civil_comments_toxicity_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1720_civil_comments_toxicity_classification.task1724_civil_comments_insult_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1724_civil_comments_insult_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1724_civil_comments_insult_classification.Algerian-Youtube-Comments
Algerian Youtube Comments
55,365 raw YouTube comments on Algeria-related videos for Darija social-text modeling, from the Algerian NLP Collective. Counted 2026-09-17 via the Hub datasets-server (/info?dataset=algerian-nlp/Algerian-Youtube-Comments: 55,365 train rows) and re-counted row-by-row with datasets streaming (load_dataset("algerian-nlp/Algerian-Youtube-Comments", split="train", streaming=True): 55,365 rows).
The default config answers: how do Algerians actually write in… See the full description on the dataset page: https://huggingface.co/datasets/algerian-nlp/Algerian-Youtube-Comments.digikala-commentstask1722_civil_comments_threat_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1722_civil_comments_threat_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1722_civil_comments_threat_classification.hacker_news_with_comments
Dataset Card for [Dataset Name]
Dataset Summary
Hacker news until 2015 with comments. Collect from Google BigQuery open dataset. We didn't do any pre-processing except remove HTML tags.
Supported Tasks and Leaderboards
Comment Generation; News analysis with comments; Other comment-based NLP tasks.
Languages
English
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/Linkseed/hacker_news_with_comments.reddit-comments-sample
Reddit Comment Trees Sample — Initial snapshot
Initial sample: 43,913 posts and 59,874 comments across three communities.
This is a selected, structurally checked snapshot, with incomplete subreddit coverage. It is not a complete three-month archive.
Overview
Posts and associated comments from r/LocalLLaMA, r/wallstreetbets, and r/SkincareAddiction.
The requested post window is June 1–August 31, 2026 in Asia/Shanghai, with UTC bounds 2026-05-31 16:00:00… See the full description on the dataset page: https://huggingface.co/datasets/mashu-data/reddit-comments-sample.truexa-comments
Truha audience comments
45,918 real audience comments from the public satirical news channel
«Труха⚡️Україна» (June 30 – July 22, 2026), collected from posts and
cleaned (ads, links, mentions, ultra-short fragments, duplicates and
phone numbers dropped — 50,000 raw → 45,918 clean).
The audience writes in both Ukrainian and Russian (a mix, not a
clean split — a share of comments code-switch between the two), so the
corpus carries language: [ru, uk].
These are the "taste anchor"… See the full description on the dataset page: https://huggingface.co/datasets/hausmer/truexa-comments.DC_inside_comments
DC_inside_comments
This dataset contains 110,000 raw comments collected from DC Inside. It is intended for unsupervised learning or pretraining purposes.
Dataset Summary
Data Type: Unlabeled raw comments
Number of Examples: 110,000
Source: DC Inside
Related Dataset
For labeled data and multi-task annotated examples, please refer to the KoMultiText dataset.
How to Load the Dataset
from datasets import load_dataset
# Load the unlabeled dataset… See the full description on the dataset page: https://huggingface.co/datasets/Dasool/DC_inside_comments.jira-comments-nsp
Dataset Card for Dataset Name
Dataset Summary
Dataset contains pairs of sentences with next_sentence_label for NSP. Sentences was given from public jira projects dataset. Next sentence is always next sentence in one comment or sentence from reply to the comment.
Supported Tasks and Leaderboards
NSP, MLM
Languages
English
Dataset Structure
sentence_a, sentence_b, next_sentence_label
Source Data… See the full description on the dataset page: https://huggingface.co/datasets/pheepa/jira-comments-nsp.LocalLLaMA-comments
LocalLLaMA-comments
A companion dataset to pszemraj/LocalLLaMA-posts. Time frame is in sync (up through Tue Mar 3 9PM EST 2026)
youtube-comments-sentiment
YouTube Comments Sentiment Dataset
375 labeled YouTube comments for sentiment analysis and toxicity detection research.
Dataset Structure
Fields
comment: Raw comment text (includes emojis, informal language)
sentiment: positive / negative / neutral
toxic: true / false
video_category: Content category of the source video
Splits
train: 300 examples
test: 75 examples
task1725_civil_comments_severtoxicity_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1725_civil_comments_severtoxicity_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1725_civil_comments_severtoxicity_classification.telegram-post-comments_data
Telegram Post-Comment Pairs (Russian)
Описание
Датасет пар «пост — комментарий» из публичных русскоязычных Telegram-каналов. Предназначен для задачи стилевой адаптации генеративных языковых моделей: модель получает текст поста и должна сгенерировать комментарий, стилистически согласованный с реальными пользовательскими откликами.
Структура
Split
Примеров
train
8 246
validation
1 034
test
1 020
Каждый пример содержит два поля:
post — текст… See the full description on the dataset page: https://huggingface.co/datasets/r9zhenka/telegram-post-comments_data.toxic_commentsillegal-job-titles-comments
My Job Sounds Illegal Dataset
A dataset of ~42,000 humorous social media comments collected from a viral reel prompt:
"Comment your job but make it sound illegal."
The dataset contains creative, exaggerated, and comedic descriptions of professions and daily work activities written to sound suspicious, criminal, or absurd while remaining harmless and humorous.
Examples include:
"I manipulate vulnerable people into buying things they don't need."
"I convince tiny humans to obey me… See the full description on the dataset page: https://huggingface.co/datasets/BibbyResearch/illegal-job-titles-comments.
