CoolFace
21 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01HuggingFaceGECLM /REDDIT_comments Dataset Card for "REDDIT_comments" Dataset Summary Comments of 50 high-quality subreddits, extracted from the REDDIT PushShift data dumps (from 2006 to Jan 2023). Supported Tasks These comments can be used for text generation and language modeling, as well as dialogue modeling. Dataset Structure Data Splits Each split corresponds to a specific subreddit in the following list: "tifu", "explainlikeimfive", "WritingPrompts", "changemyview"… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceGECLM/REDDIT_comments.texttext-generation100M<n<1B24 likes9.7k downloads4y agoHugging Face02TheGreatRambler /mm2_level_comments Mario Maker 2 level comments Part of the Mario Maker 2 Dataset Collection Dataset Description The Mario Maker 2 level comment dataset consists of 31.9 million level comments from Nintendo's online service totaling around 20GB of data. The dataset was created using the self-hosted Mario Maker 2 api over the course of 1 month in February 2022. How to use it The Mario Maker 2 level comment dataset is a very large dataset so for most use cases it is recommended to… See the full description on the dataset page: https://huggingface.co/datasets/TheGreatRambler/mm2_level_comments.tabularother10M<n<100M3 likes760 downloads4y agoHugging Face03bigcode /stack-dedup-alt-commentsgated Dataset Description This is the Python, Java and JavaScript subsets of The Stack (v1.1) after cleaning* and agressive deduplication from stack-dedup-alt-decontaminate with filtering on comment to code ratio with minimum of 0.01 and maximum of 0.8. The additional comments filtering removes 26.5% of the dataset's volume which goes from 215GB of text to 170GB. (*) cleaning: near deduplication + PII redaction + line length & percentage of alphanumeric characters filtering + data… See the full description on the dataset page: https://huggingface.co/datasets/bigcode/stack-dedup-alt-comments.tabulartext-generation10M<n<100M0 likes286 downloads3y agoHugging Face04Lots-of-LoRAs /task1721_civil_comments_obscenity_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1721_civil_comments_obscenity_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1721_civil_comments_obscenity_classification.texttext-generationn<1K1 likes230 downloads2y agoHugging Face05Lots-of-LoRAs /task1723_civil_comments_sexuallyexplicit_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1723_civil_comments_sexuallyexplicit_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1723_civil_comments_sexuallyexplicit_classification.texttext-generationn<1K0 likes226 downloads2y agoHugging Face06Lots-of-LoRAs /task1720_civil_comments_toxicity_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1720_civil_comments_toxicity_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1720_civil_comments_toxicity_classification.texttext-generationn<1K0 likes184 downloads2y agoHugging Face07Lots-of-LoRAs /task1724_civil_comments_insult_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1724_civil_comments_insult_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1724_civil_comments_insult_classification.texttext-generation1K<n<10K0 likes165 downloads2y agoHugging Face08algerian-nlp /Algerian-Youtube-Comments Algerian Youtube Comments 55,365 raw YouTube comments on Algeria-related videos for Darija social-text modeling, from the Algerian NLP Collective. Counted 2026-09-17 via the Hub datasets-server (/info?dataset=algerian-nlp/Algerian-Youtube-Comments: 55,365 train rows) and re-counted row-by-row with datasets streaming (load_dataset("algerian-nlp/Algerian-Youtube-Comments", split="train", streaming=True): 55,365 rows). The default config answers: how do Algerians actually write in… See the full description on the dataset page: https://huggingface.co/datasets/algerian-nlp/Algerian-Youtube-Comments.texttext-generation10K<n<100K0 likes165 downloads5d agoHugging Face09EhsanShahbazi /digikala-commentsgatedtabulartext-classification10M<n<100M0 likes117 downloads9mo agoHugging Face10Lots-of-LoRAs /task1722_civil_comments_threat_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1722_civil_comments_threat_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1722_civil_comments_threat_classification.texttext-generationn<1K0 likes107 downloads2y agoHugging Face11Linkseed /hacker_news_with_comments Dataset Card for [Dataset Name] Dataset Summary Hacker news until 2015 with comments. Collect from Google BigQuery open dataset. We didn't do any pre-processing except remove HTML tags. Supported Tasks and Leaderboards Comment Generation; News analysis with comments; Other comment-based NLP tasks. Languages English Data Fields [More Information Needed] Data Splits [More Information Needed] Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/Linkseed/hacker_news_with_comments.tabulartext-generation1M<n<10M6 likes96 downloads4y agoHugging Face12mashu-data /reddit-comments-sample Reddit Comment Trees Sample — Initial snapshot Initial sample: 43,913 posts and 59,874 comments across three communities. This is a selected, structurally checked snapshot, with incomplete subreddit coverage. It is not a complete three-month archive. Overview Posts and associated comments from r/LocalLLaMA, r/wallstreetbets, and r/SkincareAddiction. The requested post window is June 1–August 31, 2026 in Asia/Shanghai, with UTC bounds 2026-05-31 16:00:00… See the full description on the dataset page: https://huggingface.co/datasets/mashu-data/reddit-comments-sample.tabulartext-generation100K<n<1M0 likes67 downloads10d agoHugging Face13hausmer /truexa-comments Truha audience comments 45,918 real audience comments from the public satirical news channel «Труха⚡️Україна» (June 30 – July 22, 2026), collected from posts and cleaned (ads, links, mentions, ultra-short fragments, duplicates and phone numbers dropped — 50,000 raw → 45,918 clean). The audience writes in both Ukrainian and Russian (a mix, not a clean split — a share of comments code-switch between the two), so the corpus carries language: [ru, uk]. These are the "taste anchor"… See the full description on the dataset page: https://huggingface.co/datasets/hausmer/truexa-comments.tabulartext-classification10K<n<100K0 likes56 downloads7d agoHugging Face14Dasool /DC_inside_comments DC_inside_comments This dataset contains 110,000 raw comments collected from DC Inside. It is intended for unsupervised learning or pretraining purposes. Dataset Summary Data Type: Unlabeled raw comments Number of Examples: 110,000 Source: DC Inside Related Dataset For labeled data and multi-task annotated examples, please refer to the KoMultiText dataset. How to Load the Dataset from datasets import load_dataset # Load the unlabeled dataset… See the full description on the dataset page: https://huggingface.co/datasets/Dasool/DC_inside_comments.texttext-generation100K<n<1M0 likes37 downloads2y agoHugging Face15pheepa /jira-comments-nsp Dataset Card for Dataset Name Dataset Summary Dataset contains pairs of sentences with next_sentence_label for NSP. Sentences was given from public jira projects dataset. Next sentence is always next sentence in one comment or sentence from reply to the comment. Supported Tasks and Leaderboards NSP, MLM Languages English Dataset Structure sentence_a, sentence_b, next_sentence_label Source Data… See the full description on the dataset page: https://huggingface.co/datasets/pheepa/jira-comments-nsp.texttext-generationn<1K0 likes36 downloads4y agoHugging Face16pszemraj /LocalLLaMA-comments LocalLLaMA-comments A companion dataset to pszemraj/LocalLLaMA-posts. Time frame is in sync (up through Tue Mar 3 9PM EST 2026) tabulartext-generation1M<n<10M1 likes32 downloads7mo agoHugging Face17Okcan /youtube-comments-sentiment YouTube Comments Sentiment Dataset 375 labeled YouTube comments for sentiment analysis and toxicity detection research. Dataset Structure Fields comment: Raw comment text (includes emojis, informal language) sentiment: positive / negative / neutral toxic: true / false video_category: Content category of the source video Splits train: 300 examples test: 75 examples text-classificationn<1K0 likes18 downloads2mo agoHugging Face18Lots-of-LoRAs /task1725_civil_comments_severtoxicity_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1725_civil_comments_severtoxicity_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1725_civil_comments_severtoxicity_classification.texttext-generationn<1K0 likes17 downloads2y agoHugging Face19r9zhenka /telegram-post-comments_data Telegram Post-Comment Pairs (Russian) Описание Датасет пар «пост — комментарий» из публичных русскоязычных Telegram-каналов. Предназначен для задачи стилевой адаптации генеративных языковых моделей: модель получает текст поста и должна сгенерировать комментарий, стилистически согласованный с реальными пользовательскими откликами. Структура Split Примеров train 8 246 validation 1 034 test 1 020 Каждый пример содержит два поля: post — текст… See the full description on the dataset page: https://huggingface.co/datasets/r9zhenka/telegram-post-comments_data.texttext-generation10K<n<100K0 likes15 downloads6mo agoHugging Face20hramphul /toxic_commentstext-classification1M<n<10M0 likes8 downloads2y agoHugging Face21BibbyResearch /illegal-job-titles-commentsgated My Job Sounds Illegal Dataset A dataset of ~42,000 humorous social media comments collected from a viral reel prompt: "Comment your job but make it sound illegal." The dataset contains creative, exaggerated, and comedic descriptions of professions and daily work activities written to sound suspicious, criminal, or absurd while remaining harmless and humorous. Examples include: "I manipulate vulnerable people into buying things they don't need." "I convince tiny humans to obey me… See the full description on the dataset page: https://huggingface.co/datasets/BibbyResearch/illegal-job-titles-comments.text-generation10K<n<100K3 likes4 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.