CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01HuggingFaceGECLM /REDDIT_comments Dataset Card for "REDDIT_comments" Dataset Summary Comments of 50 high-quality subreddits, extracted from the REDDIT PushShift data dumps (from 2006 to Jan 2023). Supported Tasks These comments can be used for text generation and language modeling, as well as dialogue modeling. Dataset Structure Data Splits Each split corresponds to a specific subreddit in the following list: "tifu", "explainlikeimfive", "WritingPrompts", "changemyview"… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceGECLM/REDDIT_comments.texttext-generation100M<n<1B24 likes9.8k downloads4y agoHugging Face02google /civil_comments Dataset Card for "civil_comments" Dataset Summary The comments in this dataset come from an archive of the Civil Comments platform, a commenting plugin for independent news sites. These public comments were created from 2015 - 2017 and appeared on approximately 50 English-language news sites across the world. When Civil Comments shut down in 2017, they chose to make the public comments available in a lasting open archive to enable future research. The original data… See the full description on the dataset page: https://huggingface.co/datasets/google/civil_comments.tabulartext-classification1M<n<10M40 likes8.9k downloads3y agoHugging Face03fddemarco /pushshift-reddit-comments Dataset Card for "pushshift-reddit" More Information needed tabular1B<n<10B27 likes3.5k downloads3y agoHugging Face04Tungtom2004 /850_commentsimage1K<n<10K0 likes3.1k downloads14d agoHugging Face05touati-kamel /Algerian-Youtube-Commentstext10K<n<100K0 likes1.9k downloads7d agoHugging Face06AISE-TUDelft /leading-comments Dataset Card for Dataset Name Disclaimer: This dataset will be released shortly. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More Information Needed] Paper [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/AISE-TUDelft/leading-comments.text100M<n<1B0 likes1.3k downloads2y agoHugging Face07Jalik /TrojanLoc-gpt-4.1-embeddings-wo-comments0 likes1.1k downloads11mo agoHugging Face08TheGreatRambler /mm2_level_comments Mario Maker 2 level comments Part of the Mario Maker 2 Dataset Collection Dataset Description The Mario Maker 2 level comment dataset consists of 31.9 million level comments from Nintendo's online service totaling around 20GB of data. The dataset was created using the self-hosted Mario Maker 2 api over the course of 1 month in February 2022. How to use it The Mario Maker 2 level comment dataset is a very large dataset so for most use cases it is recommended to… See the full description on the dataset page: https://huggingface.co/datasets/TheGreatRambler/mm2_level_comments.tabularother10M<n<100M3 likes763 downloads4y agoHugging Face09community-datasets /danish_political_comments Dataset Card for DanishPoliticalComments Dataset Summary The dataset consists of 9008 sentences that are labeled with fine-grained polarity in the range from -2 to 2 (negative to positive). The quality of the fine-grained is not cross-validated and is therefore subject to uncertainties; however, the simple polarity has been cross-validated and therefore is considered to be more correct. Supported Tasks and Leaderboards [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/danish_political_comments.texttext-classification1K<n<10K3 likes760 downloads2y agoHugging Face10Jkatzy /code-comments-small Comment Dataset Opening comments extracted from code datasets with CommentMiner and ML4SE-toolkit. Files are grouped as <dataset>/<language>/part-*.parquet. The Hugging Face dataset card declares one config per source dataset and one split-safe language name per language. Each row contains dataset, record_id, opening_comment, language, path, repo, extracted_at, and metadata. For Parquet exports, metadata is stored as a JSON string so every source dataset shares one stable… See the full description on the dataset page: https://huggingface.co/datasets/Jkatzy/code-comments-small.0 likes676 downloads2mo agoHugging Face11ro-h /regulatory_commentsUnited States governmental agencies often make proposed regulations open to the public for comment. Proposed regulations are organized into "dockets". This project will use Regulation.gov public API to aggregate and clean public comments for dockets that mention opioid use. Each example will consist of one docket, and include metadata such as docket id, docket title, etc. Each docket entry will also include information about the top 10 comments, including comment metadata and comment text.text-classificationn<1K46 likes503 downloads3y agoHugging Face12davanstrien /test-sync-commentstextn<1K0 likes450 downloads3y agoHugging Face13RadeAI /Digikala_comments_products Digikala Dataset (Comments & Products) Unveiling Insights into Digikala's Product Universe and Customer Sentiments Welcome to the Digikala (Comments & Products) dataset, a treasure trove of information offering a comprehensive glimpse into the dynamic online marketplace of Digikala. Dive into a world comprising over 1.2 million products and more than 6 million comments, providing invaluable insights into consumer sentiments and market trends. About the Dataset… See the full description on the dataset page: https://huggingface.co/datasets/RadeAI/Digikala_comments_products.12 likes400 downloads3y agoHugging Face14alvanlii /reddit-comments-uwaterloo--- Generated Part of README Below --- Dataset Overview The goal is to have an open dataset of r/uwaterloo submissions, leveraging PRAW and the Reddit API to get downloads. Posts are here Comments are here Creation Details This dataset was created by alvanlii/dataset-creator-reddit-uwaterloo Update Frequency The dataset is updated custom with the most recent update being 2024-12-12 23:00:00 UTC+0000 where we added 72 new rows. Licensing… See the full description on the dataset page: https://huggingface.co/datasets/alvanlii/reddit-comments-uwaterloo.tabular1M<n<10M2 likes373 downloads2y agoHugging Face15nixiesearch /hackernews-comments Hackernews Comments Dataset A dataset of all HN API items from id=0 till id=41422887 (so from 2006 till 02 Sep 2024). The dataset is build by scraping the HN API according to its official schema and docs. Scraper code is also available on github: nixiesearch/hnscrape Dataset contents No cleaning, validation or filtering was performed. The resulting data files are raw JSON API response dumps in zstd-compressed JSONL files. An example payload: { "by": "goldfish"… See the full description on the dataset page: https://huggingface.co/datasets/nixiesearch/hackernews-comments.tabular10M<n<100M3 likes346 downloads2y agoHugging Face16fibonacciai /Digikala-Commentstabular1K<n<10K1 likes322 downloads1y agoHugging Face17aictsharif /persian-sentiment-comments Dataset Card for Persian Sentiment Comments Dataset Description persian-sentiment-comments is a curated dataset of Persian (Farsi) user comments, each annotated with a sentiment label. The dataset is designed for sentiment analysis tasks and is particularly suitable for training and evaluating machine learning models on Persian text data. The comments are collected from a variety of sources, including online product reviews and user feedback. Each comment is paired with a… See the full description on the dataset page: https://huggingface.co/datasets/aictsharif/persian-sentiment-comments.texttext-classification10K<n<100K0 likes299 downloads1y agoHugging Face18bigcode /stack-dedup-alt-commentsgated Dataset Description This is the Python, Java and JavaScript subsets of The Stack (v1.1) after cleaning* and agressive deduplication from stack-dedup-alt-decontaminate with filtering on comment to code ratio with minimum of 0.01 and maximum of 0.8. The additional comments filtering removes 26.5% of the dataset's volume which goes from 215GB of text to 170GB. (*) cleaning: near deduplication + PII redaction + line length & percentage of alphanumeric characters filtering + data… See the full description on the dataset page: https://huggingface.co/datasets/bigcode/stack-dedup-alt-comments.tabulartext-generation10M<n<100M0 likes290 downloads3y agoHugging Face19RebuttalAgent /Comments_200K Comments-200k 📚 1. Introduction Raw reviews often intermix critical points with extraneous content, such as salutations and summaries. Directly feeding this unprocessed text into a model introduces significant noise and redundancy, which can compromise the precision of the generated rebuttal. Furthermore, due to diverse reviewer writing styles and varying conference formats, comments are typically presented in an unstructured manner. Therefore, to address these… See the full description on the dataset page: https://huggingface.co/datasets/RebuttalAgent/Comments_200K.2 likes271 downloads11mo agoHugging Face20gorpeliates /swe-bench-commentsIncludes the SWE-bench dataset along with the corresponding comments for each instance_id text1K<n<10K0 likes232 downloads1y agoHugging Face21Lots-of-LoRAs /task1721_civil_comments_obscenity_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1721_civil_comments_obscenity_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1721_civil_comments_obscenity_classification.texttext-generationn<1K1 likes227 downloads2y agoHugging Face22Lots-of-LoRAs /task1723_civil_comments_sexuallyexplicit_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1723_civil_comments_sexuallyexplicit_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1723_civil_comments_sexuallyexplicit_classification.texttext-generationn<1K0 likes224 downloads2y agoHugging Face23AlexSham /Toxic_Russian_Commentshttps://www.kaggle.com/datasets/alexandersemiletov/toxic-russian-comments 0 - neutral user comments 1 - toxic user comments Toxic Russian Comments Dataset This dataset contains labelled comments from the popular Russian social network ok.ru. The data was used in a competition where participants had to automatically label each comment with at least one of the four predefined classes. The classes represent different levels of toxicity. The competition was held on the All Cups platform. Each… See the full description on the dataset page: https://huggingface.co/datasets/AlexSham/Toxic_Russian_Comments.texttext-classification100K<n<1M10 likes210 downloads2y agoHugging Face24mteb /danish_political_commentstext1K<n<10K0 likes209 downloads1y agoHugging Face25Jennyyahoo /RedditDataFor353_2026_COMMENTStabular1M<n<10M0 likes209 downloads1mo agoHugging Face26Morteza-Shahrabi-Farahani /Detecting-toxic-comments1 likes198 downloads3y agoHugging Face27nicohrubec /codebase-content-SWE-bench_Verified-with-comments-and-teststext100K<n<1M0 likes188 downloads1y agoHugging Face28Lots-of-LoRAs /task1720_civil_comments_toxicity_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1720_civil_comments_toxicity_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1720_civil_comments_toxicity_classification.texttext-generationn<1K0 likes172 downloads2y agoHugging Face29Heliosoph /Jigsaw-Toxic-Comments Jigsaw Toxic Comments — canonical 2017–2018 Kaggle release A verbatim mirror of the Jigsaw Toxic Comment Classification Challenge training and test data, repackaged as three gzip-compressed CSVs. No rows added, removed, or reordered relative to the upstream Kaggle release — only the hosting moved and each file was gzipped. Re-hosted under Heliosoph for ingestion-pipeline stability — the upstream files live behind Kaggle's competition-rules click-through and require an… See the full description on the dataset page: https://huggingface.co/datasets/Heliosoph/Jigsaw-Toxic-Comments.tabulartext-classification100K<n<1M0 likes171 downloads3mo agoHugging Face30algerian-nlp /Algerian-Youtube-Comments Algerian Youtube Comments 55,365 raw YouTube comments on Algeria-related videos for Darija social-text modeling, from the Algerian NLP Collective. Counted 2026-09-17 via the Hub datasets-server (/info?dataset=algerian-nlp/Algerian-Youtube-Comments: 55,365 train rows) and re-counted row-by-row with datasets streaming (load_dataset("algerian-nlp/Algerian-Youtube-Comments", split="train", streaming=True): 55,365 rows). The default config answers: how do Algerians actually write in… See the full description on the dataset page: https://huggingface.co/datasets/algerian-nlp/Algerian-Youtube-Comments.texttext-generation10K<n<100K0 likes163 downloads5d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.