datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Digikala-CommentsJigsaw-Toxic-Comments
Jigsaw Toxic Comments — canonical 2017–2018 Kaggle release
A verbatim mirror of the Jigsaw Toxic Comment Classification Challenge training and test data, repackaged as three gzip-compressed CSVs. No rows added, removed, or reordered relative to the upstream Kaggle release — only the hosting moved and each file was gzipped.
Re-hosted under Heliosoph for ingestion-pipeline stability — the upstream files live behind Kaggle's competition-rules click-through and require an… See the full description on the dataset page: https://huggingface.co/datasets/Heliosoph/Jigsaw-Toxic-Comments.hacker_news_with_comments
Dataset Card for [Dataset Name]
Dataset Summary
Hacker news until 2015 with comments. Collect from Google BigQuery open dataset. We didn't do any pre-processing except remove HTML tags.
Supported Tasks and Leaderboards
Comment Generation; News analysis with comments; Other comment-based NLP tasks.
Languages
English
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/Linkseed/hacker_news_with_comments.reddit-comments-demoCharlie-Kirk-comments
Dataset Summary
This dataset is used in the following Github repository: "Charlie-Kirk-comments-sentiment-analysis".
This dataset contains Reddit comments related to Charlie Kirk, founder of Turning Point USA, from 11-2024 to 10-2025,
which includes the event of his death in 10th September 2025.
Charlie Kirk was a highly polarizing political figure in the United States, often drawing both strong support and harsh criticism.
Following his death, many media platforms, included… See the full description on the dataset page: https://huggingface.co/datasets/Proyecto-charlie-kirk-reddit/Charlie-Kirk-comments.suicide-comments-es
Dataset Summary
The dataset consists of comments on Reddit, Twitter, and inputs/outputs of the Alpaca dataset translated to Spanish language and classified as suicidal ideation/behavior and non-suicidal.
Dataset Structure
The dataset has 10050 rows (777 considered as Suicidal Ideation/Behavior and 9273 considered Not Suicidal).
Dataset fields
Text: User comment.
Label: 1 if suicidal ideation/behavior; 0 if not suicidal comment.
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2023/suicide-comments-es.ahsenwaheed_youtube-comments-spam-dataset
Youtube Comments Spam Dataset
Predicting YouTube Comment Spam: An Insightful Dataset for Text Classification
Dataset Info
Source: Kaggle
Original Size: 0.16 MB
Kaggle Downloads: 4,105
Files: 1
Files
Youtube-Spam-Dataset.csv
Mirrored from Kaggle
fpt-comments-sentiment
fpt-comments-sentiment
This is a hand-curated Vietnamese feedback dataset built from FPT-related discussions across Facebook communities, FuOverflow, Reddit/VOZ-style forums, FPT web pages, and other student-community sources. (Inspired by the NEU-ESC dataset)
The data is curated by manually collected examples with Facebook group comments scraped using Playwright, then cleaned, redacted, hand-reviewed, and labeled with support from LLM-assisted review.
Each row contains a… See the full description on the dataset page: https://huggingface.co/datasets/lducc/fpt-comments-sentiment.vietnamese-social-comments
🇻🇳 Bộ dữ liệu phân loại bình luận tiếng Việt
Bộ dữ liệu này bao gồm 4.896 bình luận tiếng Việt được thu thập từ nhiều nền tảng mạng xã hội phổ biến như TikTok, Facebook, YouTube,...Mỗi bình luận được gán nhãn theo 2 cấp độ:
label: thể hiện cảm xúc hoặc thái độ tổng thể.
category: phân loại chi tiết theo ngữ nghĩa hoặc mục đích cụ thể của câu.
🔖 Cấu trúc dữ liệu
Trường
Kiểu dữ liệu
Mô tả
comment
string
Văn bản bình luận (có thể viết tắt, không dấu… See the full description on the dataset page: https://huggingface.co/datasets/vanhai123/vietnamese-social-comments.Reddit_Comments_DatasetJigsaw-Toxic-Comments
Jigsaw Toxic Comments — canonical 2017–2018 Kaggle release
A verbatim mirror of the Jigsaw Toxic Comment Classification Challenge training and test data, repackaged as three gzip-compressed CSVs. No rows added, removed, or reordered relative to the upstream Kaggle release — only the hosting moved and each file was gzipped.
Re-hosted under Heliosoph for ingestion-pipeline stability — the upstream files live behind Kaggle's competition-rules click-through and require an… See the full description on the dataset page: https://huggingface.co/datasets/preethi16102005/Jigsaw-Toxic-Comments.toxic-comments
Toxic-comments (Teeny-Tiny Castle)
This dataset is part of a tutorial tied to the Teeny-Tiny Castle, an open-source repository containing educational tools for AI Ethics and Safety research.
How to Use
from datasets import load_dataset
dataset = load_dataset("AiresPucrs/toxic_content", split = 'train')
hungarian-toxic-comments
Hungarian Toxic Comments
The first openly available Hungarian dataset for toxic comment classification, introduced in:
Hatvani, P., & Yang, Z. Gy. (2025). Automated detection of toxic comments in Hungarian. Annales Mathematicae et Informaticae, 61, 108-117. DOI: 10.33039/ami.2025.10.007
Dataset Description
This dataset contains 654 manually annotated Hungarian-language comments collected from social media and political news forums. Each comment is annotated across five… See the full description on the dataset page: https://huggingface.co/datasets/RabidUmarell/hungarian-toxic-comments.Jigsaw-Toxic-Comments
Jigsaw Toxic Comments — canonical 2017–2018 Kaggle release
A verbatim mirror of the Jigsaw Toxic Comment Classification Challenge training and test data, repackaged as three gzip-compressed CSVs. No rows added, removed, or reordered relative to the upstream Kaggle release — only the hosting moved and each file was gzipped.
Re-hosted under Heliosoph for ingestion-pipeline stability — the upstream files live behind Kaggle's competition-rules click-through and require an… See the full description on the dataset page: https://huggingface.co/datasets/peu123/Jigsaw-Toxic-Comments.Nostalgic_Sentiment_Analysis_of_YouTube_Comments_Data
Dataset Summary
The dataset is a collection of Youtube Comments and it was captured using the YouTube Data API.
The data set consists of 1500 nostalgic and non-nostalgic comments in English.
Languages
The language of the data is English.
Citation
If you find this dataset usefull for your study, please cite the paper as followed:
@article{postalcioglu2020comparison,
title={Comparison of Neural Network Models for Nostalgic Sentiment Analysis of YouTube… See the full description on the dataset page: https://huggingface.co/datasets/Senem/Nostalgic_Sentiment_Analysis_of_YouTube_Comments_Data.instagram-comments-sentimentvietnamese-caucu-comments
Vietnamese Cau Cuu Facebook Comments
Dataset Summary
This dataset contains Vietnamese Facebook comments collected from a natural-disaster discussion thread and auto-labeled for binary emergency detection.
The target task is to detect whether a comment is a real-time rescue request (cau_cuu) versus a non-emergency comment (khong_phai_cau_cuu).
This release is intended as a bootstrap dataset for triage modeling and should be treated as a weakly supervised resource. Human… See the full description on the dataset page: https://huggingface.co/datasets/dat201204/vietnamese-caucu-comments.spam_ham_commentsyoutube_top_popular_videos_commentsyoutube-commentsthis is a very bad dataset. a better one comming soon.
commentscivil_comments_cleanFrench_Youtube_CommentsCe dataset contient un scraping de commentaires Youtube sur des chaînes "grands publics" destinées aux jeunes.
Nous avons notamment scrapé 187269 commentaires sous 29 vidéos de Squeezie.
L'autre fichier, qui contient 191856 commentaires, contient pour une bonne part les commentaires sous 39 vidéos de Michou.
Le dataset, en l'état actuel n'est pas nettoyé , c'est donné comme c'est sorti de l'API !
Pour une partie, j'ai supprimé la colonne 'username'. Mais elle est reconstructible de plusieurs… See the full description on the dataset page: https://huggingface.co/datasets/GwendalTsang/French_Youtube_Comments.persian_commercial_comments_filingdz-sentiment-yt-comments
A Sentiment Analysis Dataset for the Algerian Dialect of Arabic
This dataset consists of 50,016 samples of comments extracted from Algerian YouTube channels. It is manually annotated with 3 classes (the label column) and is not balanced. Here are the number of rows of each class:
0 (Negative): 17,033 (34.06%)
1 (Neutral): 11,136 (22.26%)
2 (Positive): 21,847 (43.68%)
Please note that there are some swear words in the dataset, so please use it with caution.
Citation… See the full description on the dataset page: https://huggingface.co/datasets/Abdou/dz-sentiment-yt-comments.reddit_comments_sp500
Reddit Finance Comments for S&P 500 Companies
Overview
This dataset contains Reddit comments collected from posts discussing S&P 500 companies, with a focus on high-quality and relevant finance conversations.Each comment is linked to its original post, and detailed metadata for both comments and posts is included.
Original posts are sourced from the emilpartow/reddit_finance_posts_sp500 dataset.Comments were collected and filtered for quality, then thoroughly cleaned to… See the full description on the dataset page: https://huggingface.co/datasets/emilpartow/reddit_comments_sp500.vkplay-smuta-commentsNOTE:
датасет актуален в период (04.04.2024-06.04.2024)
данные отзывов агрегированны с количеством купленных/запущенных игр на аккаунте (если профиль не скрыт)
YouTube-Comments-Dataset-45k-rows
YouTube Comments Dataset with Sentiment, Toxicity, and Spam Labels (45K Rows)
📘 Overview
This dataset contains over 45,000 real-world YouTube comments collected from a diverse set of YouTube channels across genres such as entertainment, news, devotional content, and education. Each comment has been automatically annotated with:
label_sentiment: positive, neutral, or negative
label_toxicity: toxic or non-toxic
label_spam: spam or not spam
The dataset is suitable… See the full description on the dataset page: https://huggingface.co/datasets/atakshat11/YouTube-Comments-Dataset-45k-rows.RO_offensive_political_commentsyoutube-comments-intent-sentiment
