datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Algerian-Youtube-Commentsyoutube-comment-sentiment
YouTube Comments Sentiment Analysis Dataset (1M+ Labeled Comments)
Overview
This dataset comprises over one million YouTube comments, each annotated with sentiment labels—Positive, Neutral, or Negative. The comments span a diverse range of topics including programming, news, sports, politics and more, and are enriched with comprehensive metadata to facilitate various NLP and sentiment analysis tasks.
How to use:
import pandas as pd
df =… See the full description on the dataset page: https://huggingface.co/datasets/AmaanP314/youtube-comment-sentiment.Algerian-Youtube-Comments
Algerian Youtube Comments
55,365 raw YouTube comments on Algeria-related videos for Darija social-text modeling, from the Algerian NLP Collective. Counted 2026-09-17 via the Hub datasets-server (/info?dataset=algerian-nlp/Algerian-Youtube-Comments: 55,365 train rows) and re-counted row-by-row with datasets streaming (load_dataset("algerian-nlp/Algerian-Youtube-Comments", split="train", streaming=True): 55,365 rows).
The default config answers: how do Algerians actually write in… See the full description on the dataset page: https://huggingface.co/datasets/algerian-nlp/Algerian-Youtube-Comments.youtube-comment-sentiment
YouTube Comments Sentiment Analysis Dataset (1M+ Labeled Comments)
Overview
This dataset comprises over one million YouTube comments, each annotated with sentiment labels—Positive, Neutral, or Negative. The comments span a diverse range of topics including programming, news, sports, politics and more, and are enriched with comprehensive metadata to facilitate various NLP and sentiment analysis tasks.
How to use:
import pandas as pd
df =… See the full description on the dataset page: https://huggingface.co/datasets/vnkat/youtube-comment-sentiment.Nostalgic_Sentiment_Analysis_of_YouTube_Comments_Data
Dataset Summary
The dataset is a collection of Youtube Comments and it was captured using the YouTube Data API.
The data set consists of 1500 nostalgic and non-nostalgic comments in English.
Languages
The language of the data is English.
Citation
If you find this dataset usefull for your study, please cite the paper as followed:
@article{postalcioglu2020comparison,
title={Comparison of Neural Network Models for Nostalgic Sentiment Analysis of YouTube… See the full description on the dataset page: https://huggingface.co/datasets/Senem/Nostalgic_Sentiment_Analysis_of_YouTube_Comments_Data.ahsenwaheed_youtube-comments-spam-dataset
Youtube Comments Spam Dataset
Predicting YouTube Comment Spam: An Insightful Dataset for Text Classification
Dataset Info
Source: Kaggle
Original Size: 0.16 MB
Kaggle Downloads: 4,105
Files: 1
Files
Youtube-Spam-Dataset.csv
Mirrored from Kaggle
youtube-commentsthis is a very bad dataset. a better one comming soon.
youtube-comments-v2youtube_comment_sentiment_dataset_preprocessing
유튜브 댓글 긍정, 부정, 중립 분류 데이터셋
0: 긍정, 1: 부정, 2: 중립
Youtube_shorts_comments
Fine-tuned distilgpt2 on this dataset
The average amount of emojis in a YouTube short comment is 4.59 (Based on this dataset)
12 millions view 😂😂😂 good god 🤦♂️
youtube-comments-180k@misc {matthew_mitton_2025,
author = { {Matthew Mitton} },
title = { youtube-comments-180k (Revision dee1941) },
year = 2025,
url = {\url{https://huggingface.co/datasets/breadlicker45/youtube-comments-180k} },
doi = { 10.57967/hf/4742 },
publisher = { Hugging Face }
}
youtube_top_popular_videos_commentsyoutube-bot-comments
Note for users, README.md generated by Claude 4 Sonnet.
Dataset Card for YouTube Korean Bot Comment Dataset
This dataset contains YouTube comments from the top 50 Korean videos, classified to identify bot-generated content for research in automated content detection and Korean natural language processing.
Dataset Details
Dataset Description
This dataset consists of 185,830 YouTube comments collected from the top 50 Korean videos and classified for bot… See the full description on the dataset page: https://huggingface.co/datasets/MisileLab/youtube-bot-comments.youtube-bot-comments-v2
Dataset Card for youtube-bot-comment-v2
This dataset contains Korean YouTube comments labeled for bot detection, focusing on identifying automated comments that promote adult content or gambling websites.
Dataset Details
Dataset Description
This dataset consists of Korean YouTube comments collected from top South Korean videos, with binary classification labels indicating whether each comment is generated by a bot or a human user. The dataset specifically… See the full description on the dataset page: https://huggingface.co/datasets/MisileLab/youtube-bot-comments-v2.youtube-comment-sentiment
YouTube Comments Sentiment Analysis Dataset (1M+ Labeled Comments)
Overview
This dataset comprises over one million YouTube comments, each annotated with sentiment labels—Positive, Neutral, or Negative. The comments span a diverse range of topics including programming, news, sports, politics and more, and are enriched with comprehensive metadata to facilitate various NLP and sentiment analysis tasks.
How to use:
import pandas as pd
df =… See the full description on the dataset page: https://huggingface.co/datasets/krish12209/youtube-comment-sentiment.shawgpt-youtube-commentsDataset for ShawGPT, a fine-tuned data science YouTube comment responder.
Video link: https://youtu.be/XpoKB3usmKc
Blog link: https://medium.com/towards-data-science/qlora-how-to-fine-tune-an-llm-on-a-single-gpu-4e44d6b5be32
youtube_comment_sentiment_dataset_preprocessingFrench_Youtube_CommentsCe dataset contient un scraping de commentaires Youtube sur des chaînes "grands publics" destinées aux jeunes.
Nous avons notamment scrapé 187269 commentaires sous 29 vidéos de Squeezie.
L'autre fichier, qui contient 191856 commentaires, contient pour une bonne part les commentaires sous 39 vidéos de Michou.
Le dataset, en l'état actuel n'est pas nettoyé , c'est donné comme c'est sorti de l'API !
Pour une partie, j'ai supprimé la colonne 'username'. Mais elle est reconstructible de plusieurs… See the full description on the dataset page: https://huggingface.co/datasets/GwendalTsang/French_Youtube_Comments.ChatGPT-Sentiment-Analysis-YouTube-Comments-Datasetbiagpt-youtube-commentsyoutube_comment_sentiment_dataset_preprocessingyoutube_comment_sentiment_dataset_preprocessingTurkish-Youtube-Comments
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/yusiqo/Turkish-Youtube-Comments.test-shawgpt-youtube-commentsyoutube_comments_sentiment_analysisyoutube-comments-sentiment
YouTube Comments Sentiment Dataset
375 labeled YouTube comments for sentiment analysis and toxicity detection research.
Dataset Structure
Fields
comment: Raw comment text (includes emojis, informal language)
sentiment: positive / negative / neutral
toxic: true / false
video_category: Content category of the source video
Splits
train: 300 examples
test: 75 examples
youtube-comments50-YouTube-Commentsyoutube_comment_sentiment_dataset_preprocessingyoutube_comment_sentiment_dataset_preprocessing
