datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
multilingual-sentiments
Multilingual Sentiments Dataset
A collection of multilingual sentiments datasets grouped into 3 classes -- positive, neutral, negative.
Most multilingual sentiment datasets are either 2-class positive or negative, 5-class ratings of products reviews (e.g. Amazon multilingual dataset) or multiple classes of emotions. However, to an average person, sometimes positive, negative and neutral classes suffice and are more straightforward to perceive and annotate. Also, a positive/negative… See the full description on the dataset page: https://huggingface.co/datasets/tyqiangz/multilingual-sentiments.arabic-sentiments2finance-financialmodelingprep-stock-news-sentiments-rss-feed
Dataset Card for "finance-financialmodelingprep-stock-news-sentiments-rss-feed"
More Information needed
Sentiments-FinBERT-PT-BR
Dataset
A manually annotated dataset was created to enable supervised training for the FinBERT-PT-BR model, which focuses on sentiment analysis of Brazilian Portuguese financial texts.
More than 1.4 million financial news texts in Portuguese were collected and used for the initial language modeling phase. From this corpus, a sample of 1,000 texts was manually annotated with sentiment labels.
Annotation Process
Three annotators participated in the process.
All texts were… See the full description on the dataset page: https://huggingface.co/datasets/lucas-leme/Sentiments-FinBERT-PT-BR.swati-sentiments-corpus
Swati Sentiment Corpus
Dataset Description
This dataset contains sentiment-labeled text data in Swati for binary sentiment classification (Positive/Negative). Sentiments are extracted and processed from the English meanings of the sentences using DistilBERT for sentiment classification. The dataset is part of a larger collection of African language sentiment analysis resources.
Dataset Statistics
Total samples: 96,002
Positive sentiment: 56515 (58.9%)
Negative… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/swati-sentiments-corpus.hinglish-youtube-sentiments-dataset
Hinglish YouTube Comments Sentiment Dataset
A manually annotated dataset of 3,190 Hinglish YouTube comments for 3-class sentiment classification. Hinglish is the code-mixed Hindi-English language used by hundreds of millions of Indians online — written in Roman script, mixing Hindi and English words fluidly within the same sentence.
This dataset was created because no sufficiently large, cleanly annotated Hinglish sentiment dataset existed for YouTube comment data specifically.… See the full description on the dataset page: https://huggingface.co/datasets/shae2977/hinglish-youtube-sentiments-dataset.bitcoin-news-sentiments-latest
Bitcoin News Headlines with Directional Sentiment Labels
This dataset relabels the headlines from Bitcoin News Sentiment Dataset
by filipemunizz + google news headlines I pulled(2024-July 2026). This release consists of two parts: a full relabel of the original headlines using DeepSeek V4 Flash and an explicit directional rubric, and a set of 500 synthetic headlines added afterward to correct two specific weaknesses found during model validation.
Manual review of the original… See the full description on the dataset page: https://huggingface.co/datasets/HiyawErtiro/bitcoin-news-sentiments-latest.twi-sentiments-corpus-400k
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
This dataset is made available because of Ghana NLP's volunteer driven research work. Please consider contributing to any of our projects on Github
Twi Sentiment Corpus
Dataset Description
This dataset contains… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/twi-sentiments-corpus-400k.lingala-sentiments-corpus
Lingala Sentiment Corpus
Dataset Description
This dataset contains sentiment-labeled text data in Lingala for binary sentiment classification (Positive/Negative). Sentiments are extracted and processed from the English meanings of the sentences using DistilBERT for sentiment classification. The dataset is part of a larger collection of African language sentiment analysis resources.
Dataset Statistics
Total samples: 427,979
Positive sentiment: 251923 (58.9%)… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/lingala-sentiments-corpus.sentiments_instructionrundi-sentiments-corpus
Rundi Sentiment Corpus
Dataset Description
This dataset contains sentiment-labeled text data in Rundi for binary sentiment classification (Positive/Negative). Sentiments are extracted and processed from the English meanings of the sentences using DistilBERT for sentiment classification. The dataset is part of a larger collection of African language sentiment analysis resources.
Dataset Statistics
Total samples: 372,663
Positive sentiment: 209740 (56.3%)… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/rundi-sentiments-corpus.SentimentSynth
SentimentSynth Dataset
Overview
The SentimentSynth dataset is a collection of text samples expressing various sentiments, ranging from joy and excitement to stress and sadness. These samples are generated to simulate human-like expressions of emotions in different contexts.
Citation
If you use the SentimentSynth dataset in your work, please cite it as:
@misc {helpingai_2024,
author = { {HelpingAI} },
title = { SentimentSynth (Revision… See the full description on the dataset page: https://huggingface.co/datasets/OEvortex/SentimentSynth.sentimentsAllSides-sentimentsmossi-sentiments-corpus
Mossi Sentiment Corpus
Dataset Description
This dataset contains sentiment-labeled text data in Mossi for binary sentiment classification (Positive/Negative). Sentiments are extracted and processed from the English meanings of the sentences using DistilBERT for sentiment classification. The dataset is part of a larger collection of African language sentiment analysis resources.
Dataset Statistics
Total samples: 125,695
Positive sentiment: 74409 (59.2%)… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/mossi-sentiments-corpus.multilingual-sentiments
multilingual-sentiments
Multilingual sentiments (tyqiangz): 3-way sentiment in 12 languages, merged from public corpora (source).
Original data: tyqiangz/multilingual-sentiments. Repackaged as parquet for tasksource by scripts/upload_repackaged.py.
pedi-sentiments-corpus
Pedi Sentiment Corpus
Dataset Description
This dataset contains sentiment-labeled text data in Pedi for binary sentiment classification (Positive/Negative). Sentiments are extracted and processed from the English meanings of the sentences using DistilBERT for sentiment classification. The dataset is part of a larger collection of African language sentiment analysis resources.
Dataset Statistics
Total samples: 422,975
Positive sentiment: 255703 (60.5%)
Negative… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/pedi-sentiments-corpus.iwslt17_google_trans_scores_sentimentssentiments_dataset_azerbaijaniSentiments Dataset in Azerbaijani
Description
A collection of sentiments datasets in Azerbaijani language grouped into 3 classes: positive, neutral, negative. Data was collected from various sources such as social networks, reviews. It was created in 2024 and contains 42k text with label.
Format
The dataset is provided in comma-separated values (CSV) format. Each row represented on a new line with the following fields separated by commas:
text: user review / comment
labels: sentiment label… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/sentiments_dataset_azerbaijani.umbundu-sentiments-corpus
Umbundu Sentiment Corpus
Dataset Description
This dataset contains sentiment-labeled text data in Umbundu for binary sentiment classification (Positive/Negative). Sentiments are extracted and processed from the English meanings of the sentences using DistilBERT for sentiment classification. The dataset is part of a larger collection of African language sentiment analysis resources.
Dataset Statistics
Total samples: 83,350
Positive sentiment: 48940 (58.7%)… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/umbundu-sentiments-corpus.amharic-sentiments-corpus
Amharic Sentiment Corpus
Dataset Description
This dataset contains sentiment-labeled text data in Amharic for binary sentiment classification (Positive/Negative). Sentiments are extracted and processed from the English meanings of the sentences using DistilBERT for sentiment classification. The dataset is part of a larger collection of African language sentiment analysis resources.
Dataset Statistics
Total samples: 1,199,999
Positive sentiment: 667555 (55.6%)… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/amharic-sentiments-corpus.swahili-sentiments-corpus
Swahili Sentiment Corpus
Dataset Description
This dataset contains sentiment-labeled text data in Swahili for binary sentiment classification (Positive/Negative). Sentiments are extracted and processed from the English meanings of the sentences using DistilBERT for sentiment classification. The dataset is part of a larger collection of African language sentiment analysis resources.
Dataset Statistics
Total samples: 1,500,000
Positive sentiment: 808179 (53.9%)… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/swahili-sentiments-corpus.wolof-sentiments-corpus
Wolof Sentiment Corpus
Dataset Description
This dataset contains sentiment-labeled text data in Wolof for binary sentiment classification (Positive/Negative). Sentiments are extracted and processed from the English meanings of the sentences using DistilBERT for sentiment classification. The dataset is part of a larger collection of African language sentiment analysis resources.
Dataset Statistics
Total samples: 320,609
Positive sentiment: 179014 (55.8%)… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/wolof-sentiments-corpus.African-Languages_Sentiments
African Languages Sentiment Dataset (Hausa, Yorùbá, Swahili)
A stitched multi-source sentiment classification dataset combining three
independently collected sentiment corpora for Hausa, Yorùbá, and Swahili,
built for the Adaption Labs AutoScientist Challenge
(Language category).
Companion model: fine-tuned weights trained on the adapted version of this dataset via
AutoScientist are released separately at… See the full description on the dataset page: https://huggingface.co/datasets/gospelgit/African-Languages_Sentiments.xhosa-sentiments-corpus
Xhosa Sentiment Corpus
Dataset Description
This dataset contains sentiment-labeled text data in Xhosa for binary sentiment classification (Positive/Negative). Sentiments are extracted and processed from the English meanings of the sentences using DistilBERT for sentiment classification. The dataset is part of a larger collection of African language sentiment analysis resources.
Dataset Statistics
Total samples: 1,499,997
Positive sentiment: 841785 (56.1%)… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/xhosa-sentiments-corpus.comments-sentimentslabel: {'Neutral': 0, 'Positive': 1, 'Negative': 2}
kabuverdianu-sentiments-corpus
Kabuverdianu Sentiment Corpus
Dataset Description
This dataset contains sentiment-labeled text data in Kabuverdianu for binary sentiment classification (Positive/Negative). Sentiments are extracted and processed from the English meanings of the sentences using DistilBERT for sentiment classification. The dataset is part of a larger collection of African language sentiment analysis resources.
Dataset Statistics
Total samples: 99,553
Positive sentiment: 59631… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/kabuverdianu-sentiments-corpus.igbo-sentiments-corpus
Igbo Sentiment Corpus
Dataset Description
This dataset contains sentiment-labeled text data in Igbo for binary sentiment classification (Positive/Negative). Sentiments are extracted and processed from the English meanings of the sentences using DistilBERT for sentiment classification. The dataset is part of a larger collection of African language sentiment analysis resources.
Dataset Statistics
Total samples: 188,595
Positive sentiment: 102837 (54.5%)
Negative… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/igbo-sentiments-corpus.zulu-sentiments-corpus
Zulu Sentiment Corpus
Dataset Description
This dataset contains sentiment-labeled text data in Zulu for binary sentiment classification (Positive/Negative). Sentiments are extracted and processed from the English meanings of the sentences using DistilBERT for sentiment classification. The dataset is part of a larger collection of African language sentiment analysis resources.
Dataset Statistics
Total samples: 187,435
Positive sentiment: 102512 (54.7%)
Negative… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/zulu-sentiments-corpus.tumbuka-sentiments-corpus
Tumbuka Sentiment Corpus
Dataset Description
This dataset contains sentiment-labeled text data in Tumbuka for binary sentiment classification (Positive/Negative). Sentiments are extracted and processed from the English meanings of the sentences using DistilBERT for sentiment classification. The dataset is part of a larger collection of African language sentiment analysis resources.
Dataset Statistics
Total samples: 190,542
Positive sentiment: 109522 (57.5%)… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/tumbuka-sentiments-corpus.
