datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
arabic-sentiments2Sentiments-FinBERT-PT-BR
Dataset
A manually annotated dataset was created to enable supervised training for the FinBERT-PT-BR model, which focuses on sentiment analysis of Brazilian Portuguese financial texts.
More than 1.4 million financial news texts in Portuguese were collected and used for the initial language modeling phase. From this corpus, a sample of 1,000 texts was manually annotated with sentiment labels.
Annotation Process
Three annotators participated in the process.
All texts were… See the full description on the dataset page: https://huggingface.co/datasets/lucas-leme/Sentiments-FinBERT-PT-BR.hinglish-youtube-sentiments-dataset
Hinglish YouTube Comments Sentiment Dataset
A manually annotated dataset of 3,190 Hinglish YouTube comments for 3-class sentiment classification. Hinglish is the code-mixed Hindi-English language used by hundreds of millions of Indians online — written in Roman script, mixing Hindi and English words fluidly within the same sentence.
This dataset was created because no sufficiently large, cleanly annotated Hinglish sentiment dataset existed for YouTube comment data specifically.… See the full description on the dataset page: https://huggingface.co/datasets/shae2977/hinglish-youtube-sentiments-dataset.African-Languages_Sentiments
African Languages Sentiment Dataset (Hausa, Yorùbá, Swahili)
A stitched multi-source sentiment classification dataset combining three
independently collected sentiment corpora for Hausa, Yorùbá, and Swahili,
built for the Adaption Labs AutoScientist Challenge
(Language category).
Companion model: fine-tuned weights trained on the adapted version of this dataset via
AutoScientist are released separately at… See the full description on the dataset page: https://huggingface.co/datasets/gospelgit/African-Languages_Sentiments.hindi-sentimentslabel: {'Neutral': 0, 'Positive': 1, 'Negative': 2}
Bug_Reports_with_Sentimentsvnmese_sentimentsSentimentssentiment-sample-dataset
Sentiment Analysis Sample Dataset
A small curated dataset of text samples labeled with sentiment polarity for educational purposes.
Dataset Description
This dataset contains 15 text samples annotated with sentiment labels (positive, neutral, negative) along with confidence scores.
Features
id: Unique identifier (1-15)
sentence: Text excerpt (15-50 characters)
sentiment: Sentiment classification (positive/neutral/negative)
confidence: Confidence score between… See the full description on the dataset page: https://huggingface.co/datasets/ASHOKSIMHADRI/sentiment-sample-dataset.multilingual-sentimentscuba_hotels_sentiments_essentimentsanalysis
