datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
NLU-Sentiment-Analysis
SEA Sentiment Analysis
SEA Sentiment Analysis evaluates a model's ability to identify the sentiment polarity of a text. It is sampled from NusaX for Indonesian, Javanese, and Sundanese, IndicSentiment for Tamil, Wisesight Sentiment for Thai, and UIT-VSFC for Vietnamese.
Supported Tasks and Leaderboards
SEA Sentiment Analysis is designed for evaluating chat or instruction-tuned large language models (LLMs). It is part of the SEA-HELM leaderboard from AI Singapore.… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/NLU-Sentiment-Analysis.SentimentAnalysisHindi
SentimentAnalysisHindi
An MTEB dataset
Massive Text Embedding Benchmark
Hindi Sentiment Analysis Dataset
Task category
t2c
Domains
Reviews, Written
Reference
https://huggingface.co/datasets/OdiaGenAI/sentiment_analysis_hindi
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_task("SentimentAnalysisHindi")
evaluator = mteb.MTEB([task])
model =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/SentimentAnalysisHindi.saraiki-sentiment-analysis-datasetbengali_sentiment_analysis
Bengali Sentiment Analysis
Context
The dataset contains 3307 Negative reviews and 8500 Positive reviews collected and manually annotated from Youtube Bengali drama.
Positive_Label=1 and Negative_Label=0
Acknowledgements
Sazzed, Salim (2021), “Bangla ( Bengali ) sentiment analysis classification benchmark dataset corpus”, Mendeley Data, V4, doi: 10.17632/p6zc7krs37.4
sentiment-analysis-in-commodity-market-gold
Dataset Card for Sentiment Analysis of Commodity News (Gold)
This is a news dataset for the commodity market which has been manually annotated for 10,000+ news headlines across multiple dimensions into various classes. The dataset has been sampled from a period of 20+ years (2000-2021).
The dataset was curated by Ankur Sinha and Tanmay Khandait and is detailed in their paper "Impact of News on the Commodity Market: Dataset and Results." It is currently published by the authors on… See the full description on the dataset page: https://huggingface.co/datasets/SaguaroCapital/sentiment-analysis-in-commodity-market-gold.portuguese_sentiment_analysisThis dataset is based on the dataset originally posted in Kaggle
aspect-based-sentiment-analysis-uzbeksentiment_analysis_data
Dataset Card for "sentiment_analysis_data"
More Information needed
sentiment-analysis-pt
Sentiment Analysis PT (Teeny-Tiny Castle)
This dataset is part of a tutorial tied to the Teeny-Tiny Castle, an open-source repository containing educational tools for AI Ethics and Safety research.
How to Use
from datasets import load_dataset
dataset = load_dataset("AiresPucrs/sentiment-analysis-pt", split = 'train')
sentiment-analysis-ind-classification
SentimentAnalysis_ind_Classification
Deduplicated copy of kornwtp/sentiment-analysis-ind-classification.
Splits
split
rows
train
10,082
sentiment-analysis
Sentiment Analysis (Teeny-Tiny Castle)
This dataset is part of a tutorial tied to the Teeny-Tiny Castle, an open-source repository containing educational tools for AI Ethics and Safety research.
How to Use
from datasets import load_dataset
dataset = load_dataset("AiresPucrs/sentiment-analysis", split = 'train')
tweets_pt_sentiment_analysis
Dataset Card for "tweets_pt_sentiment_analysis"
More Information needed
financial_sentiment_analysis_train_compilation
Dataset Card for "financial_sentiment_analysis_train_compilation"
More Information needed
twitter-sentiment-analysis
Twitter Sentiment Analysis: Prabowo's First 100 Days
Dataset Overview
This dataset contains tweets related to President Prabowo Subianto's first 100 days in office in Indonesia (2024-2029). The tweets have been preprocessed and classified into three sentiment categories using a fine-tuned BERT model for Indonesian language (IndoBERT).
Dataset Details
Language: Indonesian
Source: Twitter/X
Time period: First 100 days of President Prabowo's… See the full description on the dataset page: https://huggingface.co/datasets/KidzRizal/twitter-sentiment-analysis.CaSSA-catalan-structured-sentiment-analysis
Dataset Card for CaSSA, the Catalan Structured Sentiment Analysis dataset
Dataset Summary
The CaSSA dataset is a corpus of 6,400 reviews and forum messages annotated with polar expressions. Each piece of text is annotated with all the expressions of polarity that it contains. For each polar expression, we annotated the expression itself, the target (the object of the expression), and the source (the subject expressing the sentiment). 25,453 polar expressions have been… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/CaSSA-catalan-structured-sentiment-analysis.roman-urdu-sentiment-analysisSentiment_Analysis_SLUE-VoxCelebsentiment_analysis_training_testamazon_reviews_multi_fr_prompt_sentiment_analysis
amazon_reviews_multi_fr_prompt_sentiment_analysis
Summary
amazon_reviews_multi_fr_prompt_sentiment_analysis is a subset of the Dataset of French Prompts (DFP).It contains 5,880,000 rows that can be used for a binary sentiment analysis task.The original data (without prompts) comes from the dataset amazon_reviews_multi by Keung et al. where only the French part has been kept.A list of prompts (see below) was then applied in order to build the input and target columns and… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/amazon_reviews_multi_fr_prompt_sentiment_analysis.uzbek-sentiment-analysis
Uzum Market Sentiment Analysis Dataset
About the Dataset
This dataset is created for performing sentiment analysis on comments from Uzum Market. The dataset allows for evaluating the sentiments in the comments and categorizing them into various ratings.
Data Structure
The data is provided in JSON format with the following structure:
{
"features": ["normalized_review_text", "rating"],
"num_rows": 352151
}
Ratings are defined as follows:
rating_to_label… See the full description on the dataset page: https://huggingface.co/datasets/risqaliyevds/uzbek-sentiment-analysis.sentiment-analysis-for-financial-news
Sentiment Analysis for Financial News (FinancialPhraseBank-style)
This dataset contains financial news headlines labeled with sentiment from the perspective of a retail investor.
Columns
sentiment: one of negative, neutral, positive
news_headline: financial news headline text
Source / Context
The original dataset is commonly referenced as FinancialPhraseBank and is widely used for financial sentiment benchmarking.
Usage
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/prithvi1029/sentiment-analysis-for-financial-news.Sentiment_Analysis_Tweetsmyanmar-social-media-sentiment-analysis-dataset
Myanmar Social Media Sentiment Analysis Dataset
A Myanmar language dataset for sentiment analysis of social media content, translated from an English source dataset.
Dataset Description
This dataset contains social media text with sentiment annotations translated into Myanmar language. It is derived from the original Social Media Sentiments Analysis Dataset on Kaggle, with texts professionally translated to Myanmar language while preserving the sentiment labels.… See the full description on the dataset page: https://huggingface.co/datasets/chuuhtetnaing/myanmar-social-media-sentiment-analysis-dataset.txsa_twitter_sentiment_analysissentiment-analysis-ind-classificationCOVID_Vaccine_Tweet_sentiment_analysis_roberta
Dataset Card for "COVID_Vaccine_Tweet_sentiment_analysis_roberta"
More Information needed
sentiment-analysis-finetune
Dataset Card for "sentiment-analysis-finetune"
More Information needed
allocine_fr_prompt_sentiment_analysis
allocine_fr_prompt_sentiment_analysis
Summary
allocine_fr_prompt_sentiment_analysis is a subset of the Dataset of French Prompts (DFP).It contains 5,600,000 rows that can be used for a binary sentiment analysis task.The original data (without prompts) comes from the dataset allocine by Blard.A list of prompts (see below) was then applied in order to build the input and target columns and thus obtain the same format as the xP3 dataset by Muennighoff et al.… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/allocine_fr_prompt_sentiment_analysis.txsa_twitter_sentiment_analysis_fullInflation_News_Sentiment_Analysis
