datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tweet_sentiment_extraction
TweetSentimentExtractionClassification
An MTEB dataset
Massive Text Embedding Benchmark
Task category
t2c
Domains
Social, Written
Reference
https://www.kaggle.com/competitions/tweet-sentiment-extraction/overview
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["TweetSentimentExtractionClassification"])
evaluator = mteb.MTEB(task)
model =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/tweet_sentiment_extraction.sentiment-dksfThe Sentiment DKSF (Digikala/Snappfood comments) is a dataset for sentiment analysis.
twitter-financial-news-sentiment
Dataset Description
The Twitter Financial News dataset is an English-language dataset containing an annotated corpus of finance-related tweets. This dataset is used to classify finance-related tweets for their sentiment.
The dataset holds 11,932 documents annotated with 3 labels:
sentiments = {
"LABEL_0": "Bearish",
"LABEL_1": "Bullish",
"LABEL_2": "Neutral"
}
The data was collected using the Twitter API. The current dataset supports the multi-class classification… See the full description on the dataset page: https://huggingface.co/datasets/zeroshot/twitter-financial-news-sentiment.NLU-Sentiment-Analysis
SEA Sentiment Analysis
SEA Sentiment Analysis evaluates a model's ability to identify the sentiment polarity of a text. It is sampled from NusaX for Indonesian, Javanese, and Sundanese, IndicSentiment for Tamil, Wisesight Sentiment for Thai, and UIT-VSFC for Vietnamese.
Supported Tasks and Leaderboards
SEA Sentiment Analysis is designed for evaluating chat or instruction-tuned large language models (LLMs). It is part of the SEA-HELM leaderboard from AI Singapore.… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/NLU-Sentiment-Analysis.sentiment140Sentiment140 consists of Twitter messages with emoticons, which are used as noisy labels for
sentiment classification. For more detailed information please refer to the paper.fingpt-sentiment-train
Dataset Card for "fingpt-sentiment-train"
More Information needed
poem_sentiment
Dataset Card for Gutenberg Poem Dataset
Dataset Summary
Poem Sentiment is a sentiment dataset of poem verses from Project Gutenberg.
This dataset can be used for tasks such as sentiment classification or style transfer for poems.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
The text in the dataset is in English (en).
Dataset Structure
Data Instances
Example of one instance in the dataset.
{'id': 0… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/poem_sentiment.descriptiveness-sentiment-trl-style
TRL's Sentiment and Descriptiveness Preference Dataset
The dataset comes from https://arxiv.org/abs/1909.08593, one of the earliest RLHF work from OpenAI.
We preprocess the dataset using our standard prompt, chosen, rejected format.
Reproduce this dataset
Download the descriptiveness_sentiment.py from the https://huggingface.co/datasets/trl-internal-testing/descriptiveness-sentiment-trl-style/tree/0.1.0.
Run python examples/datasets/descriptiveness_sentiment.py… See the full description on the dataset page: https://huggingface.co/datasets/trl-internal-testing/descriptiveness-sentiment-trl-style.tweet_sentiment_multilingual
Dataset Card for cardiffnlp/tweet_sentiment_multilingual
Dataset Summary
Tweet Sentiment Multilingual consists of sentiment analysis dataset on Twitter in 8 different lagnuages.
arabic
english
french
german
hindi
italian
portuguese
spanish
Supported Tasks and Leaderboards
text_classification: The dataset can be trained using a SentenceClassification model from HuggingFace transformers.
Dataset Structure
Data Instances
An instance from… See the full description on the dataset page: https://huggingface.co/datasets/cardiffnlp/tweet_sentiment_multilingual.banking_sentiment_setfit
Dataset Card for "banking_sentiment_setfit"
More Information needed
my_sentimenttwitter-airline-sentiment
Dataset Card for Twitter US Airline Sentiment
Dataset Summary
This data originally came from Crowdflower's Data for Everyone library.
As the original source says,
A sentiment analysis job about the problems of each major U.S. airline. Twitter data was scraped from February of 2015 and contributors were asked to first classify positive, negative, and neutral tweets, followed by categorizing negative reasons (such as "late flight" or "rude service").
The data we're… See the full description on the dataset page: https://huggingface.co/datasets/osanseviero/twitter-airline-sentiment.Arabic_Sentiment_Twitter_Corpus
Dataset Card for [Dataset Name]
Dataset Summary
[More Information Needed]
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation
Curation Rationale
[More Information Needed]
Source Data… See the full description on the dataset page: https://huggingface.co/datasets/arbml/Arabic_Sentiment_Twitter_Corpus.wisesight_sentiment
Dataset Card for wisesight_sentiment
Dataset Summary
Wisesight Sentiment Corpus: Social media messages in Thai language with sentiment label (positive, neutral, negative, question)
Released to public domain under Creative Commons Zero v1.0 Universal license.
Labels: {"pos": 0, "neu": 1, "neg": 2, "q": 3}
Size: 26,737 messages
Language: Central Thai
Style: Informal and conversational. With some news headlines and advertisement.
Time period: Around 2016 to early 2019. With… See the full description on the dataset page: https://huggingface.co/datasets/pythainlp/wisesight_sentiment.sentiment-bankingvoxceleb-sentimentPersian_sentimentcrypto-market-sentiment-observations
Instrumetriq: Crypto Market Activity & Sentiment Context Dataset
Time-aligned observational snapshots of crypto market activity and social sentiment across 270+ assets, designed to contextualize market structure, liquidity, and attention dynamics.
Observational data only. No trading advice, predictions, or signal generation.
Dataset Description
This dataset provides weekly Sunday snapshots from Instrumetriq's continuous monitoring pipeline. Each snapshot… See the full description on the dataset page: https://huggingface.co/datasets/Instrumetriq/crypto-market-sentiment-observations.multilingual-sentiments
Multilingual Sentiments Dataset
A collection of multilingual sentiments datasets grouped into 3 classes -- positive, neutral, negative.
Most multilingual sentiment datasets are either 2-class positive or negative, 5-class ratings of products reviews (e.g. Amazon multilingual dataset) or multiple classes of emotions. However, to an average person, sometimes positive, negative and neutral classes suffice and are more straightforward to perceive and annotate. Also, a positive/negative… See the full description on the dataset page: https://huggingface.co/datasets/tyqiangz/multilingual-sentiments.multilingual-sentiment-classification
MultilingualSentimentClassification
An MTEB dataset
Massive Text Embedding Benchmark
Sentiment classification dataset with binary
(positive vs negative sentiment) labels. Includes 30 languages and dialects.
Task category
t2c
DomainsReviews, Written
Reference
https://huggingface.co/datasets/mteb/multilingual-sentiment-classification
How to evaluate on this task
You can evaluate an embedding model on this dataset using the… See the full description on the dataset page: https://huggingface.co/datasets/mteb/multilingual-sentiment-classification.turkish-sentiment-analysis-dataset
Dataset
This dataset contains positive , negative and notr sentences from several data sources given in the references. In the most sentiment models , there are only two labels; positive and negative. However , user input can be totally notr sentence. For such cases there were no data I could find. Therefore I created this dataset with 3 class. Positive and negative sentences are listed below. Notr examples are extraced from turkish wiki dump. In addition, added some random text… See the full description on the dataset page: https://huggingface.co/datasets/winvoker/turkish-sentiment-analysis-dataset.Arabic_Sentiment_Twitter_Corpus
Dataset Card for "Arabic_Sentiment_Twitter_Corpus"
Source:
https://www.kaggle.com/datasets/mksaad/arabic-sentiment-twitter-corpus
amazon-reviews-sentiment-analysis
Dataset Card for amazon reviews for sentiment analysis
Dataset Summary
One of the most important problems in e-commerce is the correct calculation of the points given to after-sales products. The solution to this problem is to provide greater customer satisfaction for the e-commerce site, product prominence for sellers, and a seamless shopping experience for buyers. Another problem is the correct ordering of the comments given to the products. The prominence of misleading… See the full description on the dataset page: https://huggingface.co/datasets/hugginglearners/amazon-reviews-sentiment-analysis.tweet_sentiment_multilingual
TweetSentimentClassification
An MTEB dataset
Massive Text Embedding Benchmark
A multilingual Sentiment Analysis dataset consisting of tweets in 8 different languages.
Task category
t2c
Domains
Social, Written
Reference
https://aclanthology.org/2022.lrec-1.27
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["TweetSentimentClassification"])
evaluator =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/tweet_sentiment_multilingual.AfriSenti-twitter-sentimentAfriSenti is the largest sentiment analysis benchmark dataset for under-represented African languages---covering 110,000+ annotated tweets in 14 African languages (Amharic, Algerian Arabic, Hausa, Igbo, Kinyarwanda, Moroccan Arabic, Mozambican Portuguese, Nigerian Pidgin, Oromo, Swahili, Tigrinya, Twi, Xitsonga, and yoruba).snappfood-sentiment-analysisfiqa-sentiment-classification
Dataset Name
Dataset Description
This dataset is based on the task 1 of the Financial Sentiment Analysis in the Wild (FiQA) challenge. It follows the same settings as described in the paper 'A Baseline for Aspect-Based Sentiment Analysis in Financial Microblogs and News'. The dataset is split into three subsets: train, valid, test with sizes 822, 117, 234 respectively.
Dataset Structure
_id: ID of the data point
sentence: The sentence
target: The target of the… See the full description on the dataset page: https://huggingface.co/datasets/TheFinAI/fiqa-sentiment-classification.climate_sentiment
Dataset Card for climate_sentiment
Dataset Summary
We introduce an expert-annotated dataset for classifying climate-related sentiment of climate-related paragraphs in corporate disclosures.
Supported Tasks and Leaderboards
The dataset supports a ternary sentiment classification task of whether a given climate-related paragraph has sentiment opportunity, neutral, or risk.
Languages
The text in the dataset is in English.
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/climatebert/climate_sentiment.digikala-sentiment-analysisacme-sentiment-511pin4k
Customer Sentiment Corpus
Samples: 25000
License: mit
Language: en
Description
Sentiment-labeled customer feedback corpus.
Usage
Intended for sentiment classification of customer feedback.
