datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sentiment_analysisSENTIPOLC 2016 dataset
The SENTIPOLC 2016 dataset contains 9410 tweets annotated for subjectivity, overall and literal polarity, and irony.
The dataset has been created and used in the context of the SENTIPOLC 2016 task (http://www.di.unito.it/~tutreeb/sentipolc-evalita16/index.html), organized as part of the EVALITA 2016 evaluation campaign.
Original files available here:
https://live.european-language-grid.eu/catalogue/corpus/7479/download/
If you find this dataset useful please cite:… See the full description on the dataset page: https://huggingface.co/datasets/evalitahf/sentiment_analysis.sentiment-analysis-for-financial-news-v2sentiment_analysis_hindiConventions followed to decide the polarity: -
labels consisting of a single value are left undisturbed, i.e. if label = 'pos', then it'll be pos
labels consisting of multiple values separated by '&' are processed. If all the labels are the same ('pos&pos&pos' or 'neg&neg'), then the shortened form of the multiple label is assigned as the final label. For example, if label = 'pos&pos&pos', then final label will be 'pos'.
labels consisting of mixed values ('pos&neg&pos' or 'neg&neu&pos') are… See the full description on the dataset page: https://huggingface.co/datasets/OdiaGenAI/sentiment_analysis_hindi.synthetic-persian-chatbot-conversational-sentiment-analysis-anger
Dataset Summary
Synthetic Persian Chatbot Conversational SA – Anger is a Persian (Farsi) dataset created for the Classification task, with a focus on detecting the emotion "anger" in chatbot conversations. It is part of the FaMTEB (Farsi Massive Text Embedding Benchmark). The dataset was synthetically generated using GPT-4o-mini and is derived from the broader Synthetic Persian Chatbot Conversational Sentiment Analysis dataset.
Language(s): Persian (Farsi)
Task(s): Classification… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/synthetic-persian-chatbot-conversational-sentiment-analysis-anger.synthetic-persian-chatbot-conversational-sentiment-analysis-fear
Dataset Summary
Synthetic Persian Chatbot Conversational SA – Fear is a Persian (Farsi) dataset for the Classification task, focused on detecting the expression of "fear" in user-chatbot conversations. It is part of the FaMTEB (Farsi Massive Text Embedding Benchmark). The dataset was synthetically generated using GPT-4o-mini and is a subset of the Synthetic Persian Chatbot Conversational Sentiment Analysis dataset.
Language(s): Persian (Farsi)
Task(s): Classification (Emotion… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/synthetic-persian-chatbot-conversational-sentiment-analysis-fear.synthetic-persian-chatbot-conversational-sentiment-analysis-sadness
Dataset Summary
Synthetic Persian Chatbot Conversational SA – Sadness is a Persian (Farsi) dataset for the Classification task, focused on detecting the expression of "sadness" in user-chatbot conversations. It is part of the FaMTEB (Farsi Massive Text Embedding Benchmark). This dataset was synthetically generated using GPT-4o-mini and is a subset of the broader Synthetic Persian Chatbot Conversational Sentiment Analysis dataset.
Language(s): Persian (Farsi)
Task(s):… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/synthetic-persian-chatbot-conversational-sentiment-analysis-sadness.synthetic-persian-chatbot-conversational-sentiment-analysis-happiness
Dataset Summary
Synthetic Persian Chatbot Conversational SA – Happiness is a Persian (Farsi) dataset for the Classification task, specifically focused on detecting the emotion "happiness" in user-chatbot conversations. It is part of the FaMTEB (Farsi Massive Text Embedding Benchmark). This dataset was synthetically generated using GPT-4o-mini, and is a subset of the broader Synthetic Persian Chatbot Conversational Sentiment Analysis collection.
Language(s): Persian (Farsi)… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/synthetic-persian-chatbot-conversational-sentiment-analysis-happiness.synthetic-persian-chatbot-conversational-sentiment-analysis-friendship
Dataset Summary
Synthetic Persian Chatbot Conversational SA – Friendship is a Persian (Farsi) dataset created for the Classification task, with a focus on detecting the emotion "friendship" in chatbot conversations. It is part of the FaMTEB (Farsi Massive Text Embedding Benchmark). The dataset was synthetically generated using GPT-4o-mini and is derived from the broader Synthetic Persian Chatbot Conversational Sentiment Analysis dataset.
Language(s): Persian (Farsi)
Task(s):… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/synthetic-persian-chatbot-conversational-sentiment-analysis-friendship.twitter-sentiment-analysis
🐦 Twitter Sentiment Analysis (bdstar/twitter-sentiment-analysis)
🧠 Overview
A refined and merged version of Twitter text sentiment datasets, providing a clean and well-balanced dataset for sentiment classification across three sentiment categories:positive, negative, and neutral.
This dataset is split into three parts — train, test, and validation — each sourced from highly reputable open datasets.It is designed for training, evaluating, and benchmarking NLP models for… See the full description on the dataset page: https://huggingface.co/datasets/bdstar/twitter-sentiment-analysis.synthetic-persian-chatbot-conversational-sentiment-analysis-tone-chatbot-classification
Dataset Summary
Synthetic Persian Chatbot Conversational SA – Chatbot Tone Classification(SynPerChatbotConvSAToneChatbotClassification) is a Persian (Farsi) dataset created for the Classification task. It focuses on identifying the chatbot’s conversational tone—formal, casual, or childish—in dialogue exchanges that include emotional content. This dataset is part of the FaMTEB (Farsi Massive Text Embedding Benchmark) and was synthetically generated using GPT-4o-mini.
Language(s):… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/synthetic-persian-chatbot-conversational-sentiment-analysis-tone-chatbot-classification.synthetic-persian-chatbot-conversational-sentiment-analysis-jealousy
Dataset Summary
Synthetic Persian Chatbot Conversational SA – Jealousy is a Persian (Farsi) dataset for the Classification task, focused on detecting the expression of "jealousy" in user-chatbot conversations. It is part of the FaMTEB (Farsi Massive Text Embedding Benchmark). This dataset was synthetically generated using GPT-4o-mini and is a subset of the broader Synthetic Persian Chatbot Conversational Sentiment Analysis dataset.
Language(s): Persian (Farsi)
Task(s):… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/synthetic-persian-chatbot-conversational-sentiment-analysis-jealousy.chatbot-conversational-sentiment-analysis-tone-user-classification
Dataset Summary
Synthetic Persian Chatbot Conversational SA – User Tone Classification(SynPerChatbotConvSAToneUserClassification) is a Persian (Farsi) dataset created for the Classification task. It focuses on identifying the user’s conversational tone—formal, casual, or childish—in emotionally rich chatbot interactions. This dataset is part of the FaMTEB (Farsi Massive Text Embedding Benchmark) and was synthetically generated using GPT-4o-mini.
Language(s): Persian (Farsi)… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/chatbot-conversational-sentiment-analysis-tone-user-classification.synthetic-persian-chatbot-conversational-sentiment-analysis-surprise
Dataset Summary
Synthetic Persian Chatbot Conversational SA – Surprise is a Persian (Farsi) dataset for the Classification task, focused on detecting the expression of "surprise" in user-chatbot conversations. It is part of the FaMTEB (Farsi Massive Text Embedding Benchmark). The dataset was synthetically generated using GPT-4o-mini and is a subset of the Synthetic Persian Chatbot Conversational Sentiment Analysis dataset.
Language(s): Persian (Farsi)
Task(s): Classification… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/synthetic-persian-chatbot-conversational-sentiment-analysis-surprise.synthetic-persian-chatbot-conversational-sentiment-analysis-satisfaction
Dataset Summary
Synthetic Persian Chatbot Conversational Sentiment Analysis – Satisfaction is a Persian (Farsi) dataset developed for the Classification task, specifically focused on detecting the emotion of satisfaction in chatbot conversations. It is part of the FaMTEB (Farsi Massive Text Embedding Benchmark) and was synthetically generated using the GPT-4o-mini language model.
Language(s): Persian (Farsi)
Task(s): Classification (Emotion Detection – Satisfaction)
Source:… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/synthetic-persian-chatbot-conversational-sentiment-analysis-satisfaction.Sentiment-Analysis-ComplexExcellent — congrats on getting the repo ready 🚀
Here’s a professional Hugging Face Dataset Card (README.md) you can paste directly into your repository.
This is written to match HF best practices and serious research usage.
📘 README.md
👉 Copy everything below into your README.md
Sentiment-Analysis-Complex
🧠 Overview
Sentiment-Analysis-Complex is a large-scale synthetic sentiment analysis dataset designed for benchmarking modern NLP models under… See the full description on the dataset page: https://huggingface.co/datasets/NNEngine/Sentiment-Analysis-Complex.synthetic-persian-chatbot-conversational-sentiment-analysis-love
Dataset Summary
Synthetic Persian Chatbot Conversational SA – Love is a Persian (Farsi) dataset for the Classification task, specifically focused on detecting the emotion "love" in user-chatbot conversations. It is part of the FaMTEB (Farsi Massive Text Embedding Benchmark). This dataset was synthetically generated using GPT-4o-mini and is a subset of the broader Synthetic Persian Chatbot Conversational Sentiment Analysis collection.
Language(s): Persian (Farsi)
Task(s):… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/synthetic-persian-chatbot-conversational-sentiment-analysis-love.flipkart_sentiment_analysisindonesian-financial-sentiment-analysischat-sentiment-analysis
A Sentiment Analsysis Dataset for Finetuning Large Models in Chat-style
More details can be found at https://github.com/l294265421/chat-sentiment-analysis
Supported Tasks
Aspect Term Extraction (ATE)
Opinion Term Extraction (OTE)
Aspect Term-Opinion Term Pair Extraction (AOPE)
Aspect term, Sentiment, Opinion term Triplet Extraction (ASOTE)
Aspect Category Detection (ACD)
Aspect Category-Sentiment Pair Extraction (ACSA)
Aspect-Category-Opinion-Sentiment (ACOS) Quadruple… See the full description on the dataset page: https://huggingface.co/datasets/yuncongli/chat-sentiment-analysis.Emotional_Sentiment_AnalysisEmotional Sentiment Analysis Dataset for LLaMA-2 Fine-tuning
(The formatted version can be directly used for fine tuning which contain only the formatted text, while the dataset.csv contain all the text, emotion, response and the formatted text)
This dataset contains conversational data for training and fine-tuning language models for emotional sentiment analysis and response generation. The dataset includes user inputs, their corresponding emotional states, and tailored chatbot responses… See the full description on the dataset page: https://huggingface.co/datasets/VaisakhKrishna/Emotional_Sentiment_Analysis.bitcoin-sentiment-analysisfinancial-sentiment-analysis-dataset
💰 Financial Sentiment Analysis Dataset
A dataset containing financial news headlines and social media posts labeled with sentiment.
Designed for:
Financial NLP research
Market sentiment analysis
Trading signal modeling
📊 Dataset Statistics
Train: 30,000 samples
Validation: 5,000 samples
Test: 5,000 samples
Total: 40,000 samples
📄 Data Format
{
"text": "Tesla stock surges after strong earnings report.",
"sentiment": "positive",
"source": "news"… See the full description on the dataset page: https://huggingface.co/datasets/Caplin43/financial-sentiment-analysis-dataset.SLC_Sentiment_AnalysisThis is information about the dataset
twitter-sentiment-analysis
🐦 Twitter Sentiment Analysis (bdstar/twitter-sentiment-analysis)
🧠 Overview
A refined and merged version of Twitter text sentiment datasets, providing a clean and well-balanced dataset for sentiment classification across three sentiment categories:positive, negative, and neutral.
This dataset is split into three parts — train, test, and validation — each sourced from highly reputable open datasets.It is designed for training, evaluating, and benchmarking NLP models for… See the full description on the dataset page: https://huggingface.co/datasets/akhiljoe143/twitter-sentiment-analysis.tweet-sentiment-analysis-from-kaggleindonesian-financial-sentiment-analysissentiment-analysis-sesentiment-analysis-UITsentiment_analysis_sharegpt_jsonreview_sentiment_analysis_prompt
