datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tweet_sentiment_extraction
TweetSentimentExtractionClassification
An MTEB dataset
Massive Text Embedding Benchmark
Task category
t2c
Domains
Social, Written
Reference
https://www.kaggle.com/competitions/tweet-sentiment-extraction/overview
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["TweetSentimentExtractionClassification"])
evaluator = mteb.MTEB(task)
model =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/tweet_sentiment_extraction.senator-tweetsangry-tweets
Dataset Card for AngryTweets
Dataset Summary
This dataset consists of anonymised Danish Twitter data that has been annotated for sentiment analysis through crowd-sourcing. All credits go to the authors of the following paper, who created the dataset:
Pauli, Amalie Brogaard, et al. "DaNLP: An open-source toolkit for Danish Natural Language Processing." Proceedings of the 23rd Nordic Conference on Computational Linguistics (NoDaLiDa). 2021
Supported Tasks and… See the full description on the dataset page: https://huggingface.co/datasets/DDSC/angry-tweets.political-leaning-tweets-100k
🗳️ political-leaning-tweets-100k
Châtelet AI presents a 100,000+ dataset of tweets labelled for political leaning: neutral, liberal, conservative.Labels are machine-generated using a SOTA thinking-enabled LLM. The dataset is intended for research on political language modelling, ideology detection, robustness, and safety evaluation.
📦 Dataset Card
Name: chatelet/political-leaning-tweets-100k
Publisher: Châtelet AI
Licence: MIT with additional restrctions against… See the full description on the dataset page: https://huggingface.co/datasets/chatelet/political-leaning-tweets-100k.spanish-tweets
spanish-tweets
A big corpus of tweets for pretraining embeddings and language models
Dataset Summary
A big dataset of (mostly) Spanish tweets for pre-training language models (or other representations).
Supported Tasks and Leaderboards
Language Modeling
Languages
Mostly Spanish, but some Portuguese, English, and other languages.
Dataset Structure
Data Fields
tweet_id: id of the tweet
user_id: id of the user
text:… See the full description on the dataset page: https://huggingface.co/datasets/pysentimiento/spanish-tweets.twitter100m_tweets
Dataset Card for "twitter100m_tweets"
Dataset with tweets for this post.
DOI: 10.5281/zenodo.15086029
tweet_sentiment_multilingual
Dataset Card for cardiffnlp/tweet_sentiment_multilingual
Dataset Summary
Tweet Sentiment Multilingual consists of sentiment analysis dataset on Twitter in 8 different lagnuages.
arabic
english
french
german
hindi
italian
portuguese
spanish
Supported Tasks and Leaderboards
text_classification: The dataset can be trained using a SentenceClassification model from HuggingFace transformers.
Dataset Structure
Data Instances
An instance from… See the full description on the dataset page: https://huggingface.co/datasets/cardiffnlp/tweet_sentiment_multilingual.tweets_hate_speech_detection
Dataset Card for Tweets Hate Speech Detection
Dataset Summary
The objective of this task is to detect hate speech in tweets. For the sake of simplicity, we say a tweet contains hate speech if it has a racist or sexist sentiment associated with it. So, the task is to classify racist or sexist tweets from other tweets.
Formally, given a training sample of tweets and labels, where label ‘1’ denotes the tweet is racist/sexist and label ‘0’ denotes the tweet is not… See the full description on the dataset page: https://huggingface.co/datasets/tweets-hate-speech-detection/tweets_hate_speech_detection.tweet_sentiment_multilingual
TweetSentimentClassification
An MTEB dataset
Massive Text Embedding Benchmark
A multilingual Sentiment Analysis dataset consisting of tweets in 8 different languages.
Task category
t2c
Domains
Social, Written
Reference
https://aclanthology.org/2022.lrec-1.27
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["TweetSentimentClassification"])
evaluator =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/tweet_sentiment_multilingual.tweetsentbr_fewshot
tweetSentBR (Few-shot)
This dataset is a subset of the tweetSentBR, it contains only 75 samples from the training set and all 2.000+ instances of the test set.This is meant for evaluating language models in a few-shot setting on the 🚀 Open Portuguese LLM Leaderboard with the portuguese fork of the Eleuther AI Language Model Evaluation Harness
For the complete dataset with 15.000+ annotated tweets go to https://bitbucket.org/HBrum/tweetsentbr or contact the paper authors:… See the full description on the dataset page: https://huggingface.co/datasets/eduagarcia/tweetsentbr_fewshot.arabic_tweets_dialectstweets_sample_2026
Tweets Sample 2026 — Newspaper-Vocabulary Reference Corpus
A large, deliberately untargeted sample of public posts from X/Twitter, collected via
Nitter by sweeping a 65,689-term newspaper vocabulary
rather than a topical keyword set.
It is built as a background / reference corpus: a baseline of "what was being said in
general" against which a topically targeted collection can be contrasted. It is the reference
arm of a narrative-detection study, not a curated dataset about any… See the full description on the dataset page: https://huggingface.co/datasets/SinclairSchneider/tweets_sample_2026.tweet-stock-synthetic-retrieval_beirThis is a copy of https://huggingface.co/datasets/jinaai/tweet-stock-synthetic-retrieval reformatted into the BEIR format. For any further information like license, please refer to the original dataset.
Disclaimer
This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/tweet-stock-synthetic-retrieval_beir.trump_tweetsArabic-Tweets
Dataset Card for Dataset Arabic-Tweets
Dataset Summary
This dataset has been collected from twitter which is more than 41 GB of clean data of Arabic Tweets with nearly 4-billion Arabic words (12-million unique Arabic words).
Languages
Arabic
Source Data
Twitter
Example on data loading using streaming:
from datasets import load_dataset
dataset = load_dataset("pain/Arabic-Tweets",split='train', streaming=True)
print(next(iter(dataset)))… See the full description on the dataset page: https://huggingface.co/datasets/pain/Arabic-Tweets.ArSAS_An_Arabic_Speech-Act_and_Sentiment_Corpus_of_Tweets
ArSAS: An Arabic Speech-Act and Sentiment Corpus of Tweets
Dataset Card for "ArSAS: An Arabic Speech-Act and Sentiment Corpus of Tweets"
Note About Sentiment_label_confidence
"Crowdflower provides a confidence score with each annotated tweet that represents the confidence in the quality of the label. For a three annotators per tweet setup, the confidence score would range between 0.3 and 1 according to two factors: 1) annotator quality level; and 2) agreement… See the full description on the dataset page: https://huggingface.co/datasets/Qanadil/ArSAS_An_Arabic_Speech-Act_and_Sentiment_Corpus_of_Tweets.tweet_sentiment_extraction
Tweet Sentiment Extraction
Source: https://www.kaggle.com/c/tweet-sentiment-extraction/data
financial-tweets-sentiment
Financial Sentiment Analysis Dataset
Overview
This dataset is a comprehensive collection of tweets focused on financial topics, meticulously curated to assist in sentiment analysis in the domain of finance and stock markets. It serves as a valuable resource for training machine learning models to understand and predict sentiment trends based on social media discourse, particularly within the financial sector.
Data Description
The dataset comprises tweets… See the full description on the dataset page: https://huggingface.co/datasets/TimKoornstra/financial-tweets-sentiment.ai-tweetsmo-customer-support-tweets-945k
Customer Support on Twitter Dataset 945k
Dataset Description
Context
This dataset provides a large corpus of real-world English conversations between consumers and customer support agents on Twitter, designed to drive innovation in Natural Language Processing (NLP) by providing data that better matches the actual language used in contemporary customer support interactions.
Content
Initially, the data included complex threads of conversations… See the full description on the dataset page: https://huggingface.co/datasets/MohammadOthman/mo-customer-support-tweets-945k.ASTD_Arabic_Sentiment_Tweets_Dataset
Citation
(https://aclanthology.org/D15-1299/)
disaster_tweetsDLT-Tweets
DLT-Tweets
[Paper] •
[Code]
Dataset Description
Dataset Summary
DLT-Tweets is a large-scale corpus of social media posts related to Distributed Ledger Technology (DLT). This dataset is part of the larger DLT-Corpus collection, designed to support NLP research, social computing studies, and public discourse analysis in the DLT domain. It was introduced in the paper DLT-Corpus: A Large-Scale Text Collection for the Distributed Ledger Technology Domain.… See the full description on the dataset page: https://huggingface.co/datasets/ExponentialScience/DLT-Tweets.covid_tweets_japanese
Dataset Card for COVID-19 日本語Twitterデータセット (COVID-19 Japanese Twitter Dataset)
Dataset Summary
53,640 Japanese tweets with annotation if a tweet is related to COVID-19 or not. The annotation is by majority decision by 5 - 10 crowd workers. Target tweets include "COVID" or "コロナ". The period of the tweets is from around January 2020 to around June 2020. The original tweets are not contained. Please use Twitter API to get them, for example.
Supported Tasks and… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/covid_tweets_japanese.tweets_ar_en_parallel Twitter users often post parallel tweets—tweets that contain the same content but are
written in different languages. Parallel tweets can be an important resource for developing
machine translation (MT) systems among other natural language processing (NLP) tasks. This
resource is a result of a generic method for collecting parallel tweets. Using the method,
we compiled a bilingual corpus of English-Arabic parallel tweets and a list of Twitter accounts
who post English-Arabic tweets regularly. Additionally, we annotate a subset of Twitter accounts
with their countries of origin and topic of interest, which provides insights about the population
who post parallel tweets.trump-tweetsThis is a clone of the Trump Tweet Kaggle dataset found here: https://www.kaggle.com/datasets/headsortails/trump-twitter-archive
COP26_Energy_Transition_TweetsTweet_Siniflandirma
References
alper bayram
stock-market-tweets-data
Stock Market Tweets Data
Overview
This dataset is the same as the Stock Market Tweets Data on IEEE by Bruno Taborda.
Data Description
This dataset contains 943,672 tweets collected between April 9 and July 16, 2020, using the S&P 500 tag (#SPX500), the references to the top 25 companies in the S&P 500 index, and the Bloomberg tag (#stocks).
Dataset Structure
created_at: The exact time this tweet was posted.
text: The text of the tweet, providing… See the full description on the dataset page: https://huggingface.co/datasets/StephanAkkerman/stock-market-tweets-data.antisemitic-tweets
