datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TwitterHjerneRetrieval
TwitterHjerneRetrieval
An MTEB dataset
Massive Text Embedding Benchmark
Danish question asked on Twitter with the Hashtag #Twitterhjerne ('Twitter brain') and their corresponding answer.
Task category
t2t
Domains
Social, Written
Reference
https://huggingface.co/datasets/sorenmulli/da-hashtag-twitterhjerne
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/TwitterHjerneRetrieval.twittersemeval2015-pairclassification
TwitterSemEval2015
An MTEB dataset
Massive Text Embedding Benchmark
Paraphrase-Pairs of Tweets from the SemEval 2015 workshop.
Task category
t2t
Domains
Social, Written
Reference
https://alt.qcri.org/semeval2015/task1/
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["TwitterSemEval2015"])
evaluator = mteb.MTEB(task)
model = mteb.get_model(YOUR_MODEL)… See the full description on the dataset page: https://huggingface.co/datasets/mteb/twittersemeval2015-pairclassification.twitterurlcorpus-pairclassification
TwitterURLCorpus
An MTEB dataset
Massive Text Embedding Benchmark
Paraphrase-Pairs of Tweets.
Task category
t2t
Domains
Social, Written
Reference
https://languagenet.github.io/
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["TwitterURLCorpus"])
evaluator = mteb.MTEB(task)
model = mteb.get_model(YOUR_MODEL)
evaluator.run(model)
To learn more about how to run… See the full description on the dataset page: https://huggingface.co/datasets/mteb/twitterurlcorpus-pairclassification.twitter-financial-news-sentiment
Dataset Description
The Twitter Financial News dataset is an English-language dataset containing an annotated corpus of finance-related tweets. This dataset is used to classify finance-related tweets for their sentiment.
The dataset holds 11,932 documents annotated with 3 labels:
sentiments = {
"LABEL_0": "Bearish",
"LABEL_1": "Bullish",
"LABEL_2": "Neutral"
}
The data was collected using the Twitter API. The current dataset supports the multi-class classification… See the full description on the dataset page: https://huggingface.co/datasets/zeroshot/twitter-financial-news-sentiment.twitter-financial-news-topic
Dataset Description
The Twitter Financial News dataset is an English-language dataset containing an annotated corpus of finance-related tweets. This dataset is used to classify finance-related tweets for their topic.
The dataset holds 21,107 documents annotated with 20 labels:
topics = {
"LABEL_0": "Analyst Update",
"LABEL_1": "Fed | Central Banks",
"LABEL_2": "Company | Product News",
"LABEL_3": "Treasuries | Corporate Debt",
"LABEL_4": "Dividend"… See the full description on the dataset page: https://huggingface.co/datasets/zeroshot/twitter-financial-news-topic.twitter100m_tweets
Dataset Card for "twitter100m_tweets"
Dataset with tweets for this post.
DOI: 10.5281/zenodo.15086029
da-hashtag-twitterhjerne
Dataset Card for "da-hashtag-twitterhjerne"
Danish questions asked on Twitter using the Hashtag "#Twitterhjerne" ('Twitter brain') and their answers.
For each question tweet 2-6 answer tweets are included.
Further details can be found in Section 4.2.3 in the thesis.
Produced by: Søren Vejlgaard Holm under supervision of Lars Kai Hansen and Martin Carsten Nielsen.
Usable for: Question Answering Evaluation.
Contact: Søren Vejlgaard Holm at swiho@dtu.dk or swh@alvenir.ai.
twitter-airline-sentiment
Dataset Card for Twitter US Airline Sentiment
Dataset Summary
This data originally came from Crowdflower's Data for Everyone library.
As the original source says,
A sentiment analysis job about the problems of each major U.S. airline. Twitter data was scraped from February of 2015 and contributors were asked to first classify positive, negative, and neutral tweets, followed by categorizing negative reasons (such as "late flight" or "rude service").
The data we're… See the full description on the dataset page: https://huggingface.co/datasets/osanseviero/twitter-airline-sentiment.Arabic_Sentiment_Twitter_Corpus
Dataset Card for [Dataset Name]
Dataset Summary
[More Information Needed]
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation
Curation Rationale
[More Information Needed]
Source Data… See the full description on the dataset page: https://huggingface.co/datasets/arbml/Arabic_Sentiment_Twitter_Corpus.customer-support-on-twitter-conversationArabic_Sentiment_Twitter_Corpus
Dataset Card for "Arabic_Sentiment_Twitter_Corpus"
Source:
https://www.kaggle.com/datasets/mksaad/arabic-sentiment-twitter-corpus
AfriSenti-twitter-sentimentAfriSenti is the largest sentiment analysis benchmark dataset for under-represented African languages---covering 110,000+ annotated tweets in 14 African languages (Amharic, Algerian Arabic, Hausa, Igbo, Kinyarwanda, Moroccan Arabic, Mozambican Portuguese, Nigerian Pidgin, Oromo, Swahili, Tigrinya, Twi, Xitsonga, and yoruba).SAKE-TwitterTwitterArtistsviewer: true
Dataset Card for TwitterArtists (Pixel-Dust)
This dataset is a collection of art and media scraped from various artists and profiles across X (formerly Twitter) and Instagram. It is primarily focused on furry art and similar stylized content, intended for use in training or fine-tuning generative models.
Data Collection & Annotation
Source Data
The images were collected from social media profiles of numerous artists. While the bulk of… See the full description on the dataset page: https://huggingface.co/datasets/Pixel-Dust/TwitterArtists.ajgt_twitter_ar
Dataset Card for Arabic Jordanian General Tweets
Dataset Summary
Arabic Jordanian General Tweets (AJGT) Corpus consisted of 1,800 tweets annotated as positive and negative. Modern Standard Arabic (MSA) or Jordanian dialect.
Supported Tasks and Leaderboards
The dataset was published on this paper.
Languages
The dataset is based on Arabic.
Dataset Structure
Data Instances
A binary datset with with negative and positive… See the full description on the dataset page: https://huggingface.co/datasets/komari6/ajgt_twitter_ar.twitterCustomer_Support_on_Twittertwitter-chinese_av11-2025.12.23-2008047140334133610-zUpR1Z5yWO2iUL9S-part2hate_speech_twitter
Dataset Card for Dataset Name
The dataset is designed to analyze and address hate speech within online platforms. It consists of two sets: the training and testing sets. The two datasets have been labeled and categorized instances of hate speech into nine distinct categories.
Dataset Description
The dataset comprises three key features: tweets, labels (with hate speech denoted as 1 and non-hate speech as 0), and categories (behavior, class, disability, ethnicity, gender… See the full description on the dataset page: https://huggingface.co/datasets/thefrankhsu/hate_speech_twitter.pegos-twitter-streamTwitterHateSpeechtwitter-trending-hashtags
Twitter/X Trending Hashtags (2020-2025)
A comprehensive dataset of trending hashtags on Twitter/X from 2020 to 2025, containing 12,036 unique trend entries across six years, capturing major world events, cultural moments, and viral phenomena.
📊 Dataset Description
This dataset captures trending hashtags from Twitter/X (formerly Twitter) by analyzing Wayback Machine snapshots of trends24.in, providing insights into breaking news, viral content, cultural moments, and… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/twitter-trending-hashtags.twitter-year-splits
Dataset Card for "twitter-year-splits"
More Information needed
twitter-sentiment-analysisThe Twitter Sentiment Analysis Dataset contains 1,578,627 classified tweets, each row is marked as 1 for positive sentiment and 0 for negative sentiment.
The dataset is based on data from the following two sources:
University of Michigan Sentiment Analysis competition on Kaggle
Twitter Sentiment Corpus by Niek Sanders
Finally, I randomly selected a subset of them, applied a cleaning process, and divided them between the test and train subsets, keeping a balance between
the number of positive and negative tweets within each of these subsets.twitter-WheresMyMedia-2025.09.19-1968885375524393449-WDq9EYn8riCUFJY9-part1snapshot-twitter-2022-09-03
Snapshot Twitter
We no longer able to snapshot due to API changes.
description
minimum timestamp, 2022-04-17T16:30:07.000Z2.
maximum timestamp, 2022-09-03T09:23:52.000Z
7075025 rows
full attributes,
{
"datetime": "2022-04-18T05:57:04",
"datetime_gmt8": "2022-04-18T13:57:04",
"data_text": "kekal halal kak https://t.co/YHKqszqPnS",
"body": "kekal halal kak https://t.co/YHKqszqPnS",
"screen_name": "Luke_Sebastian2",
"followers_count": 10413,
"friends_count":… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/snapshot-twitter-2022-09-03.twitter_hate_speech_classificationtwitter_disasternlp_twitter_analysisNaijaSenti-TwitterNaijaSenti is the first large-scale human-annotated Twitter sentiment dataset for the four most widely spoken languages in Nigeria — Hausa, Igbo, Nigerian-Pidgin, and Yorùbá — consisting of around 30,000 annotated tweets per language, including a significant fraction of code-mixed tweets.
