datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
twitter100m_tweets
Dataset Card for "twitter100m_tweets"
Dataset with tweets for this post.
DOI: 10.5281/zenodo.15086029
tweets_sample_2026
Tweets Sample 2026 — Newspaper-Vocabulary Reference Corpus
A large, deliberately untargeted sample of public posts from X/Twitter, collected via
Nitter by sweeping a 65,689-term newspaper vocabulary
rather than a topical keyword set.
It is built as a background / reference corpus: a baseline of "what was being said in
general" against which a topically targeted collection can be contrasted. It is the reference
arm of a narrative-detection study, not a curated dataset about any… See the full description on the dataset page: https://huggingface.co/datasets/SinclairSchneider/tweets_sample_2026.trump_tweetstweet-stock-synthetic-retrieval_beirThis is a copy of https://huggingface.co/datasets/jinaai/tweet-stock-synthetic-retrieval reformatted into the BEIR format. For any further information like license, please refer to the original dataset.
Disclaimer
This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/tweet-stock-synthetic-retrieval_beir.ArSAS_An_Arabic_Speech-Act_and_Sentiment_Corpus_of_Tweets
ArSAS: An Arabic Speech-Act and Sentiment Corpus of Tweets
Dataset Card for "ArSAS: An Arabic Speech-Act and Sentiment Corpus of Tweets"
Note About Sentiment_label_confidence
"Crowdflower provides a confidence score with each annotated tweet that represents the confidence in the quality of the label. For a three annotators per tweet setup, the confidence score would range between 0.3 and 1 according to two factors: 1) annotator quality level; and 2) agreement… See the full description on the dataset page: https://huggingface.co/datasets/Qanadil/ArSAS_An_Arabic_Speech-Act_and_Sentiment_Corpus_of_Tweets.disaster_tweetsDisaster_TweetsDLT-Tweets
DLT-Tweets
[Paper] •
[Code]
Dataset Description
Dataset Summary
DLT-Tweets is a large-scale corpus of social media posts related to Distributed Ledger Technology (DLT). This dataset is part of the larger DLT-Corpus collection, designed to support NLP research, social computing studies, and public discourse analysis in the DLT domain. It was introduced in the paper DLT-Corpus: A Large-Scale Text Collection for the Distributed Ledger Technology Domain.… See the full description on the dataset page: https://huggingface.co/datasets/ExponentialScience/DLT-Tweets.trump-tweetsThis is a clone of the Trump Tweet Kaggle dataset found here: https://www.kaggle.com/datasets/headsortails/trump-twitter-archive
COP26_Energy_Transition_Tweetstweets_correctiv_and_factcheckvaccines-tweets-ind-classification
VaccinesTweets_ind_Classification
Deduplicated copy of kornwtp/vaccines-tweets-ind-classification.
Splits
split
rows
train
4,056
Hate-Speech-Tweetstweets_pt_sentiment_analysis
Dataset Card for "tweets_pt_sentiment_analysis"
More Information needed
disaster_tweets
README.md
data:
train.csv
validation.csv
test.csv
drug-use-raw-tweets
Drug Use Raw Tweets
Full pool of tweets collected via GoXCrap and
Corpus Creator for the
drug-use-corpus project. This is the raw,
uncategorized collection: none of these tweets have been manually labeled as Positive/Negative for
drug-use content. It is published so other researchers can continue the manual categorization process,
extend it to other substances, or use it as a source pool for related tasks.
drug-use-corpus (the labeled EPB corpus used to train the models in this… See the full description on the dataset page: https://huggingface.co/datasets/lhbelfanti/drug-use-raw-tweets.portuguese-tweetselon-tweetsAppleVisionPro_Tweets
Apple Vision Pro Tweets Dataset
Overview
The Apple Vision Pro Tweets Dataset is a collection of tweets related to Apple Vision Pro from January 01 2024 to March 16 2024, scraped from X using the Twitter API. The dataset includes various attributes associated with each tweet, such as the tweet text, author information, engagement metrics, and metadata.
Content
id: Unique identifier for each tweet.
tweetText: The text content of the tweet.
tweetURL: URL link… See the full description on the dataset page: https://huggingface.co/datasets/divyasharma0795/AppleVisionPro_Tweets.tweets_about_german_politicians_jan_feb_2025_with_party_and_sentimentarabic_tweets_classification
Dataset Card for "arabic_tweets_classification"
More Information needed
COVID-19-vaccine-attitude-tweets
Dataset Card for COVID-19-vaccine-attitude-tweets
Dataset Summary
The dataset consists of 2564 manually annotated tweets related to COVID-19 vaccines. The dataset can be used to discover the attitude expressed in the tweet towards the subject of COVID-19 vaccines. Tweets are in English. The dataset was curated in such a way as to maximize the likelihood of tweets with a strong emotional tone. We have assumed the existence of three classes:
PRO (label 0): positive, the… See the full description on the dataset page: https://huggingface.co/datasets/webimmunization/COVID-19-vaccine-attitude-tweets.stock_market_tweets
Overview
This file contains over 1.7m public tweets about Apple, Amazon, Google, Microsoft and Tesla stocks, published between 01/01/2015 and 31/12/2019.
synthetic_social_persona_tweets
Synthetic Social Persona Tweets Dataset
This dataset contains synthetic social media posts generated by various language models.
This dataset is only meant to be used for quick and dirty experiments i.e. it's a toy dataset.
Every column/field in this dataset is generated by an LLM.
The code/prompts used to create this dataset can be found here.
The dataset was built to be used for some fine-tuning experiments with ModernBert for one of my blog posts/tutorials.
Each row in the… See the full description on the dataset page: https://huggingface.co/datasets/chrislevy/synthetic_social_persona_tweets.tweets_dataset_jan_feb_big_deduplicatedvaccines-tweets-ind-classificationNatural-Language-Processing-with-Disaster-Tweets-0.84033bitcoin-daily-raw-news-tweets-financials
Daily Bitcoin Multimodal & LLM-Augmented Financial Dataset
Dataset Description
This is a rich, time-series dataset designed for multimodal analysis and forecasting of the Bitcoin market. It aggregates a wide array of daily data from early 2015 to the end of 2022, with each row representing a single day.
The dataset combines several data dimensions:
Textual Data: Raw text from news articles, social media (tweets and Reddit), providing daily public discourse and sentiment.… See the full description on the dataset page: https://huggingface.co/datasets/tahamajs/bitcoin-daily-raw-news-tweets-financials.amazon-help-tweetsamazon-help-tweets-english
