datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
twitter-airline-sentiment
Dataset Card for Twitter US Airline Sentiment
Dataset Summary
This data originally came from Crowdflower's Data for Everyone library.
As the original source says,
A sentiment analysis job about the problems of each major U.S. airline. Twitter data was scraped from February of 2015 and contributors were asked to first classify positive, negative, and neutral tweets, followed by categorizing negative reasons (such as "late flight" or "rude service").
The data we're… See the full description on the dataset page: https://huggingface.co/datasets/osanseviero/twitter-airline-sentiment.SEW-TWIST
SEW-TWIST G1 Teleoperation Dataset
This dataset contains offline teleoperation trajectories for the Unitree G1 humanoid robot generated using the SEW-MIMIC controller [1] and LaFAN1 BVH motion capture data [2].
The dataset was generated by replaying BVH motion capture sequences through a MuJoCo simulation of the G1 robot and logging the resulting robot state trajectories in a format compatible with TWIST-style imitation learning pipelines [3].
Each trajectory is stored as a .pkl… See the full description on the dataset page: https://huggingface.co/datasets/introvoyz042/SEW-TWIST.pegos-twitter-streamSEW-TWIST
SEW-TWIST G1 Teleoperation Dataset
This dataset contains offline teleoperation trajectories for the Unitree G1 humanoid robot generated using the SEW-MIMIC controller [1] and LaFAN1 BVH motion capture data [2].
The dataset was generated by replaying BVH motion capture sequences through a MuJoCo simulation of the G1 robot and logging the resulting robot state trajectories in a format compatible with TWIST-style imitation learning pipelines [3].
Each trajectory is stored as a .pkl… See the full description on the dataset page: https://huggingface.co/datasets/pshinde612/SEW-TWIST.twitter-trending-hashtags
Twitter/X Trending Hashtags (2020-2025)
A comprehensive dataset of trending hashtags on Twitter/X from 2020 to 2025, containing 12,036 unique trend entries across six years, capturing major world events, cultural moments, and viral phenomena.
📊 Dataset Description
This dataset captures trending hashtags from Twitter/X (formerly Twitter) by analyzing Wayback Machine snapshots of trends24.in, providing insights into breaking news, viral content, cultural moments, and… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/twitter-trending-hashtags.TwitterHateSpeechtwin-cities-public-records
Twin Cities public records, joined
25 datasets · 1,575,384 rows · free, CC BY 4.0 · mirrored from brickandmortar.dev
A city emits records constantly — parcels, recorded sales, assessments, permits, licences, inspections, 911 calls, cleanup sites, flood zones, federal loans, wages, census measures — and almost nobody joins them. These are the joined slices, published as files rather than as an API you have to ask for a key to. The join is the work; the data is free.
This is a… See the full description on the dataset page: https://huggingface.co/datasets/brickandmortar/twin-cities-public-records.TwitterFaveGraph
MiCRO: Multi-interest Candidate Retrieval Online
This repo contains the TwitterFaveGraph dataset from our paper MiCRO: Multi-interest Candidate Retrieval Online.
[PDF]
[HuggingFace Datasets]
This work is licensed under a Creative Commons Attribution 4.0 International License.
TwitterFaveGraph
TwitterFaveGraph is a bipartite directed graph of user nodes to Tweet nodes where an edge represents a "fave" engagement. Each edge is binned into predetermined time chunks which… See the full description on the dataset page: https://huggingface.co/datasets/Twitter/TwitterFaveGraph.twitter_disastertwitter-misinformation
Dataset Card for Twitter Misinformation Dataset
Dataset Description
Dataset Summary
This dataset is a compilation of several existing datasets focused on misinformation detection, disaster-related tweets, and fact-checking. It combines data from multiple sources to create a comprehensive dataset for training misinformation detection models. This dataset has been utilized in research studying backdoor attacks in textual content, notably in "Claim-Guided Textual… See the full description on the dataset page: https://huggingface.co/datasets/roupenminassian/twitter-misinformation.TwitterFollowGraph
kNN-Embed: Locally Smoothed Embedding Mixtures For Multi-interest Candidate Retrieval
This repo contains the TwitterFaveGraph dataset from our paper kNN-Embed: Locally Smoothed Embedding Mixtures For Multi-interest Candidate Retrieval.
[PDF]
[HuggingFace Datasets]
This work is licensed under a Creative Commons Attribution 4.0 International License.
TwitterFollowGraph
TwitterFollowGraph is a bipartite directed graph of users (consumer) nodes to author (producer) nodes… See the full description on the dataset page: https://huggingface.co/datasets/Twitter/TwitterFollowGraph.twitter-dataset-tesla
Dataset Card for Twitter Dataset: Tesla
Dataset Summary
This dataset contains all the Tweets regarding #Tesla or #tesla till 12/07/2022 (dd-mm-yyyy). It can be used for sentiment analysis research purpose or used in other NLP tasks or just for fun.
It contains 10,000 recent Tweets with the user ID, the hashtags used in the Tweets, and other important features.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information… See the full description on the dataset page: https://huggingface.co/datasets/hugginglearners/twitter-dataset-tesla.SignedGraphs
Learning Stance Embeddings from Signed Social Graphs
This repo contains the datasets from our paper Learning Stance Embeddings from Signed Social Graphs.
[PDF]
[HuggingFace Datasets]
This work is licensed under a Creative Commons Attribution 4.0 International License.
Overview
A key challenge in social network analysis is understanding the position, or stance, of people in the graph on a large set of topics. In such social graphs, modeling (dis)agreement patterns… See the full description on the dataset page: https://huggingface.co/datasets/Twitter/SignedGraphs.twitter-sentiment-meta-analysis
Twitter Sentiment Meta-Analysis Dataset
Dataset Description
This dataset contains sentiment analysis results for English tweets collected between September 2009 and January 2010. The tweets were processed and analyzed using 10 different sentiment classifiers, with the final sentiment score derived from principal component analysis (PCA).
Source Data
Original Data: Cheng-Caverlee-Lee Twitter Scrape (Sept 2009 - Jan 2010)
Number of Tweets: 138 690
Language:… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/twitter-sentiment-meta-analysis.twitter-human-botstwitter_bot_detectionforge-pump-digital-twin-synthetic
Forge Pump Digital Twin — Synthetic
A deterministic, clean-room tabular baseline for pump surrogate modelling, edge-runtime conformance, and advisory anomaly examples. It contains 40,000 synthetic rows split into 28,000 train, 6,000 validation, and 6,000 test rows.
This dataset contains no plant telemetry, customer data, equipment identifiers, CAD, BOMs, nameplates, vendor curves, or values copied from a private repository. Every constant is an illustrative engineering proxy. It… See the full description on the dataset page: https://huggingface.co/datasets/sankalpsthakur/forge-pump-digital-twin-synthetic.Twitter-COVID-19General description:
This dataset comprisses a set of tweets crawled during the COVID-19 pandemic (from March 2020 to June 2021). Tweets are located in two different regions: Spain and USA. This adds value to the collection, as it contains data in two languages.
This data was used as part of a broader study that aimed to determine the evolution of different personality traits and disorders during the pandemic. Thus, weak labels for different dimensions, such as sentiment, personality… See the full description on the dataset page: https://huggingface.co/datasets/citiusLTL/Twitter-COVID-19.twitter_racism_datasettwittersentiment-llama-3.1-405B-labelsFiltered and processed subset of mteb/tweet_sentiment_extraction
First 5000 entries were gathered for the train subset, and then 5001-6000 for test.
Blanks were removed from this subset, and further filtered to remove innapropriate content via Llama 3.1 405B's inherent harmful/explicit content flagging during the below label processing.
This results in a split ofTrain: 4992Test: 998
Original labels have been kept, and further labels have been generated using Llama 3.1 405B, via the prompt:… See the full description on the dataset page: https://huggingface.co/datasets/AdamLucek/twittersentiment-llama-3.1-405B-labels.Twitter-Conversations-Sentiment-Dataset
Twitter Sentiment Dataset
Sample English-only tweet sentiment dataset. Each row represents a single tweet with anonymized text and conversation structure.
This is a sample dataset. To access the full version or request any custom dataset tailored to your needs, contact DataHive at contact@datahive.ai.
Files Included
dataset.csv – tweets data
What’s included
Anonymized tweet text
Conversation linkage via root_id and parent_id
3-class sentiment label (positive… See the full description on the dataset page: https://huggingface.co/datasets/datahiveai/Twitter-Conversations-Sentiment-Dataset.twitter-nlpaayushmishra1512_twitchdata
Top Streamers on Twitch
This contains data of Top 1000 Streamers from past year.
Dataset Info
Source: Kaggle
Original Size: 0.04 MB
Kaggle Downloads: 14,390
Files: 1
Files
twitchdata-update.csv
Mirrored from Kaggle
twitteremoTwitterEmo 1.0 dataset.
Full description available in the following publication:
Bogdanowicz, S., Cwynar, H., Zwierzchowska, A., Klamra, C., Kieraś, W., Kobyliński, Ł. (2023). TwitterEmo: Annotating Emotions and Sentiment in Polish Twitter. In: Mikyška, J., de Mulatier, C., Paszynski, M., Krzhizhanovskaya, V.V., Dongarra, J.J., Sloot, P.M. (eds) Computational Science – ICCS 2023. ICCS 2023. Lecture Notes in Computer Science, vol 14074. Springer, Cham.… See the full description on the dataset page: https://huggingface.co/datasets/clarin-pl/twitteremo.Social_Media_And_Twitter_Mental_Health_Datasettwice_kr_financial_mmlu_cls
FinancialMMLU-CLS-ko
Multiple-choice questions, where a question and answer choices are provided to find the correct answer.
An open dataset generated and verified by GPT, based on financial public websites and Wikipedia.
Utilizing the open dataset allganize/financial-mmlu-ko (original source: public websites, Wikipedia).
twitter_strawmantwitter_airlines_pos_neg_smallajgt_twitter_artwice_kr_financial_mcqa_cls
FinancialMCQA-CLS-ko
Multiple-choice questions, where a question and answer choices are provided to find the correct answer.
Utilizing the open dataset FINNUMBER/QA_Instruction (original source: public websites, Wikipedia).
