datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
twitter-financial-news-sentiment
Dataset Description
The Twitter Financial News dataset is an English-language dataset containing an annotated corpus of finance-related tweets. This dataset is used to classify finance-related tweets for their sentiment.
The dataset holds 11,932 documents annotated with 3 labels:
sentiments = {
"LABEL_0": "Bearish",
"LABEL_1": "Bullish",
"LABEL_2": "Neutral"
}
The data was collected using the Twitter API. The current dataset supports the multi-class classification… See the full description on the dataset page: https://huggingface.co/datasets/zeroshot/twitter-financial-news-sentiment.twitter-financial-news-topic
Dataset Description
The Twitter Financial News dataset is an English-language dataset containing an annotated corpus of finance-related tweets. This dataset is used to classify finance-related tweets for their topic.
The dataset holds 21,107 documents annotated with 20 labels:
topics = {
"LABEL_0": "Analyst Update",
"LABEL_1": "Fed | Central Banks",
"LABEL_2": "Company | Product News",
"LABEL_3": "Treasuries | Corporate Debt",
"LABEL_4": "Dividend"… See the full description on the dataset page: https://huggingface.co/datasets/zeroshot/twitter-financial-news-topic.twitter-airline-sentiment
Dataset Card for Twitter US Airline Sentiment
Dataset Summary
This data originally came from Crowdflower's Data for Everyone library.
As the original source says,
A sentiment analysis job about the problems of each major U.S. airline. Twitter data was scraped from February of 2015 and contributors were asked to first classify positive, negative, and neutral tweets, followed by categorizing negative reasons (such as "late flight" or "rude service").
The data we're… See the full description on the dataset page: https://huggingface.co/datasets/osanseviero/twitter-airline-sentiment.pegos-twitter-streamhate_speech_twitter
Dataset Card for Dataset Name
The dataset is designed to analyze and address hate speech within online platforms. It consists of two sets: the training and testing sets. The two datasets have been labeled and categorized instances of hate speech into nine distinct categories.
Dataset Description
The dataset comprises three key features: tweets, labels (with hate speech denoted as 1 and non-hate speech as 0), and categories (behavior, class, disability, ethnicity, gender… See the full description on the dataset page: https://huggingface.co/datasets/thefrankhsu/hate_speech_twitter.twitter-trending-hashtags
Twitter/X Trending Hashtags (2020-2025)
A comprehensive dataset of trending hashtags on Twitter/X from 2020 to 2025, containing 12,036 unique trend entries across six years, capturing major world events, cultural moments, and viral phenomena.
📊 Dataset Description
This dataset captures trending hashtags from Twitter/X (formerly Twitter) by analyzing Wayback Machine snapshots of trends24.in, providing insights into breaking news, viral content, cultural moments, and… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/twitter-trending-hashtags.nlp_twitter_analysisTwitterHateSpeechtwitter_indonesia_sarcastic
Twitter Indonesia Sarcastic
Twitter Indonesia Sarcastic is a dataset intended for sarcasm detection in the Indonesian language. This dataset is introduced in Khotijah et al. (2020), whereby Indonesian tweets are collected and labeled as either sarcastic or non-sarcastic. We took the raw data, and performed several cleaning procedures such as: sentence order re-reversal, deduplication with minHash LSH, PII masking to remove usernames, hashtags, emails, URLs, and finally a random… See the full description on the dataset page: https://huggingface.co/datasets/w11wo/twitter_indonesia_sarcastic.ENGLISH_TWI_PARALLEL_TEXT
GhanaNLP Twi and English Parallel Data
Twi_to_English
• 1 MB • XLS
English_to_Twi
• 1 MB • XLS
The GhanaNLP Twi dataset contains sentence pairs in Twi and English, designed to support translation models between these two languages. Twi is a Ghanaian local language that lacks extensive digital resources, making this dataset useful for… See the full description on the dataset page: https://huggingface.co/datasets/Ghana-NLP/ENGLISH_TWI_PARALLEL_TEXT.twin-cities-public-records
Twin Cities public records, joined
25 datasets · 1,575,384 rows · free, CC BY 4.0 · mirrored from brickandmortar.dev
A city emits records constantly — parcels, recorded sales, assessments, permits, licences, inspections, 911 calls, cleanup sites, flood zones, federal loans, wages, census measures — and almost nobody joins them. These are the joined slices, published as files rather than as an API you have to ask for a key to. The join is the work; the data is free.
This is a… See the full description on the dataset page: https://huggingface.co/datasets/brickandmortar/twin-cities-public-records.indonesian-twitter-hate-speech-cleaned
Dataset Card for indonesian-twitter-hate-speech-cleaned
Dataset Summary
Cleaned Indonesian Twitter Hate Speech is a curated dataset consisting of Indonesian-language tweets labeled as either hate or neutral. The dataset was collected through a combination of direct scraping from Twitter and aggregation from multiple publicly available GitHub repositories.
The data has been cleaned to remove duplicates, irrelevant content, and non-textual noise, making it suitable for… See the full description on the dataset page: https://huggingface.co/datasets/haipradana/indonesian-twitter-hate-speech-cleaned.TwitterDuygulanguage:
tr
negatif 54%
pozitif 46%
reddit-blogspot-twittertwitter_disastertwitter-misinformation
Dataset Card for Twitter Misinformation Dataset
Dataset Description
Dataset Summary
This dataset is a compilation of several existing datasets focused on misinformation detection, disaster-related tweets, and fact-checking. It combines data from multiple sources to create a comprehensive dataset for training misinformation detection models. This dataset has been utilized in research studying backdoor attacks in textual content, notably in "Claim-Guided Textual… See the full description on the dataset page: https://huggingface.co/datasets/roupenminassian/twitter-misinformation.twitter-dataset-tesla
Dataset Card for Twitter Dataset: Tesla
Dataset Summary
This dataset contains all the Tweets regarding #Tesla or #tesla till 12/07/2022 (dd-mm-yyyy). It can be used for sentiment analysis research purpose or used in other NLP tasks or just for fun.
It contains 10,000 recent Tweets with the user ID, the hashtags used in the Tweets, and other important features.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information… See the full description on the dataset page: https://huggingface.co/datasets/hugginglearners/twitter-dataset-tesla.twinviews-13k
Dataset Card for TwinViews-13k
This dataset contains 13,855 pairs of left-leaning and right-leaning political statements matched by topic. The dataset was generated using GPT-3.5 Turbo and has been audited to ensure quality and ideological balance. It is designed to facilitate the study of political bias in reward models and language models, with a focus on the relationship between truthfulness and political views.
Dataset Details
Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/wwbrannon/twinviews-13k.SignedGraphs
Learning Stance Embeddings from Signed Social Graphs
This repo contains the datasets from our paper Learning Stance Embeddings from Signed Social Graphs.
[PDF]
[HuggingFace Datasets]
This work is licensed under a Creative Commons Attribution 4.0 International License.
Overview
A key challenge in social network analysis is understanding the position, or stance, of people in the graph on a large set of topics. In such social graphs, modeling (dis)agreement patterns… See the full description on the dataset page: https://huggingface.co/datasets/Twitter/SignedGraphs.TWI_ENGLISH_PARALLEL_TEXT
GhanaNLP Twi and English Parallel Data
Twi_to_English
• 1 MB • XLS
English_to_Twi
• 1 MB • XLS
The GhanaNLP Twi dataset contains sentence pairs in Twi and English, designed to support translation models between these two languages. Twi is a Ghanaian local language that lacks extensive digital resources, making this dataset useful for… See the full description on the dataset page: https://huggingface.co/datasets/Ghana-NLP/TWI_ENGLISH_PARALLEL_TEXT.decision-twin-v0-1
Decision Twin v0.1 encoder seed dataset
This dataset turns the user’s confirmed purchase and creative preferences into sentence-pair classification groups. Each context is phrased as a first-person request that a buying or creative assistant could receive.
Main files
File
Purpose
decision_twin_encoder.csv
Same rows in CSV form.
train.csv, validation.csv, test.csv
Group-safe 80/10/10 CSV splits, generated with seed 42.
decision_groups.jsonl
Group IDs… See the full description on the dataset page: https://huggingface.co/datasets/JohnGorri/decision-twin-v0-1.twitter-sentiment-meta-analysis
Twitter Sentiment Meta-Analysis Dataset
Dataset Description
This dataset contains sentiment analysis results for English tweets collected between September 2009 and January 2010. The tweets were processed and analyzed using 10 different sentiment classifiers, with the final sentiment score derived from principal component analysis (PCA).
Source Data
Original Data: Cheng-Caverlee-Lee Twitter Scrape (Sept 2009 - Jan 2010)
Number of Tweets: 138 690
Language:… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/twitter-sentiment-meta-analysis.twitter-hate-speech-en-240ksamplesThis dataset is a combination of the three datasets listed below:
tdavidson/hate_speech_offensive
LennardZuendorf/Dynamically-Generated-Hate-Speech-Dataset
ucberkeley-dlab/measuring-hate-speech
It has only two columns, "tweet" and "labels", and 242738 rows of uncleaned data.
twitter-human-botstwitter_sentiment_analysistwitter-financial-news-sentiment
Dataset Description
The Twitter Financial News dataset is an English-language dataset containing an annotated corpus of finance-related tweets. This dataset is used to classify finance-related tweets for their sentiment.
The dataset holds 11,932 documents annotated with 3 labels:
sentiments = {
"LABEL_0": "Bearish",
"LABEL_1": "Bullish",
"LABEL_2": "Neutral"
}
The data was collected using the Twitter API. The current dataset supports the multi-class… See the full description on the dataset page: https://huggingface.co/datasets/ROLEX-2007/twitter-financial-news-sentiment.large-twitter-tweets-sentiment
Dataset Card for "Large twitter tweets sentiment analysis"
Dataset Description
Dataset Summary
This dataset is a collection of tweets formatted in a tabular data structure, annotated for sentiment analysis.
Each tweet is associated with a sentiment label, with 1 indicating a Positive sentiment and 0 for a Negative sentiment.
Languages
The tweets in English.
Dataset Structure
Data Instances
An instance of the dataset includes… See the full description on the dataset page: https://huggingface.co/datasets/gxb912/large-twitter-tweets-sentiment.Twitter_Sinhala_Hate_Speechtwitter_bot_detectionindonesian-twitter-hate-speech-cleaned
Dataset Card for indonesian-twitter-hate-speech-cleaned
Dataset Summary
Cleaned Indonesian Twitter Hate Speech is a curated dataset consisting of Indonesian-language tweets labeled as either hate or neutral. The dataset was collected through a combination of direct scraping from Twitter and aggregation from multiple publicly available GitHub repositories.
The data has been cleaned to remove duplicates, irrelevant content, and non-textual noise, making it suitable… See the full description on the dataset page: https://huggingface.co/datasets/egdrga/indonesian-twitter-hate-speech-cleaned.
