datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
trump_tweetsdisaster_tweetsDisaster_TweetsCOP26_Energy_Transition_TweetsTweet_Siniflandirma
References
alper bayram
stock-market-tweets-data
Stock Market Tweets Data
Overview
This dataset is the same as the Stock Market Tweets Data on IEEE by Bruno Taborda.
Data Description
This dataset contains 943,672 tweets collected between April 9 and July 16, 2020, using the S&P 500 tag (#SPX500), the references to the top 25 companies in the S&P 500 index, and the Bloomberg tag (#stocks).
Dataset Structure
created_at: The exact time this tweet was posted.
text: The text of the tweet, providing… See the full description on the dataset page: https://huggingface.co/datasets/StephanAkkerman/stock-market-tweets-data.financial-tweets-crypto
Financial Tweets - Cryptocurrency
This dataset is part of the scraped financial tweets that I collected from a variety of financial influencers on Twitter, all the datasets can be found here:
Crypto: https://huggingface.co/datasets/StephanAkkerman/financial-tweets-crypto
Stocks (and forex): https://huggingface.co/datasets/StephanAkkerman/financial-tweets-stocks
Other (Tweet without cash tags): https://huggingface.co/datasets/StephanAkkerman/financial-tweets-other
Data… See the full description on the dataset page: https://huggingface.co/datasets/StephanAkkerman/financial-tweets-crypto.tweets_cleancrypto-stock-tweets
Crypto & Stock Tweets
Overview
This dataset is a combination of publically available financial tweets.
Datset Size
Stock Tweets: 2,624,314
Crypto Tweets: 5,748,725
Bitcoin Tweets: 4,820,915
Sources
This dataset is a combination of data from various reputable sources, each contributing a unique perspective on financial tweets:
Stock Market Tweets Data: 923,673 rows of stock tweets
Stock Market Tweets: 1,700,641 rows of stock tweets
Crypto Tweets:… See the full description on the dataset page: https://huggingface.co/datasets/StephanAkkerman/crypto-stock-tweets.Gulf-Arabic-Tweets-2018-2020
Dataset Summary
This is a pre-processed (cleaned) Twitter Gulf Arabic dialect 2018-2020 dataset. Pleasre refer to the source, and data cleaning code and algorithm Github.
Languages
Arabic
Source Data
Twitter
hindi-english-code-mixed-tweets-sentimentHate-Speech-Tweetsdisaster_tweets
README.md
data:
train.csv
validation.csv
test.csv
financial-tweets-stocksdrug-use-raw-tweets
Drug Use Raw Tweets
Full pool of tweets collected via GoXCrap and
Corpus Creator for the
drug-use-corpus project. This is the raw,
uncategorized collection: none of these tweets have been manually labeled as Positive/Negative for
drug-use content. It is published so other researchers can continue the manual categorization process,
extend it to other substances, or use it as a source pool for related tasks.
drug-use-corpus (the labeled EPB corpus used to train the models in this… See the full description on the dataset page: https://huggingface.co/datasets/lhbelfanti/drug-use-raw-tweets.elon-tweetsfinancial-tweets
Financial Tweets
This dataset is a comprehensive collection of all the tweets from my Discord bot that keeps track of financial influencers on Twitter.
The data includes a variety of information, such as the tweet and the price of the tickers in that tweet at the time of posting.
This dataset can be used for a variety of tasks, such as sentiment analysis and masked language modelling (MLM).
We used this dataset for training our FinTwitBERT model.
Overview
This… See the full description on the dataset page: https://huggingface.co/datasets/StephanAkkerman/financial-tweets.large-twitter-tweets-sentiment
Dataset Card for "Large twitter tweets sentiment analysis"
Dataset Description
Dataset Summary
This dataset is a collection of tweets formatted in a tabular data structure, annotated for sentiment analysis.
Each tweet is associated with a sentiment label, with 1 indicating a Positive sentiment and 0 for a Negative sentiment.
Languages
The tweets in English.
Dataset Structure
Data Instances
An instance of the dataset includes… See the full description on the dataset page: https://huggingface.co/datasets/gxb912/large-twitter-tweets-sentiment.stock_market_tweets
Overview
This file contains over 1.7m public tweets about Apple, Amazon, Google, Microsoft and Tesla stocks, published between 01/01/2015 and 31/12/2019.
tweets-turkishcyberbullying_tweets.csv
Cyberbullying Tweets Dataset
Overview
This dataset contains labeled tweet data used for training and evaluating text classification models to identify and mitigate digital harassment and online toxic content.
Dataset Structure
tweet_text: Raw text content extracted from tweets.
cyberbullying_type: Corresponding label/category (e.g., gender, religion, ethnicity, age, or non-cyberbullying).
How to Load in Python
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/poorvanshi04/cyberbullying_tweets.csv.indian_railways_complaint_tweetsThis dataset contains roughly 4500 tweets mentioning the Indian Railways. It was created to help researchers and developers train natural language processing models for classifying passenger feedback, analyzing sentiment, and categorizing railway complaints. The data captures genuine passenger experiences including urgent grievances, infrastructure problems, news sharing, and general appreciation.
Structure of dataset:
The dataset is structured in a simple tabular format with three… See the full description on the dataset page: https://huggingface.co/datasets/faizmubeen/indian_railways_complaint_tweets.tweets_about_fireThis dataset is human-annotated and compiled by me from various sources, predominantly tweets.
Product_Tweets_Datasettweets_emotions_elections_colombiasetfit-absa-tesla-tweetsCOVID-19-conspiracy-theories-tweets
Dataset Summary
This dataset consists of 6591 tweets generated by GPT-3.5 model. The tweets are juxtaposed with a conspiracy theory related to COVID-19 pandemic. Each item consists of a label that represents the item's output class. The possible labels are support/deny/neutral.
support: the tweet suggests support for the conspiracy theory
deny: the tweet contradicts the conspiracy theory
neutral: the tweet is mostly informative, and does not show emotions against the conspiracy… See the full description on the dataset page: https://huggingface.co/datasets/webimmunization/COVID-19-conspiracy-theories-tweets.italian_long_covid_tweetsNatural_disaster_tweetsstock-market-tweets-data
Stock Market Tweets Data
Overview
This dataset is the same as the Stock Market Tweets Data on IEEE by Bruno Taborda.
Data Description
This dataset contains 943,672 tweets collected between April 9 and July 16, 2020, using the S&P 500 tag (#SPX500), the references to the top 25 companies in the S&P 500 index, and the Bloomberg tag (#stocks).
Dataset Structure
created_at: The exact time this tweet was posted.
text: The text of the tweet, providing… See the full description on the dataset page: https://huggingface.co/datasets/VeeraThakshith/stock-market-tweets-data.
