datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
trump_tweetsdisaster_tweetsCOP26_Energy_Transition_TweetsDisaster_TweetsHate-Speech-Tweetsdrug-use-raw-tweets
Drug Use Raw Tweets
Full pool of tweets collected via GoXCrap and
Corpus Creator for the
drug-use-corpus project. This is the raw,
uncategorized collection: none of these tweets have been manually labeled as Positive/Negative for
drug-use content. It is published so other researchers can continue the manual categorization process,
extend it to other substances, or use it as a source pool for related tasks.
drug-use-corpus (the labeled EPB corpus used to train the models in this… See the full description on the dataset page: https://huggingface.co/datasets/lhbelfanti/drug-use-raw-tweets.elon-tweetsdisaster_tweets
README.md
data:
train.csv
validation.csv
test.csv
COVID-19-vaccine-attitude-tweets
Dataset Card for COVID-19-vaccine-attitude-tweets
Dataset Summary
The dataset consists of 2564 manually annotated tweets related to COVID-19 vaccines. The dataset can be used to discover the attitude expressed in the tweet towards the subject of COVID-19 vaccines. Tweets are in English. The dataset was curated in such a way as to maximize the likelihood of tweets with a strong emotional tone. We have assumed the existence of three classes:
PRO (label 0): positive, the… See the full description on the dataset page: https://huggingface.co/datasets/webimmunization/COVID-19-vaccine-attitude-tweets.stock_market_tweets
Overview
This file contains over 1.7m public tweets about Apple, Amazon, Google, Microsoft and Tesla stocks, published between 01/01/2015 and 31/12/2019.
Natural-Language-Processing-with-Disaster-Tweets-0.84033setfit-absa-tesla-tweetsitalian_long_covid_tweetsCOVID-19-conspiracy-theories-annotated-tweets
Dataset Card for Dataset Name
The dataset includes approximately 54,000 tweet IDs collected through the Twitter API between November 2019 and December 2021, along with annotations indicating whether a tweet supports a particular conspiracy theory.
Dataset Details
Curated by: Izabela Krysinska
Funded by [optional]: EEA Financial Mechanism 2014–2021. Project registration number: 2019/35/J/HS6/03498
Shared by [optional]: Izabela Krysinska
Language(s) (NLP): English… See the full description on the dataset page: https://huggingface.co/datasets/webimmunization/COVID-19-conspiracy-theories-annotated-tweets.Tweets-from-AKThis dataset contains Twitter information from AK92501
fifa-world-cup-2022-tweetsTweetscrime_tweets_in_portuguese
DataCrimeBR: Building a Dataset of Crimes Reported in Tweets in Brazil
This dataset contains 61.715 tweets related to possible crime reports, labeled with categories such as "Assalto", "Roubo", "Furto", "Assédio", "Segurança Pública", "Homicídio, and "Outros", along with sentiment analysis, toxicity analysis, and location identification.
A particular feature in the Portuguese language is that many words potentially related to crimes are used in non-criminal contexts, such as "O… See the full description on the dataset page: https://huggingface.co/datasets/miguelribeirokk/crime_tweets_in_portuguese.Conflict_TweetsThis dataset contains tweets related to the Israel-Palestine conflict from October 17, 2023, to December 17, 2023. It includes information on tweet IDs, links, text, date, likes, and comments, categorized into different ranges of like counts.
Dataset Details
Date Range: October 17, 2023 - December 17, 2023
Total Tweets: 15,478
Unique Tweets: 14,854
Data Description
The dataset consists of the following columns:
Column
Description
id
Unique identifier for the… See the full description on the dataset page: https://huggingface.co/datasets/Mehyaar/Conflict_Tweets.tweets_sentiment_analysisgpt3.5_tweetssynth-tweetsFitzwilliam-museum-tweetstweet_sentimentCorona_tweets.csvtweets_on_chatgptdf-tweets-HiQualProp3569-dehydratedtweetsbritish-museum-pompeii-live-tweetsfifa-world-cup-2022-tweets
