CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01enryu43 /twitter100m_tweets Dataset Card for "twitter100m_tweets" Dataset with tweets for this post. DOI: 10.5281/zenodo.15086029 tabular10M<n<100M35 likes1.3k downloads2y agoHugging Face02SinclairSchneider /tweets_sample_2026 Tweets Sample 2026 — Newspaper-Vocabulary Reference Corpus A large, deliberately untargeted sample of public posts from X/Twitter, collected via Nitter by sweeping a 65,689-term newspaper vocabulary rather than a topical keyword set. It is built as a background / reference corpus: a baseline of "what was being said in general" against which a topically targeted collection can be contrasted. It is the reference arm of a narrative-detection study, not a curated dataset about any… See the full description on the dataset page: https://huggingface.co/datasets/SinclairSchneider/tweets_sample_2026.tabulartext-classification10M<n<100M0 likes522 downloads2mo agoHugging Face03rguo123 /trump_tweetstabular10K<n<100K1 likes454 downloads3y agoHugging Face04jinaai /tweet-stock-synthetic-retrieval_beirThis is a copy of https://huggingface.co/datasets/jinaai/tweet-stock-synthetic-retrieval reformatted into the BEIR format. For any further information like license, please refer to the original dataset. Disclaimer This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/tweet-stock-synthetic-retrieval_beir.image1K<n<10K0 likes447 downloads1y agoHugging Face05Qanadil /ArSAS_An_Arabic_Speech-Act_and_Sentiment_Corpus_of_Tweets ArSAS: An Arabic Speech-Act and Sentiment Corpus of Tweets Dataset Card for "ArSAS: An Arabic Speech-Act and Sentiment Corpus of Tweets" Note About Sentiment_label_confidence "Crowdflower provides a confidence score with each annotated tweet that represents the confidence in the quality of the label. For a three annotators per tweet setup, the confidence score would range between 0.3 and 1 according to two factors: 1) annotator quality level; and 2) agreement… See the full description on the dataset page: https://huggingface.co/datasets/Qanadil/ArSAS_An_Arabic_Speech-Act_and_Sentiment_Corpus_of_Tweets.tabulartext-classification10K<n<100K1 likes408 downloads2y agoHugging Face06venetis /disaster_tweetstabulartext-classification1K<n<10K3 likes347 downloads4y agoHugging Face07Sachinkelenjaguri /Disaster_Tweetstabular1K<n<10K1 likes196 downloads4y agoHugging Face08ExponentialScience /DLT-Tweets DLT-Tweets [Paper] • [Code] Dataset Description Dataset Summary DLT-Tweets is a large-scale corpus of social media posts related to Distributed Ledger Technology (DLT). This dataset is part of the larger DLT-Corpus collection, designed to support NLP research, social computing studies, and public discourse analysis in the DLT domain. It was introduced in the paper DLT-Corpus: A Large-Scale Text Collection for the Distributed Ledger Technology Domain.… See the full description on the dataset page: https://huggingface.co/datasets/ExponentialScience/DLT-Tweets.tabulartext-generation10M<n<100M0 likes192 downloads7mo agoHugging Face09fschlatt /trump-tweetsThis is a clone of the Trump Tweet Kaggle dataset found here: https://www.kaggle.com/datasets/headsortails/trump-twitter-archive tabular10K<n<100K6 likes150 downloads3y agoHugging Face10DSCI511G1 /COP26_Energy_Transition_Tweetstabular10K<n<100K3 likes145 downloads5y agoHugging Face11SinclairSchneider /tweets_correctiv_and_factchecktabular1M<n<10M0 likes103 downloads4mo agoHugging Face12suwaimyo /vaccines-tweets-ind-classification VaccinesTweets_ind_Classification Deduplicated copy of kornwtp/vaccines-tweets-ind-classification. Splits split rows train 4,056 tabular1K<n<10K0 likes83 downloads28d agoHugging Face13kaifahmad /Hate-Speech-Tweetstabular10K<n<100K0 likes75 downloads3y agoHugging Face14ricardo-filho /tweets_pt_sentiment_analysis Dataset Card for "tweets_pt_sentiment_analysis" More Information needed tabular100K<n<1M1 likes73 downloads4y agoHugging Face15VuduVations /disaster_tweets README.md data: train.csv validation.csv test.csv tabular10K<n<100K0 likes65 downloads3y agoHugging Face16lhbelfanti /drug-use-raw-tweets Drug Use Raw Tweets Full pool of tweets collected via GoXCrap and Corpus Creator for the drug-use-corpus project. This is the raw, uncategorized collection: none of these tweets have been manually labeled as Positive/Negative for drug-use content. It is published so other researchers can continue the manual categorization process, extend it to other substances, or use it as a source pool for related tasks. drug-use-corpus (the labeled EPB corpus used to train the models in this… See the full description on the dataset page: https://huggingface.co/datasets/lhbelfanti/drug-use-raw-tweets.tabulartext-classification100K<n<1M0 likes53 downloads1mo agoHugging Face17fpaulino /portuguese-tweetstabular100K<n<1M4 likes51 downloads4y agoHugging Face18dawood /elon-tweetstabularn<1K1 likes51 downloads4y agoHugging Face19divyasharma0795 /AppleVisionPro_Tweets Apple Vision Pro Tweets Dataset Overview The Apple Vision Pro Tweets Dataset is a collection of tweets related to Apple Vision Pro from January 01 2024 to March 16 2024, scraped from X using the Twitter API. The dataset includes various attributes associated with each tweet, such as the tweet text, author information, engagement metrics, and metadata. Content id: Unique identifier for each tweet. tweetText: The text content of the tweet. tweetURL: URL link… See the full description on the dataset page: https://huggingface.co/datasets/divyasharma0795/AppleVisionPro_Tweets.tabulartext-classification10K<n<100K13 likes49 downloads2y agoHugging Face20SinclairSchneider /tweets_about_german_politicians_jan_feb_2025_with_party_and_sentimenttabular100K<n<1M0 likes49 downloads8mo agoHugging Face21Abdelkareem /arabic_tweets_classification Dataset Card for "arabic_tweets_classification" More Information needed tabular10K<n<100K1 likes45 downloads3y agoHugging Face22webimmunization /COVID-19-vaccine-attitude-tweets Dataset Card for COVID-19-vaccine-attitude-tweets Dataset Summary The dataset consists of 2564 manually annotated tweets related to COVID-19 vaccines. The dataset can be used to discover the attitude expressed in the tweet towards the subject of COVID-19 vaccines. Tweets are in English. The dataset was curated in such a way as to maximize the likelihood of tweets with a strong emotional tone. We have assumed the existence of three classes: PRO (label 0): positive, the… See the full description on the dataset page: https://huggingface.co/datasets/webimmunization/COVID-19-vaccine-attitude-tweets.tabulartext-classification1K<n<10K2 likes44 downloads4y agoHugging Face23mjw /stock_market_tweets Overview This file contains over 1.7m public tweets about Apple, Amazon, Google, Microsoft and Tesla stocks, published between 01/01/2015 and 31/12/2019. tabular1M<n<10M21 likes42 downloads4y agoHugging Face24chrislevy /synthetic_social_persona_tweets Synthetic Social Persona Tweets Dataset This dataset contains synthetic social media posts generated by various language models. This dataset is only meant to be used for quick and dirty experiments i.e. it's a toy dataset. Every column/field in this dataset is generated by an LLM. The code/prompts used to create this dataset can be found here. The dataset was built to be used for some fine-tuning experiments with ModernBert for one of my blog posts/tutorials. Each row in the… See the full description on the dataset page: https://huggingface.co/datasets/chrislevy/synthetic_social_persona_tweets.tabulartext-classification1K<n<10K0 likes35 downloads2y agoHugging Face25SinclairSchneider /tweets_dataset_jan_feb_big_deduplicatedtabular10M<n<100M0 likes35 downloads5mo agoHugging Face26kornwtp /vaccines-tweets-ind-classificationtabular1K<n<10K0 likes33 downloads2y agoHugging Face27cuonguyenphu /Natural-Language-Processing-with-Disaster-Tweets-0.84033tabular1K<n<10K0 likes33 downloads6d agoHugging Face28tahamajs /bitcoin-daily-raw-news-tweets-financials Daily Bitcoin Multimodal & LLM-Augmented Financial Dataset Dataset Description This is a rich, time-series dataset designed for multimodal analysis and forecasting of the Bitcoin market. It aggregates a wide array of daily data from early 2015 to the end of 2022, with each row representing a single day. The dataset combines several data dimensions: Textual Data: Raw text from news articles, social media (tweets and Reddit), providing daily public discourse and sentiment.… See the full description on the dataset page: https://huggingface.co/datasets/tahamajs/bitcoin-daily-raw-news-tweets-financials.tabular1K<n<10K0 likes32 downloads1y agoHugging Face29PETROL-BOY /amazon-help-tweetstabular10K<n<100K0 likes32 downloads14d agoHugging Face30PETROL-BOY /amazon-help-tweets-englishtabular10K<n<100K0 likes30 downloads14d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.