CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mteb /tweet_sentiment_extraction TweetSentimentExtractionClassification An MTEB dataset Massive Text Embedding Benchmark Task category t2c Domains Social, Written Reference https://www.kaggle.com/competitions/tweet-sentiment-extraction/overview How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_tasks(["TweetSentimentExtractionClassification"]) evaluator = mteb.MTEB(task) model =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/tweet_sentiment_extraction.texttext-classification10K<n<100K38 likes5.6k downloads1y agoHugging Face02m-newhauser /senator-tweetstext10K<n<100K6 likes5k downloads3y agoHugging Face03DDSC /angry-tweets Dataset Card for AngryTweets Dataset Summary This dataset consists of anonymised Danish Twitter data that has been annotated for sentiment analysis through crowd-sourcing. All credits go to the authors of the following paper, who created the dataset: Pauli, Amalie Brogaard, et al. "DaNLP: An open-source toolkit for Danish Natural Language Processing." Proceedings of the 23rd Nordic Conference on Computational Linguistics (NoDaLiDa). 2021 Supported Tasks and… See the full description on the dataset page: https://huggingface.co/datasets/DDSC/angry-tweets.texttext-classification1K<n<10K4 likes2.6k downloads3y agoHugging Face04chatelet /political-leaning-tweets-100k 🗳️ political-leaning-tweets-100k Châtelet AI presents a 100,000+ dataset of tweets labelled for political leaning: neutral, liberal, conservative.Labels are machine-generated using a SOTA thinking-enabled LLM. The dataset is intended for research on political language modelling, ideology detection, robustness, and safety evaluation. 📦 Dataset Card Name: chatelet/political-leaning-tweets-100k Publisher: Châtelet AI Licence: MIT with additional restrctions against… See the full description on the dataset page: https://huggingface.co/datasets/chatelet/political-leaning-tweets-100k.texttext-classification100K<n<1M2 likes1.8k downloads1y agoHugging Face05pysentimiento /spanish-tweets spanish-tweets A big corpus of tweets for pretraining embeddings and language models Dataset Summary A big dataset of (mostly) Spanish tweets for pre-training language models (or other representations). Supported Tasks and Leaderboards Language Modeling Languages Mostly Spanish, but some Portuguese, English, and other languages. Dataset Structure Data Fields tweet_id: id of the tweet user_id: id of the user text:… See the full description on the dataset page: https://huggingface.co/datasets/pysentimiento/spanish-tweets.text100M<n<1B14 likes1.6k downloads3y agoHugging Face06enryu43 /twitter100m_tweets Dataset Card for "twitter100m_tweets" Dataset with tweets for this post. DOI: 10.5281/zenodo.15086029 tabular10M<n<100M35 likes1.3k downloads1y agoHugging Face07cardiffnlp /tweet_sentiment_multilingual Dataset Card for cardiffnlp/tweet_sentiment_multilingual Dataset Summary Tweet Sentiment Multilingual consists of sentiment analysis dataset on Twitter in 8 different lagnuages. arabic english french german hindi italian portuguese spanish Supported Tasks and Leaderboards text_classification: The dataset can be trained using a SentenceClassification model from HuggingFace transformers. Dataset Structure Data Instances An instance from… See the full description on the dataset page: https://huggingface.co/datasets/cardiffnlp/tweet_sentiment_multilingual.texttext-classification10K<n<100K24 likes1.1k downloads4y agoHugging Face08tweets-hate-speech-detection /tweets_hate_speech_detection Dataset Card for Tweets Hate Speech Detection Dataset Summary The objective of this task is to detect hate speech in tweets. For the sake of simplicity, we say a tweet contains hate speech if it has a racist or sexist sentiment associated with it. So, the task is to classify racist or sexist tweets from other tweets. Formally, given a training sample of tweets and labels, where label ‘1’ denotes the tweet is racist/sexist and label ‘0’ denotes the tweet is not… See the full description on the dataset page: https://huggingface.co/datasets/tweets-hate-speech-detection/tweets_hate_speech_detection.texttext-classification10K<n<100K18 likes950 downloads2y agoHugging Face09mteb /tweet_sentiment_multilingual TweetSentimentClassification An MTEB dataset Massive Text Embedding Benchmark A multilingual Sentiment Analysis dataset consisting of tweets in 8 different languages. Task category t2c Domains Social, Written Reference https://aclanthology.org/2022.lrec-1.27 How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_tasks(["TweetSentimentClassification"]) evaluator =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/tweet_sentiment_multilingual.texttext-classification10K<n<100K4 likes576 downloads1y agoHugging Face10eduagarcia /tweetsentbr_fewshot tweetSentBR (Few-shot) This dataset is a subset of the tweetSentBR, it contains only 75 samples from the training set and all 2.000+ instances of the test set.This is meant for evaluating language models in a few-shot setting on the 🚀 Open Portuguese LLM Leaderboard with the portuguese fork of the Eleuther AI Language Model Evaluation Harness For the complete dataset with 15.000+ annotated tweets go to https://bitbucket.org/HBrum/tweetsentbr or contact the paper authors:… See the full description on the dataset page: https://huggingface.co/datasets/eduagarcia/tweetsentbr_fewshot.texttext-classification1K<n<10K1 likes515 downloads2y agoHugging Face11amgadhasan /arabic_tweets_dialectstexttext-classification100K<n<1M0 likes488 downloads2y agoHugging Face12SinclairSchneider /tweets_sample_2026 Tweets Sample 2026 — Newspaper-Vocabulary Reference Corpus A large, deliberately untargeted sample of public posts from X/Twitter, collected via Nitter by sweeping a 65,689-term newspaper vocabulary rather than a topical keyword set. It is built as a background / reference corpus: a baseline of "what was being said in general" against which a topically targeted collection can be contrasted. It is the reference arm of a narrative-detection study, not a curated dataset about any… See the full description on the dataset page: https://huggingface.co/datasets/SinclairSchneider/tweets_sample_2026.tabulartext-classification10M<n<100M0 likes486 downloads2mo agoHugging Face13jinaai /tweet-stock-synthetic-retrieval_beirThis is a copy of https://huggingface.co/datasets/jinaai/tweet-stock-synthetic-retrieval reformatted into the BEIR format. For any further information like license, please refer to the original dataset. Disclaimer This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/tweet-stock-synthetic-retrieval_beir.image1K<n<10K0 likes453 downloads1y agoHugging Face14rguo123 /trump_tweetstabular10K<n<100K1 likes445 downloads3y agoHugging Face15pain /Arabic-Tweets Dataset Card for Dataset Arabic-Tweets Dataset Summary This dataset has been collected from twitter which is more than 41 GB of clean data of Arabic Tweets with nearly 4-billion Arabic words (12-million unique Arabic words). Languages Arabic Source Data Twitter Example on data loading using streaming: from datasets import load_dataset dataset = load_dataset("pain/Arabic-Tweets",split='train', streaming=True) print(next(iter(dataset)))… See the full description on the dataset page: https://huggingface.co/datasets/pain/Arabic-Tweets.text100M<n<1B23 likes444 downloads3y agoHugging Face16Qanadil /ArSAS_An_Arabic_Speech-Act_and_Sentiment_Corpus_of_Tweets ArSAS: An Arabic Speech-Act and Sentiment Corpus of Tweets Dataset Card for "ArSAS: An Arabic Speech-Act and Sentiment Corpus of Tweets" Note About Sentiment_label_confidence "Crowdflower provides a confidence score with each annotated tweet that represents the confidence in the quality of the label. For a three annotators per tweet setup, the confidence score would range between 0.3 and 1 according to two factors: 1) annotator quality level; and 2) agreement… See the full description on the dataset page: https://huggingface.co/datasets/Qanadil/ArSAS_An_Arabic_Speech-Act_and_Sentiment_Corpus_of_Tweets.tabulartext-classification10K<n<100K1 likes408 downloads2y agoHugging Face17SetFit /tweet_sentiment_extraction Tweet Sentiment Extraction Source: https://www.kaggle.com/c/tweet-sentiment-extraction/data text10K<n<100K11 likes386 downloads4y agoHugging Face18TimKoornstra /financial-tweets-sentiment Financial Sentiment Analysis Dataset Overview This dataset is a comprehensive collection of tweets focused on financial topics, meticulously curated to assist in sentiment analysis in the domain of finance and stock markets. It serves as a valuable resource for training machine learning models to understand and predict sentiment trends based on social media discourse, particularly within the financial sector. Data Description The dataset comprises tweets… See the full description on the dataset page: https://huggingface.co/datasets/TimKoornstra/financial-tweets-sentiment.texttext-classification10K<n<100K24 likes330 downloads3y agoHugging Face19rescrv /ai-tweetstext1M<n<10M2 likes315 downloads2y agoHugging Face20MohammadOthman /mo-customer-support-tweets-945k Customer Support on Twitter Dataset 945k Dataset Description Context This dataset provides a large corpus of real-world English conversations between consumers and customer support agents on Twitter, designed to drive innovation in Natural Language Processing (NLP) by providing data that better matches the actual language used in contemporary customer support interactions. Content Initially, the data included complex threads of conversations… See the full description on the dataset page: https://huggingface.co/datasets/MohammadOthman/mo-customer-support-tweets-945k.texttext-generation100K<n<1M2 likes312 downloads2y agoHugging Face21Qanadil /ASTD_Arabic_Sentiment_Tweets_Dataset Citation (https://aclanthology.org/D15-1299/) texttext-classification1K<n<10K0 likes293 downloads2y agoHugging Face22venetis /disaster_tweetstabulartext-classification1K<n<10K3 likes253 downloads4y agoHugging Face23ExponentialScience /DLT-Tweets DLT-Tweets [Paper] • [Code] Dataset Description Dataset Summary DLT-Tweets is a large-scale corpus of social media posts related to Distributed Ledger Technology (DLT). This dataset is part of the larger DLT-Corpus collection, designed to support NLP research, social computing studies, and public discourse analysis in the DLT domain. It was introduced in the paper DLT-Corpus: A Large-Scale Text Collection for the Distributed Ledger Technology Domain.… See the full description on the dataset page: https://huggingface.co/datasets/ExponentialScience/DLT-Tweets.tabulartext-generation10M<n<100M0 likes195 downloads7mo agoHugging Face24community-datasets /covid_tweets_japanese Dataset Card for COVID-19 日本語Twitterデータセット (COVID-19 Japanese Twitter Dataset) Dataset Summary 53,640 Japanese tweets with annotation if a tweet is related to COVID-19 or not. The annotation is by majority decision by 5 - 10 crowd workers. Target tweets include "COVID" or "コロナ". The period of the tweets is from around January 2020 to around June 2020. The original tweets are not contained. Please use Twitter API to get them, for example. Supported Tasks and… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/covid_tweets_japanese.texttext-classification10K<n<100K2 likes187 downloads2y agoHugging Face25alt-qsri /tweets_ar_en_parallel Twitter users often post parallel tweets—tweets that contain the same content but are written in different languages. Parallel tweets can be an important resource for developing machine translation (MT) systems among other natural language processing (NLP) tasks. This resource is a result of a generic method for collecting parallel tweets. Using the method, we compiled a bilingual corpus of English-Arabic parallel tweets and a list of Twitter accounts who post English-Arabic tweets regularly. Additionally, we annotate a subset of Twitter accounts with their countries of origin and topic of interest, which provides insights about the population who post parallel tweets.translation100K<n<1M4 likes173 downloads3y agoHugging Face26fschlatt /trump-tweetsThis is a clone of the Trump Tweet Kaggle dataset found here: https://www.kaggle.com/datasets/headsortails/trump-twitter-archive tabular10K<n<100K6 likes154 downloads3y agoHugging Face27DSCI511G1 /COP26_Energy_Transition_Tweetstabular10K<n<100K3 likes144 downloads5y agoHugging Face28alperbayram /Tweet_Siniflandirma References alper bayram texttext-classification1K<n<10K2 likes134 downloads4y agoHugging Face29StephanAkkerman /stock-market-tweets-data Stock Market Tweets Data Overview This dataset is the same as the Stock Market Tweets Data on IEEE by Bruno Taborda. Data Description This dataset contains 943,672 tweets collected between April 9 and July 16, 2020, using the S&P 500 tag (#SPX500), the references to the top 25 companies in the S&P 500 index, and the Bloomberg tag (#stocks). Dataset Structure created_at: The exact time this tweet was posted. text: The text of the tweet, providing… See the full description on the dataset page: https://huggingface.co/datasets/StephanAkkerman/stock-market-tweets-data.texttext-classification100K<n<1M6 likes117 downloads3y agoHugging Face30astarostap /antisemitic-tweets0 likes115 downloads6y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.