tweets
Datasets
All datasets matching “tweets”tweet_sentiment_extraction
TweetSentimentExtractionClassification
An MTEB dataset
Massive Text Embedding Benchmark
Task category
t2c
Domains
Social, Written
Reference
https://www.kaggle.com/competitions/tweet-sentiment-extraction/overview
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["TweetSentimentExtractionClassification"])
evaluator = mteb.MTEB(task)
model =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/tweet_sentiment_extraction.senator-tweetsangry-tweets
Dataset Card for AngryTweets
Dataset Summary
This dataset consists of anonymised Danish Twitter data that has been annotated for sentiment analysis through crowd-sourcing. All credits go to the authors of the following paper, who created the dataset:
Pauli, Amalie Brogaard, et al. "DaNLP: An open-source toolkit for Danish Natural Language Processing." Proceedings of the 23rd Nordic Conference on Computational Linguistics (NoDaLiDa). 2021
Supported Tasks and… See the full description on the dataset page: https://huggingface.co/datasets/DDSC/angry-tweets.political-leaning-tweets-100k
🗳️ political-leaning-tweets-100k
Châtelet AI presents a 100,000+ dataset of tweets labelled for political leaning: neutral, liberal, conservative.Labels are machine-generated using a SOTA thinking-enabled LLM. The dataset is intended for research on political language modelling, ideology detection, robustness, and safety evaluation.
📦 Dataset Card
Name: chatelet/political-leaning-tweets-100k
Publisher: Châtelet AI
Licence: MIT with additional restrctions against… See the full description on the dataset page: https://huggingface.co/datasets/chatelet/political-leaning-tweets-100k.spanish-tweets
spanish-tweets
A big corpus of tweets for pretraining embeddings and language models
Dataset Summary
A big dataset of (mostly) Spanish tweets for pre-training language models (or other representations).
Supported Tasks and Leaderboards
Language Modeling
Languages
Mostly Spanish, but some Portuguese, English, and other languages.
Dataset Structure
Data Fields
tweet_id: id of the tweet
user_id: id of the user
text:… See the full description on the dataset page: https://huggingface.co/datasets/pysentimiento/spanish-tweets.twitter100m_tweets
Dataset Card for "twitter100m_tweets"
Dataset with tweets for this post.
DOI: 10.5281/zenodo.15086029
