datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
twitter_suicidal_risk
Twitter Suicide Risk Level Dataset
Short English tweets paired with a 0–4 suicide risk label, used for fine-tuning and
evaluating risk-level classification. This directory holds the final splits:
train.jsonl / val.jsonl / test.jsonl.
Files and size
File
Rows
Share
train.jsonl
7006
80%
val.jsonl
875
10%
test.jsonl
875
10%
Total
8756
100%
Fields
JSONL, one sample per line, three fields only:
Field
Type
Description
id… See the full description on the dataset page: https://huggingface.co/datasets/AdamLeung/twitter_suicidal_risk.twitter-sentiment-analysis
🐦 Twitter Sentiment Analysis (bdstar/twitter-sentiment-analysis)
🧠 Overview
A refined and merged version of Twitter text sentiment datasets, providing a clean and well-balanced dataset for sentiment classification across three sentiment categories:positive, negative, and neutral.
This dataset is split into three parts — train, test, and validation — each sourced from highly reputable open datasets.It is designed for training, evaluating, and benchmarking NLP models for… See the full description on the dataset page: https://huggingface.co/datasets/bdstar/twitter-sentiment-analysis.twitter_us_airline_sentiment
Twitter US Airline Sentiment Dataset
This dataset contains tweets about US airlines labeled with sentiment (negative, neutral, positive).
Dataset Details
Total tweets: 14,640
Training set: 11,712 tweets
Validation set: 1,464 tweets
Test set: 1,464 tweets
Classes: negative (0), neutral (1), positive (2)
Files
train.jsonl: Training data (11,712 examples)
validation.jsonl: Validation data (1,464 examples)
test.jsonl: Test data (1,464 examples)
Usage… See the full description on the dataset page: https://huggingface.co/datasets/viethq1906/twitter_us_airline_sentiment.nick-routledge-x-twitter-archiveNick Routledge X/Twitter Archive
A growing machine-readable collection of posts and replies authored by Nick Routledge (@nick_routledge), drawn from his public X/Twitter archive.
The complete searchable archive is available at https://archive.nickroutledge.org.
Files
posts.jsonl: one structured post per line, suitable for data pipelines and language-model research.
posts.csv: the same authored posts in spreadsheet-compatible form.
Licence
Nick Routledge’s authored post text is made available… See the full description on the dataset page: https://huggingface.co/datasets/fellowservant/nick-routledge-x-twitter-archive.Latvian-Twitter-Eater-Corpus-Sentiment
Latvian Twitter Eater Corpus - Sentiment Analysis Sub-corpus
This data set contains 5420 tweets with human-annotated sentiment as positive (pos), neutral (neu) or negative (neg). 1631 tweets are positive, 2507 - neutral and 1282 - negative.
ltec-sentiment-annotated-train.json contains tweets with human annotated sentiment
ltec-sentiment-annotated-test.json contains the test set that we used in our paper
Tweet Structure
{
"label":1… See the full description on the dataset page: https://huggingface.co/datasets/matiss/Latvian-Twitter-Eater-Corpus-Sentiment.twitter_personas
X Character Files
Generate a persona based on your X data.
Walkthrough
First, download an archive of your X data here: https://help.x.com/en/managing-your-account/how-to-download-your-x-archive
This could take 24 hours or more.
Then, run tweets2character directly from your command line:
npx tweets2character
NOTE: You need an API key to use Claude or OpenAI.
If everything is correct, you'll see a loading bar as the script processes your data and generates a character… See the full description on the dataset page: https://huggingface.co/datasets/NapthaAI/twitter_personas.chatgpt4-noisy-translation-twitter-dialect
ChatGPT 4 Noisy Translation Twitter to local dialect
Notebooks at https://github.com/mesolitica/malaysian-dataset/tree/master/translation/chatgpt4-twitter-dialect
Automatic-Sarcasm-Detection-Twitter
Automatic Sarcasm Detection
The Shared Task (2nd FigLang Workshop at ACL 2020) is now over. Thanks a lot, participants :)
Please refer to reddit and twitter sub-directories for further references on datasets.
For Twitter and Reddit, training and testing datasets are provided for sarcasm detection tasks in jsonlines format.
Each line contains a JSON object with the following fields :
label : SARCASM or NOT_SARCASM
id: String identifier for sample. This id will be required… See the full description on the dataset page: https://huggingface.co/datasets/shiv213/Automatic-Sarcasm-Detection-Twitter.Latvian-Twitter-Eater-Corpus-Translation
Latvian Twitter Eater Corpus - Translation Sub-corpus
The Latvian Twitter Eater Translation test set contains two references for English translations produced by two different translators. A separate random manual evaluation of the translations has shown that the translations in english reference 2 are slightly higher quality than the ones in english reference 1.
Publications
If you use this corpus or scripts, please cite the following paper:
Matīss Rikters, Edison… See the full description on the dataset page: https://huggingface.co/datasets/matiss/Latvian-Twitter-Eater-Corpus-Translation.ja-zh-twitter-translatetranslate by @Nekofoxtweet (me)
twitter source from @RindouMikoto
Famous-Keyword-Twitter-RepliesThe "Famous Keyword Twitter Replies Dataset" is a comprehensive collection of Twitter data that focuses on popular keywords and their associated replies. This dataset contains five essential columns that provide valuable insights into the Twitter conversation dynamics:
Keyword: This column represents the specific keyword or topic of interest that generated the original tweet. It helps identify the context or subject matter around which the conversation revolves.
Main_tweet: The main_tweet… See the full description on the dataset page: https://huggingface.co/datasets/jacksoncsie/Famous-Keyword-Twitter-Replies.twitter-sentiment-analysis
🐦 Twitter Sentiment Analysis (bdstar/twitter-sentiment-analysis)
🧠 Overview
A refined and merged version of Twitter text sentiment datasets, providing a clean and well-balanced dataset for sentiment classification across three sentiment categories:positive, negative, and neutral.
This dataset is split into three parts — train, test, and validation — each sourced from highly reputable open datasets.It is designed for training, evaluating, and benchmarking NLP models for… See the full description on the dataset page: https://huggingface.co/datasets/akhiljoe143/twitter-sentiment-analysis.tesla_twitter_labeltwittertwitter-tweets-coldstarttwitter-2kbroad-twitter-corpustwitter-tweets-promptstwitter-shortenedtwitter-tweets-coldstart-SFTTwitter_finetunetwitter-tweets-prompts-1620-linetwitter-tweets-coldstart-GRPO
