datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tweet_sentiment_extraction
Tweet Sentiment Extraction
Source: https://www.kaggle.com/c/tweet-sentiment-extraction/data
mo-customer-support-tweets-945k
Customer Support on Twitter Dataset 945k
Dataset Description
Context
This dataset provides a large corpus of real-world English conversations between consumers and customer support agents on Twitter, designed to drive innovation in Natural Language Processing (NLP) by providing data that better matches the actual language used in contemporary customer support interactions.
Content
Initially, the data included complex threads of conversations… See the full description on the dataset page: https://huggingface.co/datasets/MohammadOthman/mo-customer-support-tweets-945k.ai-tweetshate_offensive_tweets
Hate and Offensive Speech Dataset
This dataset was created using several datasets that can be found on Hugging Face:
-SetFit/hate_speech_offensive:https://huggingface.co/datasets/SetFit/hate_speech_offensive
-tweets_hate_speech_detection:https://huggingface.co/datasets/tweets_hate_speech_detection
-thefrankhsu/hate_speech_twitter:https://huggingface.co/datasets/thefrankhsu/hate_speech_twitter… See the full description on the dataset page: https://huggingface.co/datasets/MartynaKopyta/hate_offensive_tweets.tweets
BlockMesh Network
Dataset Summary
The dataset is a sample of our Twitter data collection.
It has been prepared for educational and research purposes.
It includes public tweets.
The dataset is comprised of a JSON lines.
The format is:
{
"user":"Myy23081040",
"id":"1870163769273589994",
"link":"https://x.com/Myy23081040/status/1870163769273589994",
"tweet":"Seu pai é um fofo skskks",
"date":"2024-12-21",
"reply":"0",
"retweet":"0",
"like":"2"
}
user the… See the full description on the dataset page: https://huggingface.co/datasets/blockmesh/tweets.malaysia-tweets-sentimentfrench_tweets
[!NOTE]
Dataset origin: https://www.kaggle.com/datasets/loockeeer/french-tweets-215k
tweetstockFrom the full period (Jan 1 – May 30, 2025), we extracted data corresponding to April 1, 2025 through May 31, 2025 and created this dataset.
Data Curation
Stock Data
Tickers: AAPL, TSLA, AMZN, MSFT, NVDA, GOOGL, META, INTC, SHOP, SPYG(10 stocks in total)
Period: 2025‑01‑01 to 2025‑05‑30
Source: Historical daily OHLCV (open, high, low, close, volume) via a financial data API (e.g., Yahoo Finance).
Frequency: Daily (market close).
Twitter Data
Accounts… See the full description on the dataset page: https://huggingface.co/datasets/Knovaai/tweetstock.financial-sentiment-tweets
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/ZombitX64/financial-sentiment-tweets.emotion-analysis-tweets
Dataset Card for Dataset Name
Dataset Details
Combination of data from:
https://www.kaggle.com/datasets/bhavikjikadara/emotions-dataset
https://www.kaggle.com/datasets/parulpandey/emotion-dataset
***I did not make any of this data, I simply put it all together in one file
Labels:
0 = Sadness
1 = Joy
2 = Love
3 = Anger
4 = Fear
5 = Surprise
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/echung682/emotion-analysis-tweets.tweet-sentiment-analysis-from-kaggleworld-cup-2022-tweetstweet_sentiment_extractionmo-customer-support-tweets-945k
Customer Support on Twitter Dataset 945k
Dataset Description
Context
This dataset provides a large corpus of real-world English conversations between consumers and customer support agents on Twitter, designed to drive innovation in Natural Language Processing (NLP) by providing data that better matches the actual language used in contemporary customer support interactions.
Content
Initially, the data included complex threads of conversations… See the full description on the dataset page: https://huggingface.co/datasets/shaikayub/mo-customer-support-tweets-945k.depressive_tweets.jsonltweet-scholartweetsIAmAaronWill_tweets_2173mo-customer-support-tweets-945k
Customer Support on Twitter Dataset 945k
Dataset Description
Context
This dataset provides a large corpus of real-world English conversations between consumers and customer support agents on Twitter, designed to drive innovation in Natural Language Processing (NLP) by providing data that better matches the actual language used in contemporary customer support interactions.
Content
Initially, the data included complex threads of conversations… See the full description on the dataset page: https://huggingface.co/datasets/Prady06/mo-customer-support-tweets-945k.twitter-tweets-coldstartfinance-tweets-sentimentbenthecarman-tweetstweets-18-04-23TweetsNearMSUMoorheadkanye-tweetssalvini_tweetstwitter-tweets-promptstwitter-tweets-coldstart-SFTbenthecarman-tweets2twitter-tweets-prompts-1620-line
