datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
crypto-market-sentiment-observations
Instrumetriq: Crypto Market Activity & Sentiment Context Dataset
Time-aligned observational snapshots of crypto market activity and social sentiment across 270+ assets, designed to contextualize market structure, liquidity, and attention dynamics.
Observational data only. No trading advice, predictions, or signal generation.
Dataset Description
This dataset provides weekly Sunday snapshots from Instrumetriq's continuous monitoring pipeline. Each snapshot… See the full description on the dataset page: https://huggingface.co/datasets/Instrumetriq/crypto-market-sentiment-observations.ArSAS_An_Arabic_Speech-Act_and_Sentiment_Corpus_of_Tweets
ArSAS: An Arabic Speech-Act and Sentiment Corpus of Tweets
Dataset Card for "ArSAS: An Arabic Speech-Act and Sentiment Corpus of Tweets"
Note About Sentiment_label_confidence
"Crowdflower provides a confidence score with each annotated tweet that represents the confidence in the quality of the label. For a three annotators per tweet setup, the confidence score would range between 0.3 and 1 according to two factors: 1) annotator quality level; and 2) agreement… See the full description on the dataset page: https://huggingface.co/datasets/Qanadil/ArSAS_An_Arabic_Speech-Act_and_Sentiment_Corpus_of_Tweets.US_Airline_Sentimentfinancial_news_sentiment
Dataset Card for "financial_news_sentiment"
Manually validated sentiment for ~2000 Canadian news articles.
The dataset also include a column topic which contains one of the following value:
acquisition
other
quaterly financial release
appointment to new position
dividend
corporate update
drillings results
conference
share repurchase program
grant of stocks
This was generated automatically using a zero-shot classification model and was not reviewed manually.
sentiment-analysis-in-commodity-market-gold
Dataset Card for Sentiment Analysis of Commodity News (Gold)
This is a news dataset for the commodity market which has been manually annotated for 10,000+ news headlines across multiple dimensions into various classes. The dataset has been sampled from a period of 20+ years (2000-2021).
The dataset was curated by Ankur Sinha and Tanmay Khandait and is detailed in their paper "Impact of News on the Commodity Market: Dataset and Results." It is currently published by the authors on… See the full description on the dataset page: https://huggingface.co/datasets/SaguaroCapital/sentiment-analysis-in-commodity-market-gold.ledger-market-sentiment
LEDGER Market Sentiment Prediction Data
Data used for the market sentiment prediction case study in the LEDGER paper,
linking CEO-letter rhetoric to EPS surprises and post-publication market reactions.
Dataset Description
This dataset supports research on whether the rhetoric in corporate annual report
CEO letters carries signal about future fundamentals and market reaction. It covers
six highly liquid industries (specialty chemicals, auto parts, packaged foods… See the full description on the dataset page: https://huggingface.co/datasets/artefactory/ledger-market-sentiment.multilingual-amazon-review-sentiment-processedDLT-Sentiment-News
DLT-Sentiment-News
Paper | Code
Dataset Description
Dataset Summary
DLT-Sentiment-News is a specialized sentiment analysis dataset for the Distributed Ledger Technology (DLT) domain. It addresses the lack of high-quality labeled data that captures domain-specific sentiment expressed by cryptocurrency community members.
The dataset contains 23,301 examples with 1.85 million tokens (average 79.51 tokens per example), spanning from January 2021 to May 2025. Each… See the full description on the dataset page: https://huggingface.co/datasets/ExponentialScience/DLT-Sentiment-News.financial_news_sentiment_mixte_with_phrasebank_75
Dataset Card for "financial_news_sentiment_mixte_with_phrasebank_75"
This is a customized version of the phrasebank dataset in which I kept only sentences validated by at least 75% annotators.In addition I added ~2000 articles of Canadian news where sentiment was validated manually.
The dataset also include a column topic which contains one of the following value:
acquisition
other
quaterly financial release
appointment to new position
dividend
corporate update
drillings results… See the full description on the dataset page: https://huggingface.co/datasets/Jean-Baptiste/financial_news_sentiment_mixte_with_phrasebank_75.HW1_Own_Lexicon_SentimentChinese-Herbal-Medicine-Sentiment
中药情感分析数据集 - 数据说明书
Chinese Herbal Medicine Sentiment Analysis Dataset - Datacard
数据集概述 / Dataset Overview
基本信息 / Basic Information
数据集名称 / Dataset Name: Chinese Herbal Medicine Sentiment Analysis Dataset
版本 / Version: 1.0.0
创建日期 / Created: 2025-08-26
作者 / Author: Xingqiang Chen
许可证 / License: MIT
语言 / Language: 中文 (Chinese)
领域 / Domain: 中药 / 传统中医药 (Traditional Chinese Medicine)
数据规模 / Data Scale
总样本数 / Total Samples: 234,879
唯一产品数 /… See the full description on the dataset page: https://huggingface.co/datasets/OpenModels/Chinese-Herbal-Medicine-Sentiment.tweets_pt_sentiment_analysis
Dataset Card for "tweets_pt_sentiment_analysis"
More Information needed
agent-trace-sentiment
Coding-Agent User Message Sentiment
User messages from every public format:agent-traces dataset on the Hugging Face Hub, classified as POSITIVE / NEUTRAL / NEGATIVE by a small open LLM, with a one-sentence reason for each label so you can audit any classification.
Accompanies the blog post "Your AI Coding Agent Has a Patience Cliff".
What's in here
Each row is one message from a developer to their coding agent (Claude Code, Pi, Codex, or variants).
Column
Type… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/agent-trace-sentiment.reddit-it-labor-sentiment-2020-2026
Dataset
Overview
Reddit IT Labor Market Sentiment (2020–2026)
Thesis title
The impact of the artificial intelligence bubble on the job market for new IT specialists:
An analysis of the disconnect between recruitment requirements and attitudes
This dataset is a collection of Reddit posts and comments pulled from 32 different IT-focused subreddits.
It’s built for researchers looking at shifts in the labor market, skill inflation, and how people working in IT… See the full description on the dataset page: https://huggingface.co/datasets/NetGene/reddit-it-labor-sentiment-2020-2026.twitter-sentiment-analysis
Twitter Sentiment Analysis: Prabowo's First 100 Days
Dataset Overview
This dataset contains tweets related to President Prabowo Subianto's first 100 days in office in Indonesia (2024-2029). The tweets have been preprocessed and classified into three sentiment categories using a fine-tuned BERT model for Indonesian language (IndoBERT).
Dataset Details
Language: Indonesian
Source: Twitter/X
Time period: First 100 days of President Prabowo's… See the full description on the dataset page: https://huggingface.co/datasets/KidzRizal/twitter-sentiment-analysis.fintweet-sentiment-2025
FinTweet Sentiment Dataset 2025
A curated dataset of financial tweets enriched with market data for sentiment-driven trading signal classification.
Dataset Description
This dataset combines social media signals from financial Twitter accounts with real-time market data from Interactive Brokers to create labeled samples for 3-class sentiment classification (BUY/HOLD/SELL).
Dataset Summary
Property
Value
Time Period
2024-2025
Total Samples
43,468… See the full description on the dataset page: https://huggingface.co/datasets/michael-zus/fintweet-sentiment-2025.news-sentiment-databanking_sentiment_vietnamesegerman_politicians_twitter_sentiment
Information
This dataset shows 1785 manually annotated tweets from German politicians during the election year 2021 (01.01.2021 - 31.12.2021).
The tweets were annotated by 6 academics which were separated into two different groups. So every group of 3 people annotated the sentiment of ~900 tweets. For every tweet, the majority label was built. The annotation result had a moderate Kappa agreement.
Preprocessing
The source for this version of the dataset is located here.… See the full description on the dataset page: https://huggingface.co/datasets/Alienmaster/german_politicians_twitter_sentiment.mteb-tweet_sentiment_extraction-avs_triplets
MTEB Tweet Sentiment Extraction Triplets Dataset
This dataset was used in the paper GISTEmbed: Guided In-sample Selection of Training Negatives for Text Embedding Fine-tuning. Refer to https://arxiv.org/abs/2402.16829 for details.
The code for generating the data is available at https://github.com/avsolatorio/GISTEmbed/blob/main/scripts/create_classification_dataset.py.
Citation
@article{solatorio2024gistembed,
title={GISTEmbed: Guided In-sample Selection of… See the full description on the dataset page: https://huggingface.co/datasets/avsolatorio/mteb-tweet_sentiment_extraction-avs_triplets.quirky_sentiment
Dataset Card for "quirky_sentiment"
More Information needed
swik-sentiment-labels
swik Financial Sentiment Labels
Asset-specific financial sentiment labels for 35 securities — commodities, FX, indices, and crypto.
v0.2 — Updated 2026-04-04. 50,589 total records.
What makes this different
Standard financial sentiment datasets assign generic polarity to headlines. This dataset applies
asset-specific inversion context from the swik inversion catalog — a community-maintained
knowledge base of how phrases actually move prices for each specific asset.… See the full description on the dataset page: https://huggingface.co/datasets/polibert/swik-sentiment-labels.political-social-x-us-sentiment-v1africa-synth-telecom-social-media-sentiment-datasets-nigeria
Africa Synth Telecom Social Media Sentiment Datasets Nigeria | Africa (Electric Sheep Africa metadata inventory)
Size category: 100K<n<1M - Formats: parquet - Sector: technology_digital - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-telecom-social-media-sentiment-datasets-nigeria.spanish-targeted-sentiment-headlinesquirky_sentiment_alice_easy
Dataset Card for "quirky_sentiment_alice_easy"
More Information needed
tweets_about_german_politicians_jan_feb_2025_with_party_and_sentimentsentiment140ForLlamamyanmar-social-media-sentiment-analysis-dataset
Myanmar Social Media Sentiment Analysis Dataset
A Myanmar language dataset for sentiment analysis of social media content, translated from an English source dataset.
Dataset Description
This dataset contains social media text with sentiment annotations translated into Myanmar language. It is derived from the original Social Media Sentiments Analysis Dataset on Kaggle, with texts professionally translated to Myanmar language while preserving the sentiment labels.… See the full description on the dataset page: https://huggingface.co/datasets/chuuhtetnaing/myanmar-social-media-sentiment-analysis-dataset.quirky_sentiment_bob_hard
Dataset Card for "quirky_sentiment_bob_hard"
More Information needed
