datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ambivalent_affectgspc-affect
GSPC — affect bank (AffectBench)
Council of AI measurement bank. Measurement, not certification.
Bank. Frozen split. Live n is the matching axis on GET https://councilof.ai/api/gspc, not a Hub score. Not a certificate. Art 50 (EUR-Lex): 2 August 2026 live; marking grace 2 December 2026.
Live measurement. This bank stands behind the affect row of the live GSPC board: GET https://councilof.ai/api/gspc?axis=affect (family, kind, status and n are on that row, never typed here; the… See the full description on the dataset page: https://huggingface.co/datasets/csoai/gspc-affect.t9p3c8m1-axr4e6vocal-affect-bench
VocalAffectBench
VocalAffectBench is a test-only benchmark for evaluating whether AI audio models can identify expressed vocal emotion from raw audio.
Paper: VocalAffectBench: Evaluating Vocal Emotion Recognition in AI Audio Models
The benchmark targets the expressed emotion — what the speaker conveys through vocal tone, prosody, pace, intensity, and pauses — not inferred internal state.
Contents
280 human-recorded English WAV clips, totalling 2.32 hours.
7… See the full description on the dataset page: https://huggingface.co/datasets/besimple-ai/vocal-affect-bench.AffectNetaffectnethq
Dataset Card for "affectnethq"
More Information needed
h4p7t3x2-jn6b9_tran
affectexpect/h4p7t3x2-jn6b9_tran
This dataset contains transcribed audio files organized in folders for scalability.
Dataset Structure
The dataset is organized with:
Audio files: Stored in audio_XXXXX/ folders (5000 files per folder)
Metadata: Stored in data_XXXXX/ folders as parquet files
This organization follows Hugging Face best practices for datasets with millions of files.
Statistics
Total files: 926
Total batches: 2427
Audio folders: 3
Files per… See the full description on the dataset page: https://huggingface.co/datasets/affectexpect/h4p7t3x2-jn6b9_tran.AffectNet-Mediapipe-478Points-3Dt9p3c8m1-axr4e6_sepAffectDF_EmotionSDD
AffectDF: Emotionally Expressive Speech Deepfake Benchmark
Overview
AffectDF is a large-scale benchmark for speech deepfake detection under emotionally expressive spoofing conditions. The dataset is designed to evaluate whether current speech deepfake detection (SDD) systems can generalize beyond conventional neutral-speech benchmarks to modern emotional and expressive speech attacks.
AffectDF contains approximately 260 hours of audio generated using 21 spoofing… See the full description on the dataset page: https://huggingface.co/datasets/AffectDF/AffectDF_EmotionSDD.affectnet_short
Dataset Card for "affectnet_short"
More Information needed
llm-affect-lab
LLM Affect Lab
This dataset contains the API-level results for LLM Affect Lab, a study of functional affect signatures in language model behavior.
Functional Affect Score (FAS) is a 0-1 behavioral proxy. It combines generated-token confidence, enthusiastic language, consistency across repeated samples, forced self-report computed from digit top-logprob probabilities, and length control. The goal is not to claim that models feel emotions; the goal is to measure whether different… See the full description on the dataset page: https://huggingface.co/datasets/kishan51/llm-affect-lab.affectnetLive-Affect-Analysism4r9e1x8-cd7h2affect_of_removing_misalligned_examples-eval_resultsaffectnet-comparison-resultsPoliStance_AffectDataset for training an entailment classifier to recognize approval/disapproval of politicians.
Documents are Tweets from Kawintiranon (2022), the MTSD dataset, as well as Tweets and sentences taken weekly newsletters for select politicians from the 115th, 116th, and 117th congress.
Documents are triple coded -- once from the original compilers of the dataset, once from GPT-4, and a third time to adjudicate discrepancies between the two.
Twitter handles from politicians in the dataset have… See the full description on the dataset page: https://huggingface.co/datasets/mlburnham/PoliStance_Affect.n4x7d2q9-hf1m8t3_sepaffectiamharicaffect_of_removing_misalligned_examples-judged_resultsaffect_of_removing_misalligned_examples-response_samplest9p3c8m1-axr4e6_tran
affectexpect/t9p3c8m1-axr4e6_tran
This dataset contains transcribed audio files organized in folders for scalability.
Dataset Structure
The dataset is organized with:
Audio files: Stored in audio_XXXXX/ folders (5000 files per folder)
Metadata: Stored in data_XXXXX/ folders as parquet files
This organization follows Hugging Face best practices for datasets with millions of files.
Statistics
Total files: 8,901
Total batches: 5183
Audio folders: 6
Files per… See the full description on the dataset page: https://huggingface.co/datasets/affectexpect/t9p3c8m1-axr4e6_tran.affectnet_no_contempth4p7t3x2-jn6b9_sepempathy-affective-datasetsEmpathy and Affective Computing Datasets Summary
This repository is a curated summary of existing datasets for empathy and affective computing research. It distinguishes between empathy-focused datasets (directly measuring empathic processes) and general affective computing datasets (emotion recognition, valence/arousal, etc.). This is not a new dataset but a reference guide—please access original datasets via provided links and cite their sources.
Empathy and Affective Computing… See the full description on the dataset page: https://huggingface.co/datasets/Limorgu/empathy-affective-datasets.affect_of_removing_misalligned_examples-conservative_qual_removedaffect_of_removing_misalligned_examples-judge_subsetaipsy-affect
AIPsy-Affect
Clinical affect stimuli for mechanistic interpretability research.
480 keyword-free clinical vignettes designed to study how language models process emotional content — without the confound that has compromised every prior emotion dataset.
The core problem: existing emotion datasets contain the emotion words being tested. A stimulus labeled "anger" that contains the word "furious" doesn't test emotion processing — it tests keyword detection. Every mechanistic… See the full description on the dataset page: https://huggingface.co/datasets/keidolabs/aipsy-affect.PoliStance_Affect_QTThis dataset contains quote tweets that have been hand labeled for stance towards a politician. Quote tweets are a particularly challenging classification task because they contain multiple (often contradictory) expressions from multiple authors.
Twitter handles from politicians in the dataset have been replaced by their name. Be aware that "rt @realdonaldtrump You're a liar!" has been replaced with "rt trump You're a liar!",
And means the author is retweeting Trump calling someone else a liar… See the full description on the dataset page: https://huggingface.co/datasets/mlburnham/PoliStance_Affect_QT.
