datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ProsodyEmoji
The Prosody of Emojis (ACL 2026)
Dataset Summary
Prosodic features such as pitch, timing, and intonation are central to spoken communication, conveying emotion, intent, and discourse structure. In text-based settings, emojis act as visual surrogates that add affective and pragmatic nuance.
This dataset examines how emojis influence prosodic realisation in speech and how listeners interpret prosodic cues to recover emoji meanings. It contains human speech data… See the full description on the dataset page: https://huggingface.co/datasets/GiulioZh/ProsodyEmoji.macro_prosody_sample_set
Alexandria Voice Corpus — Multilingual Macro-Prosody Telemetry
Version 1.1 — Replacement release
This pack supersedes the earlier Korean & Hindi two-language release. That release was built on a pipeline with several unresolved quality-gate bugs (documented below). This version corrects all known issues and expands to seven typologically diverse languages.
No audio is included. This is a structured acoustic feature dataset for linguistic research, speech technology, and… See the full description on the dataset page: https://huggingface.co/datasets/moonscape-software/macro_prosody_sample_set.Prosody_Breton
[!NOTE]
Dataset origin: https://cocoon.huma-num.fr/exist/crdo/meta/cocoon-3626dbec-0905-4a59-a6db-ec09054a59f7
Description originale
We collected data from Breton dialects to study their prosody and how they inform the syntax-phonology interface. The protocols comprise elicitations and free corpora from Kerne and Treger. The free corpus files include self-portraits of native Breton speakers in their eighties in 2022. The elicited files consist of the results of a… See the full description on the dataset page: https://huggingface.co/datasets/Bretagne/Prosody_Breton.multilingual_audio_alignments_prosodysocial-robotics-acoustic-prosody
Social Robotics: Acoustic Prosody (03c)
Ambient vocal tone around each task — alarming vs soothing — as corroborating context.
One layer of the Social-Affective Filter (SAF) — dehydrated social-signal metadata extracted from egocentric (first-person) video so robots can learn to read human reactions. No raw pixels and no audio. Each row is one source video, keyed by video_id; rehydrate against your own legally-obtained Ego4D copies (below).
Rows: 989 — videos in the… See the full description on the dataset page: https://huggingface.co/datasets/louisye/social-robotics-acoustic-prosody.amy-lm-synthetic-prosody-speech-datasetamy-lm-prosody-validation-benchmarkProsodyNaturalness_ProsAudit-ProtosyntaxProsodyNaturalness_ProsAudit-LexicalProsody_Semantic_Mismatchamy-lm-prosody-validation-benchmark-ultra
