datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MANGO
MANGO: A Corpus of Human Ratings for Speech
MANGO (MUSHRA Assessment corpus using Native listeners and Guidelines to understand human Opinions at scale) is the first large-scale dataset designed for evaluating Text-to-Speech (TTS) systems in Indian languages.
Key Features:
255,150 human ratings of TTS-generated outputs and ground-truth human speech.
Covers two major Indian languages: Hindi & Tamil, and English.
Based on the MUSHRA (Multiple Stimuli with Hidden Reference… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/MANGO.commonsilent-co-driver-data
Silent Co-Driver Dataset
Sample lap-time data and radio clips used for "The Silent Co-Driver" —
a tool that reads driver stress from radio calls and correlates it
with lap performance.
Files
laptimes.csv: sample lap-time data (lap number, lap time in seconds)
sample_clips/: example radio clips representing calm, stressed, and tired driver tones
Used with
openai/whisper-small (speech-to-text)
superb/wav2vec2-base-superb-er (audio emotion… See the full description on the dataset page: https://huggingface.co/datasets/Manish0134/silent-co-driver-data.AYDID-public
AYDID: Arabic Yemeni Dialect Identification Dataset (public release)
This repository contains a public sample and the held-out test set of AYDID,
the first dedicated speech corpus for Yemeni Arabic at the sub-dialectal level,
supporting both automatic speech recognition (ASR) and dialect identification (DID).
Note on scope. This release contains a representative sample plus the benchmark
test set. It is intended for evaluating models against the published baselines, not
for… See the full description on the dataset page: https://huggingface.co/datasets/mansoorSaleh/AYDID-public.
