datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
vocalgrad
VocalGrad
VocalGrad is an audio benchmark for evaluating whether a model can detect the
direction of gradual perceptual change in speech. This public release contains
the test split only.
Each example contains one audio clip and one target attribute. The task is to
answer whether that attribute increases or decreases over time.
Task
Given an audio clip and an attribute name, predict one of two labels:
increase
decrease
The ground-truth label is derived from the metadata… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-user-592888/vocalgrad.Synthetic-User-Turn-TTS
Synthetic Malaysian Telco Call-Centre Speech
Synthetic Malaysian call-centre customer utterances, as text and as speech.
The text is fully synthetic dialogue styled after real Malaysian ISP/telco ("Unifi")
call-centre recordings, containing no real customer data. The audio subsets take customer
(user) turns and voice them with a voice-conversion model, keeping only clips an ASR
round-trip confirms are accurate.
Subsets
subset
rows
content
default
4,260… See the full description on the dataset page: https://huggingface.co/datasets/Scicom-intl/Synthetic-User-Turn-TTS.Mir-1k-use-DJCM-trainingMMAU-mini-do-not-useWARNING: The original dataset is revised and pleased refer to new data source. Please refer to: MMAU-v05.15.25: https://github.com/Sakshi113/MMAU
@misc{sakshi2024mmaumassivemultitaskaudio,
title={MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark},
author={S Sakshi and Utkarsh Tyagi and Sonal Kumar and Ashish Seth and Ramaneswaran Selvakumar and Oriol Nieto and Ramani Duraiswami and Sreyan Ghosh and Dinesh Manocha},
year={2024},
eprint={2410.19168}… See the full description on the dataset page: https://huggingface.co/datasets/AudioLLMs/MMAU-mini-do-not-use.voxcpm2-native-generated-audio-user-ref
VoxCPM2 Native Generated Audio (User Ref)
Raw audio files generated from the native VoxCPM2 path in sglang-omni using a user-provided reference clip.
Contents
9 generated .wav files
metadata.json with prompt text, mode, status, size, and latency
Source Reference Audio
Reference clip used for the reference-mode generations:
https://huggingface.co/datasets/adarshxs/voxcpm2-native-test-samples/resolve/main/data/audio.wav
Files
ref_expressive.wav… See the full description on the dataset page: https://huggingface.co/datasets/adarshxs/voxcpm2-native-generated-audio-user-ref.usev-vox2-preprocessedUsethisguangzhou-daily-use-speechASR-SCCantDuSC: A Scripted Chinese Cantonese (Canton) Daily-use Speech Corpus
This open-source dataset consists of 4.06 hours of transcribed Guangzhou Cantonese scripted speech focusing on daily use sentences, where 4,060 utterances contributed by ten speakers were contained.
Source: https://magichub.com/datasets/guangzhou-cantonese-scripted-speech-corpus-daily-use-sentence/
singleWordaudio-user-studyGlobal-Conversational-Speech
Global Conversational Speech Dataset
305 hours. 18 locales. Real conversations.
Not scraped from YouTube. Not recorded by anonymous crowds who don't speak the language. Every conversation in this dataset traces back to verified native speakers we know by name.
The [Human] Standard
Most speech datasets are built the same way: scrape the internet, hire anonymous contractors, run it through automated QC, ship it. The result? Models that are confidently wrong.
We… See the full description on the dataset page: https://huggingface.co/datasets/UsergyAI/Global-Conversational-Speech.singleWordSmallscripted-malay-daily-use-speech-corpus
scripted-malay-daily-use-speech-corpus
Mirror for https://magichub.com/datasets/malay-scripted-speech-corpus-daily-use-sentence/, license is Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License
raw-donot-use-NepaliParliamentDStest-userLorem ipsummalaya-speech-malay-stt-4kus-erica-higgs-metadata1-v1SingleCommerged-dataset-4kOnline-500Manualscripted-malay-daily-use-speech-corpus-whisper-formatus-erica-higgs-metadata1-v2pine-tau2-voice-gpt41-usersimmusicality_useruser_5476d2c924204b6f9e38713118fdb9b2_datasetuser_03aa5df890b64866be4aef51a01c0a8a_datasetmalaya-speech-malay-stt-2ktotalaudio-labelled-ssSingleWordUSTSuser_35621758bf084337aad673e1cc332d6f_dataset
