datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fix_some_err
df_eval
Public evaluation-only speech deepfake detection dataset, organized like Common Voice language configs:
each config is a standard eval protocol (ASVspoof, ADD, In-the-Wild, …) with embedded audio.
Companion code: github.com/Shuo-H/df_eval
Configs are added as uploads complete. Declared configs below match currently available parquet shards on the Hub.
Load
from datasets import load_dataset
ds = load_dataset("shuohann/df_eval", name="sonar"… See the full description on the dataset page: https://huggingface.co/datasets/shuohann/fix_some_err.librispeech_asr_trainlibrispeech_asrlibrispeech_asr_testlibrispeech_asr_validationsomethingDatasetsome_audiosLata_Mangeshkar_JIEMPHASSESVietnamese-streamer-voiceI_prefer_you_over_something_GIRL_AUDIOS_SPLITED
Jubin_NautiyalI_prefer_you_over_something_BOY_AUDIOS_SPLITED
somedataset-2-lipvAI_Voice_DatasetsAI_Vocalssome_StableAudiosome_split
