datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Codemixed_New
Codemixed ASR Dataset
Unified collection of code-mixed ASR datasets.
fluent_speech_commands_synth
Dataset Card for "fluent_speech_commands_synth"
More Information needed
CodecFake
CodecFake: Enhancing Anti-Spoofing Models Against Deepfake Audios from Codec-Based Speech Synthesis Systems
Paper,
Code,
Project Page
Interspeech 2024
TL;DR: We show that better detection of deepfake speech from codec-based TTS systems can be achieved by training models on speech re-synthesized with neural audio codecs.
This dataset is released for this purpose.
See our paper and Github for more details on using our dataset.
Acknowledgement… See the full description on the dataset page: https://huggingface.co/datasets/rogertseng/CodecFake.WSC-Evalcodecfake-audio
Codecfake Dataset
Overview
The Codecfake dataset is a large-scale dataset designed for the detection of Audio Language Model (ALM)-based deepfake audio. This dataset includes millions of audio samples across two languages and various test conditions, tailored specifically for ALM-based audio detection.
Conversion
The original dataset was downloaded from Zenodo and converted to FLAC format to maintain audio quality while reducing file size. The dataset has been… See the full description on the dataset page: https://huggingface.co/datasets/ajaykarthick/codecfake-audio.librispeech_synth
Dataset Card for "librispeech_synth"
More Information needed
ASR_Code_Switch
ASR Code-Switching Benchmark
A curated benchmark of 1,200 code-switching utterances (300 per language pair)
for evaluating commercial ASR systems on multilingual speech with intra-sentential
language switching.
Paper
Benchmarking Commercial ASR Systems on Code-Switching Speech: Arabic, Persian, and German
arXiv link
Language pairs
Split
Language pair
Samples
Scripts
egyptian_arabic_english
Egyptian Arabic–English
300
Arabic + Latin… See the full description on the dataset page: https://huggingface.co/datasets/Perle-ai/ASR_Code_Switch.voice-code-bench
VoiceCodeBench
VoiceCodeBench is a test-only benchmark for evaluating whether automatic
speech recognition (ASR) systems preserve exact structured values in English
workplace speech.
Paper: VoiceCodeBench: Evaluating Exact Structured-Token Recovery in Automatic Speech Recognition
The benchmark targets cases where a transcript is software input: callback
numbers, email addresses, command-line flags, file paths, URLs, account
identifiers, dates, measurements, and similar values… See the full description on the dataset page: https://huggingface.co/datasets/besimple-ai/voice-code-bench.voxceleb1_synthvocal_imitation_synth
Dataset Card for "vocal_imitation_synth"
More Information needed
maestro_synth
Dataset Card for "maestro_synth"
More Information needed
crema_d_synth
Dataset Card for "crema_d_synth"
More Information needed
vocalset_synth
Dataset Card for "vocalset_synth"
More Information needed
librispeech_asr_test_48k_synthvocalset_synthvox_lingua_top10_synthtorgo_synthspeech_accent_archive_synthopensinger_synthlibrispeech_asr_test_synthfluent_speech_commands_femalevoxceleb1_synthNsynth-test_synthnoisy_vctk_16k_synth
Dataset Card for "noisy_vctk_16k_synth"
More Information needed
quesst14_all_synth
Dataset Card for "quesst14_all_synth"
More Information needed
easycall_synthquesst_synth
Dataset Card for "quesst_synth"
More Information needed
vox_lingua_top10_16k_synthaudioset_synth
Dataset Card for "audioset_synth"
More Information needed
esc50_synth
