datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mega-asr-noise-a5sv2
Mega-ASR Noise A5SV2
Mega-ASR Noise A5SV2 is a deterministic, English-only robustness evaluation
subset derived from
zhifeixie/Voices-in-the-Wild-2M,
the training corpus released with Mega-ASR.
We sampled from Mega-ASR-Train, rather than the standard Mega-ASR test set,
because in our experiments the standard test set was not acoustically
challenging enough to clearly discriminate among robust ASR systems. This is a
derived evaluation set; it should not be mixed into training… See the full description on the dataset page: https://huggingface.co/datasets/AirCaps/mega-asr-noise-a5sv2.nvidia-brain-noise-evaluation-dataset
Nvidia Brain Noise Evaluation Dataset
Dataset Description
This dataset contains 64 samples organized across multiple splits and 32 subsets.
The dataset includes audio data.
Dataset Structure
Subsets
This dataset includes the following subsets:
noisy-bg-snr-10: 2 samples
test: 2 samples
noisy-bg-snr-20: 2 samples
test: 2 samples
noisy-bg-snr-30: 2 samples
test: 2 samples
noisy-bg-snr-40: 2 samples
test: 2 samples
noisy-bg-snr-50: 2 samples… See the full description on the dataset page: https://huggingface.co/datasets/sujalappa/nvidia-brain-noise-evaluation-dataset.speech-brain-noise-evaluation-dataset
Speech Brain Noise Evaluation Dataset
Dataset Description
This dataset contains 2,000 samples organized across multiple splits and 20 subsets.
The dataset includes audio data.
Dataset Structure
Subsets
This dataset includes the following subsets:
noisy-bg-snr-10: 100 samples
test: 100 samples
noisy-bg-snr-30: 100 samples
test: 100 samples
noisy-bg-snr-50: 100 samples
test: 100 samples
denoised-bg-snr-10: 100 samples
test: 100 samples… See the full description on the dataset page: https://huggingface.co/datasets/sujalappa/speech-brain-noise-evaluation-dataset.real-world-noise-through-zoom
Real-World Noise Through Zoom (RWNTZ)
5 different real world noise settings (bedroom, crowded room, background music, rain, road with cars)
2 different speakers
various microphone distances (6 inches, 24 inches)
32 total samples with different phrases
recorded through Zoom to simulate real-world linguistic fieldwork scenarios
manually verified word level transcriptions
g2p phoneme trancriptions
audio to phoneme trancriptions with a variety of Wav2Vec2 based models
singaporean_district_noise
Singaporean district with noise
Dataset Description
Singaporean district speech dataset with controlled noise augmentation for ASR training
Dataset Summary
Language: EN
Task: Automatic Speech Recognition
Total Samples: 2,288
Audio Sample Rate: 16kHz
Base Dataset: Custom dataset
Processing: Noise-augmented
Dataset Structure
Data Fields
audio: Audio file (16kHz WAV format)
text: Transcription text
noise_type: Type of background noise… See the full description on the dataset page: https://huggingface.co/datasets/thucdangvan020999/singaporean_district_noise.isizulu-asr-gaussian-noise
isiZulu Speech Recognition Augmented Dataset - Gaussian Noise
Dataset Description
This dataset contains augmented speech recordings and transcriptions for isiZulu, one of South Africa's official languages.
The dataset has been optimized for use with OpenAI's Whisper ASR models.
Dataset Statistics
Number of samples: 623
Language: isiZulu (Zul)
Audio format: WAV, 16kHz, mono, 16-bit
Maximum duration: 30 seconds (truncated for Whisper compatibility)
Transcription… See the full description on the dataset page: https://huggingface.co/datasets/zionia/isizulu-asr-gaussian-noise.asr-eval-noise-clean-mix-v1-1k
ASR Eval — Restaurant Speech v1 (1K)
A 1,000-sample English speech benchmark dataset recorded in a real restaurant environment, designed to evaluate ASR systems under challenging real-world acoustic conditions.
Dataset Summary
Each audio sample was recorded in a restaurant setting, capturing natural speech alongside the ambient sounds typical of a busy dining environment — background conversations, cutlery, and general crowd noise. This makes it an ideal benchmark for… See the full description on the dataset page: https://huggingface.co/datasets/ultrasafe-ai/asr-eval-noise-clean-mix-v1-1k.singaporean_district_noise_snr_5_10
Singaporean district with noise
Dataset Description
Singaporean district speech dataset with controlled noise augmentation for ASR training
Dataset Summary
Language: EN
Task: Automatic Speech Recognition
Total Samples: 252
Audio Sample Rate: 16kHz
Base Dataset: Custom dataset
Processing: Noise-augmented
Dataset Structure
Data Fields
audio: Audio file (16kHz WAV format)
text: Transcription text
noise_type: Type of background noise… See the full description on the dataset page: https://huggingface.co/datasets/thucdangvan020999/singaporean_district_noise_snr_5_10.isixhosa-asr-gaussian-noise
isiXhosa Speech Recognition Augmented Dataset - Gaussian Noise
Dataset Description
This dataset contains augmented speech recordings and transcriptions for isiXhosa, one of South Africa's official languages.
The dataset has been optimized for use with OpenAI's Whisper ASR models.
Dataset Statistics
Number of samples: 667
Language: isiXhosa (Xho)
Audio format: WAV, 16kHz, mono, 16-bit
Maximum duration: 30 seconds (truncated for Whisper compatibility)… See the full description on the dataset page: https://huggingface.co/datasets/zionia/isixhosa-asr-gaussian-noise.singaporean_district_noise_snr_2_7
Singaporean district with noise
Dataset Description
Singaporean district speech dataset with controlled noise augmentation for ASR training
Dataset Summary
Language: EN
Task: Automatic Speech Recognition
Total Samples: 2,288
Audio Sample Rate: 16kHz
Base Dataset: Custom dataset
Processing: Noise-augmented
Dataset Structure
Data Fields
audio: Audio file (16kHz WAV format)
text: Transcription text
noise_type: Type of background noise… See the full description on the dataset page: https://huggingface.co/datasets/thucdangvan020999/singaporean_district_noise_snr_2_7.
