datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
singaporean_district_noise_snr_5_10
Singaporean district with noise
Dataset Description
Singaporean district speech dataset with controlled noise augmentation for ASR training
Dataset Summary
Language: EN
Task: Automatic Speech Recognition
Total Samples: 252
Audio Sample Rate: 16kHz
Base Dataset: Custom dataset
Processing: Noise-augmented
Dataset Structure
Data Fields
audio: Audio file (16kHz WAV format)
text: Transcription text
noise_type: Type of background noise… See the full description on the dataset page: https://huggingface.co/datasets/thucdangvan020999/singaporean_district_noise_snr_5_10.singaporean_district_noise_snr_2_7
Singaporean district with noise
Dataset Description
Singaporean district speech dataset with controlled noise augmentation for ASR training
Dataset Summary
Language: EN
Task: Automatic Speech Recognition
Total Samples: 2,288
Audio Sample Rate: 16kHz
Base Dataset: Custom dataset
Processing: Noise-augmented
Dataset Structure
Data Fields
audio: Audio file (16kHz WAV format)
text: Transcription text
noise_type: Type of background noise… See the full description on the dataset page: https://huggingface.co/datasets/thucdangvan020999/singaporean_district_noise_snr_2_7.
