datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
reprocessed_singapore_national_speech_corpus
Dataset Card for Reprocessed National Speech Corpus
NOTE: This is an Reprocessed version KaraKaraWitch from Recursal.The official download can be found here.
Dataset Details
Dataset Description
Dataset Description:
The National Speech Corpus (NSC) is the first large-scale Singapore English corpus, sponsored by the Info-communications and Media Development Authority (IMDA) of Singapore. The objective is to serve as a primary resource of open speech data for… See the full description on the dataset page: https://huggingface.co/datasets/recursal/reprocessed_singapore_national_speech_corpus.singaporean_district_noise
Singaporean district with noise
Dataset Description
Singaporean district speech dataset with controlled noise augmentation for ASR training
Dataset Summary
Language: EN
Task: Automatic Speech Recognition
Total Samples: 2,288
Audio Sample Rate: 16kHz
Base Dataset: Custom dataset
Processing: Noise-augmented
Dataset Structure
Data Fields
audio: Audio file (16kHz WAV format)
text: Transcription text
noise_type: Type of background noise… See the full description on the dataset page: https://huggingface.co/datasets/thucdangvan020999/singaporean_district_noise.singaporean_district_noise_snr_5_10
Singaporean district with noise
Dataset Description
Singaporean district speech dataset with controlled noise augmentation for ASR training
Dataset Summary
Language: EN
Task: Automatic Speech Recognition
Total Samples: 252
Audio Sample Rate: 16kHz
Base Dataset: Custom dataset
Processing: Noise-augmented
Dataset Structure
Data Fields
audio: Audio file (16kHz WAV format)
text: Transcription text
noise_type: Type of background noise… See the full description on the dataset page: https://huggingface.co/datasets/thucdangvan020999/singaporean_district_noise_snr_5_10.singaporean_district_noise_snr_2_7
Singaporean district with noise
Dataset Description
Singaporean district speech dataset with controlled noise augmentation for ASR training
Dataset Summary
Language: EN
Task: Automatic Speech Recognition
Total Samples: 2,288
Audio Sample Rate: 16kHz
Base Dataset: Custom dataset
Processing: Noise-augmented
Dataset Structure
Data Fields
audio: Audio file (16kHz WAV format)
text: Transcription text
noise_type: Type of background noise… See the full description on the dataset page: https://huggingface.co/datasets/thucdangvan020999/singaporean_district_noise_snr_2_7.
