datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
50hours_Malay_Real-world_Colloquial_Conversation_and_Monologue_Speech_Dataset
BabelSpeech: 50 Hours of Real-World Colloquial Malay ASR Speech Data
This dataset contains 50 hours of high-quality Malay colloquial ASR speech data, reflecting realistic code-switching between Malay and English, as commonly used in everyday communication in Malaysia.
Overview
Content: 50 hours of real-world colloquial Malay speech suitable for ASR fine-tuning and benchmarking.
Metadata: Stored in a separate JSON file, including audio path, duration, text, confidence… See the full description on the dataset page: https://huggingface.co/datasets/BabelSpeech/50hours_Malay_Real-world_Colloquial_Conversation_and_Monologue_Speech_Dataset.real-world-noise-through-zoom
Real-World Noise Through Zoom (RWNTZ)
5 different real world noise settings (bedroom, crowded room, background music, rain, road with cars)
2 different speakers
various microphone distances (6 inches, 24 inches)
32 total samples with different phrases
recorded through Zoom to simulate real-world linguistic fieldwork scenarios
manually verified word level transcriptions
g2p phoneme trancriptions
audio to phoneme trancriptions with a variety of Wav2Vec2 based models
