datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
IARPA_BABEL_OP3_30650hours_Malay_Real-world_Colloquial_Conversation_and_Monologue_Speech_Dataset
BabelSpeech: 50 Hours of Real-World Colloquial Malay ASR Speech Data
This dataset contains 50 hours of high-quality Malay colloquial ASR speech data, reflecting realistic code-switching between Malay and English, as commonly used in everyday communication in Malaysia.
Overview
Content: 50 hours of real-world colloquial Malay speech suitable for ASR fine-tuning and benchmarking.
Metadata: Stored in a separate JSON file, including audio path, duration, text, confidence… See the full description on the dataset page: https://huggingface.co/datasets/BabelSpeech/50hours_Malay_Real-world_Colloquial_Conversation_and_Monologue_Speech_Dataset.40hours_Indonesian_Colloquial_ASR_Speech_Dataset
BabelSpeech 50-Hour Indonesian Colloquial ASR Speech Dataset
Contains 50 hours of Indonesian colloquial ASR speech data, aligned with natural, everyday Indonesian communication patterns.
Metadata is stored in a separate JSON file, including audio path, duration, transcript confidence, signal-to-noise ratio (SNR), and DNSMOS. More metadata fields may be added in future updates.
Covered domains: technology, entertainment, travel, education, daily life, and others.
Data quality: Each… See the full description on the dataset page: https://huggingface.co/datasets/BabelSpeech/40hours_Indonesian_Colloquial_ASR_Speech_Dataset.
