datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gdpval_preference_rubricsdaily-bio-newsIndic_New_dataset_TTS
Indic TTS Dataset Hub (Mozilla)
Validated audio–text pairs for multiple Indic languages from Mozilla Common Voice.
Select the language from the Subset dropdown in the Dataset Viewer.
Columns
audio: WAV audio clip (16kHz)
text: transcription
duration: length in seconds
speaking_rate: characters per second
khmer_speech_news_dataset
Khmer Speech Dataset Processing
This repository contains scripts and instructions for preparing a Khmer speech dataset for machine learning tasks, such as automatic speech recognition (ASR). It demonstrates how to process a collection of audio files and metadata, and save them as Parquet files for efficient use in your training pipelines—without needing torchcodec.
All dataset audio and transcripts in this project are sourced from https://wmc.org.kh/, the official website of… See the full description on the dataset page: https://huggingface.co/datasets/vichetkao/khmer_speech_news_dataset.
