datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
uzbek-high-quality-10huzbek-high-quality-10h-en-translationhigh_quality_dshigh_quality_spanish_speechhigh_quality_ds_2000high_quality_ds_2500highquality_nepali_female_asr
🗣️ Nepali Female Speech Dataset (SLR43 Subset)
📖 Dataset Summary
This dataset is a restructured and Hugging Face–ready version of the OpenSLR SLR43: Multi-speaker TTS Data for Nepali (ne-NP). It consists of high-quality speech recordings from Nepali female speakers, originally curated by Google LLC.
The dataset has been reformatted for Automatic Speech Recognition (ASR) and Speech-to-Text (STT) fine-tuning, maintaining full Nepali (Devanagari) transcriptions and… See the full description on the dataset page: https://huggingface.co/datasets/Firoj112/highquality_nepali_female_asr.high_quality_ds_New_1000700hrs-hindi-english-hinglish-high-quality-tts-data
700 hrs Hindi, English, Hinglish High-Quality TTS Data
Processed speech data: Hindi, English, and Hinglish (code-mixed). Includes train/val manifests and the preprocessing script used to generate them.
Download
voice_data.zip — Password-protected archive containing:
Data directory — audio files, train.jsonl, val.jsonl
Preprocessing script — Python script used to build this dataset
The zip file is password-protected. Contact the dataset maintainers for access.… See the full description on the dataset page: https://huggingface.co/datasets/shiprocket-ai/700hrs-hindi-english-hinglish-high-quality-tts-data.high-quality-ttshindi-Single-speaker-female-hindi-high-quality
