datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Indian-languagestts-indian-languages
TTS Indian Languages Dataset
A curated Text-to-Speech training dataset of 168 segments (~66 minutes) of clean, single-speaker audio in Indian English, Hindi, and Telugu — sourced from YouTube and processed using Sarvam AI's ASR and LLM APIs.
Dataset Summary
Language
Code
Duration
Indian English
en-IN
32.4 min
Hindi
hi-IN
24.0 min
Telugu
te-IN
10.0 min
Total
66.4 min
Dataset Structure
Each row contains:
Field
Type… See the full description on the dataset page: https://huggingface.co/datasets/praneeetha/tts-indian-languages.
