datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
IndicTTS-Hindi
Hindi Indic TTS Dataset
This dataset is derived from the Indic TTS Database project, specifically using the Hindi monolingual recordings from both male and female speakers. The dataset contains high-quality speech recordings with corresponding text transcriptions, making it suitable for text-to-speech (TTS) research and development.
Dataset Details
Language: Hindi
Total Duration: ~10.33 hours (Male: 5.16 hours, Female: 5.18 hours)
Audio Format: WAV
Sampling Rate: 48000Hz… See the full description on the dataset page: https://huggingface.co/datasets/SPRINGLab/IndicTTS-Hindi.IndicTTS-EnglishIndicVoices-R_HindiSPRING_INX_Malayalam_R1IndicTTS_Bengali
Bengali Indic TTS Dataset
This dataset is derived from the Indic TTS Database project, specifically using the Bengali monolingual recordings from both male and female speakers. The dataset contains high-quality speech recordings with corresponding text transcriptions, making it suitable for text-to-speech (TTS) research and development.
Dataset Details
Language: Bengali
Total Duration: ~15.06 hours (Male: 10.05 hours, Female: 5.01 hours)
Audio Format: WAV
Sampling Rate:… See the full description on the dataset page: https://huggingface.co/datasets/SPRINGLab/IndicTTS_Bengali.BPCC_cleanedA curated subset of Bharat Parallel Corpus Collection (BPCC) for 8 Indian languages.
Translation pairs are filtered with LABSE score(>0.9) and further preprocessed.
Useful for training high-quality translation models.
IndicTTS_Gujarati
task_categories:
- text-to-speech
language:
- gj
pretty_name: Gujarati Indic TTS dataset
size_categories:
- n<1K
Gujarati Indic TTS Dataset
This dataset is derived from the Indic TTS Database project, specifically using the Gujarati monolingual recordings from both male and female speakers. The dataset contains high-quality speech recordings with corresponding text transcriptions, making it suitable for text-to-speech (TTS) research and development.
Gujarati… See the full description on the dataset page: https://huggingface.co/datasets/SPRINGLab/IndicTTS_Gujarati.Hindi-1482HrsIndicTTS_Telugu
Telugu Indic TTS Dataset
This dataset is derived from the Indic TTS Database project, specifically using the Telugu monolingual recordings from both male and female speakers. The dataset contains high-quality speech recordings with corresponding text transcriptions, making it suitable for text-to-speech (TTS) research and development.
Dataset Details
Language: Telugu
Total Duration: ~8.74 hours (Male: 4.47 hours, Female: 4.27 hours)
Audio Format: WAV
Sampling Rate:… See the full description on the dataset page: https://huggingface.co/datasets/SPRINGLab/IndicTTS_Telugu.IndicTTS_Tamil
Tamil Indic TTS Dataset
This dataset is derived from the Indic TTS Database project, specifically using the Tamil monolingual recordings from both male and female speakers. The dataset contains high-quality speech recordings with corresponding text transcriptions, making it suitable for text-to-speech (TTS) research and development.
Dataset Details
Language: Tamil
Total Duration: ~20.33 hours (Male: 10.3 hours, Female: 10.03 hours)
Audio Format: WAV
Sampling Rate:… See the full description on the dataset page: https://huggingface.co/datasets/SPRINGLab/IndicTTS_Tamil.SPRING_INX_Odia_R1IndicTTS_Assamese
Assamese Indic TTS Dataset
This dataset is derived from the Indic TTS Database project, specifically using the Assamese monolingual recordings from both male and female speakers. The dataset contains high-quality speech recordings with corresponding text transcriptions, making it suitable for text-to-speech (TTS) research and development.
Dataset Details
Language: Assamese
Total Duration: ~27.4 hours (Male: 5.16 hours, Female: 5.18 hours)
Audio Format: WAV
Sampling Rate:… See the full description on the dataset page: https://huggingface.co/datasets/SPRINGLab/IndicTTS_Assamese.SPRING_INX_Bengali_R1SPRING_INX_Bengali_R2SPRING_INX_Punjabi_R2IndicTTS_Marathi
Marathi Indic TTS Dataset
This dataset is derived from the Indic TTS Database project, specifically using the Marathi monolingual recordings from both male and female speakers. The dataset contains high-quality speech recordings with corresponding text transcriptions, making it suitable for text-to-speech (TTS) research and development.
Dataset Details
Language: Marathi
Total Duration: ~10.33 hours (Male: 5.16 hours, Female: 5.18 hours)
Audio Format: WAV
Sampling Rate:… See the full description on the dataset page: https://huggingface.co/datasets/SPRINGLab/IndicTTS_Marathi.SPRING_INX_R1IndicTTS_Kannada
Kannada Indic TTS Dataset
This dataset is derived from the Indic TTS Database project, specifically using the Kannada monolingual recordings from both male and female speakers. The dataset contains high-quality speech recordings with corresponding text transcriptions, making it suitable for text-to-speech (TTS) research and development.
Dataset Details
Language: Kannada
Total Duration: ~7.35 hours (Male: 3.4 hours, Female: 3.95 hours)
Audio Format: WAV
Sampling Rate:… See the full description on the dataset page: https://huggingface.co/datasets/SPRINGLab/IndicTTS_Kannada.IndicVoices-R_TamilSPRING_INX_Tamil_R2SPRING_INX_Gujarati_R2IndicVoices-R_BengaliSPRING_INX_Assamese_R1SPRING_INX_Marathi_R2IndicTTS_Malayalam
Malayalam Indic TTS Dataset
This dataset is derived from the Indic TTS Database project, specifically using the Malayalam monolingual recordings from both male and female speakers. The dataset contains high-quality speech recordings with corresponding text transcriptions, making it suitable for text-to-speech (TTS) research and development.
Dataset Details
Language: Malayalam
Total Duration: ~17.89 hours (Male: 9.7 hours, Female: 8.19 hours)
Audio Format: WAV
Sampling… See the full description on the dataset page: https://huggingface.co/datasets/SPRINGLab/IndicTTS_Malayalam.IndicTTS_Manipuri
Manipuri Indic TTS Dataset
This dataset is derived from the Indic TTS Database project, specifically using the Manipuri monolingual recordings from both male and female speakers. The dataset contains high-quality speech recordings with corresponding text transcriptions, making it suitable for text-to-speech (TTS) research and development.
Dataset Details
Language: Manipuri
Total Duration: ~20.75 hours (Male: 10.61 hours, Female: 10.14 hours)
Audio Format: WAV
Sampling… See the full description on the dataset page: https://huggingface.co/datasets/SPRINGLab/IndicTTS_Manipuri.mcp-agent-trajectory-benchmark
⚡ Model Context Protocol (MCP) & Advanced Tool‑Use Alignment Tiers
15-second demo: strict JSONL trajectories + 7-point rubric validation (schema stability 100%).
Schema Validation Summary
Programmatic validation of this exact trial file - reproducible from data.jsonl.
What this trial verifies — use these 50 rows to confirm, on your own stack:
Schema integrity (strict JSONL, matches the published schema)
Multi-turn / tool-use structural consistency… See the full description on the dataset page: https://huggingface.co/datasets/springofwindslabs/mcp-agent-trajectory-benchmark.IndicTTS_Odia
Odia Indic TTS Dataset
This dataset is derived from the Indic TTS Database project, specifically using the Odia monolingual recordings from both male and female speakers. The dataset contains high-quality speech recordings with corresponding text transcriptions, making it suitable for text-to-speech (TTS) research and development.
Dataset Details
Language: Odia
Total Duration: ~8.74 hours (Male: 4.47 hours, Female: 4.27 hours)
Audio Format: WAV
Sampling Rate: 48000Hz… See the full description on the dataset page: https://huggingface.co/datasets/SPRINGLab/IndicTTS_Odia.LibriSpeech-100LibriSpeech-Test
