CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01SPRINGLab /IndicTTS-Hindi Hindi Indic TTS Dataset This dataset is derived from the Indic TTS Database project, specifically using the Hindi monolingual recordings from both male and female speakers. The dataset contains high-quality speech recordings with corresponding text transcriptions, making it suitable for text-to-speech (TTS) research and development. Dataset Details Language: Hindi Total Duration: ~10.33 hours (Male: 5.16 hours, Female: 5.18 hours) Audio Format: WAV Sampling Rate: 48000Hz… See the full description on the dataset page: https://huggingface.co/datasets/SPRINGLab/IndicTTS-Hindi.audiotext-to-speech10K<n<100K36 likes1.3k downloads2y agoHugging Face02SPRINGLab /IndicTTS-Englishaudio100K<n<1M2 likes1.2k downloads2y agoHugging Face03SPRINGLab /IndicVoices-R_Hindiaudiotext-to-speech10K<n<100K11 likes631 downloads2y agoHugging Face04SPRINGLab /SPRING_INX_Malayalam_R1audio10K<n<100K0 likes548 downloads2y agoHugging Face05SPRINGLab /IndicTTS_Bengali Bengali Indic TTS Dataset This dataset is derived from the Indic TTS Database project, specifically using the Bengali monolingual recordings from both male and female speakers. The dataset contains high-quality speech recordings with corresponding text transcriptions, making it suitable for text-to-speech (TTS) research and development. Dataset Details Language: Bengali Total Duration: ~15.06 hours (Male: 10.05 hours, Female: 5.01 hours) Audio Format: WAV Sampling Rate:… See the full description on the dataset page: https://huggingface.co/datasets/SPRINGLab/IndicTTS_Bengali.audiotext-to-speech10K<n<100K3 likes536 downloads2y agoHugging Face06SPRINGLab /BPCC_cleanedA curated subset of Bharat Parallel Corpus Collection (BPCC) for 8 Indian languages. Translation pairs are filtered with LABSE score(>0.9) and further preprocessed. Useful for training high-quality translation models. texttranslation1M<n<10M1 likes502 downloads2y agoHugging Face07SPRINGLab /IndicTTS_Gujarati task_categories: - text-to-speech language: - gj pretty_name: Gujarati Indic TTS dataset size_categories: - n<1K Gujarati Indic TTS Dataset This dataset is derived from the Indic TTS Database project, specifically using the Gujarati monolingual recordings from both male and female speakers. The dataset contains high-quality speech recordings with corresponding text transcriptions, making it suitable for text-to-speech (TTS) research and development. Gujarati… See the full description on the dataset page: https://huggingface.co/datasets/SPRINGLab/IndicTTS_Gujarati.audio1K<n<10K3 likes474 downloads2y agoHugging Face08SPRINGLab /Hindi-1482Hrsaudio100K<n<1M5 likes468 downloads2y agoHugging Face09SPRINGLab /IndicTTS_Telugu Telugu Indic TTS Dataset This dataset is derived from the Indic TTS Database project, specifically using the Telugu monolingual recordings from both male and female speakers. The dataset contains high-quality speech recordings with corresponding text transcriptions, making it suitable for text-to-speech (TTS) research and development. Dataset Details Language: Telugu Total Duration: ~8.74 hours (Male: 4.47 hours, Female: 4.27 hours) Audio Format: WAV Sampling Rate:… See the full description on the dataset page: https://huggingface.co/datasets/SPRINGLab/IndicTTS_Telugu.audiotext-to-speech1K<n<10K7 likes413 downloads2y agoHugging Face10SPRINGLab /IndicTTS_Tamil Tamil Indic TTS Dataset This dataset is derived from the Indic TTS Database project, specifically using the Tamil monolingual recordings from both male and female speakers. The dataset contains high-quality speech recordings with corresponding text transcriptions, making it suitable for text-to-speech (TTS) research and development. Dataset Details Language: Tamil Total Duration: ~20.33 hours (Male: 10.3 hours, Female: 10.03 hours) Audio Format: WAV Sampling Rate:… See the full description on the dataset page: https://huggingface.co/datasets/SPRINGLab/IndicTTS_Tamil.audiotext-to-speech1K<n<10K4 likes412 downloads2y agoHugging Face11SPRINGLab /SPRING_INX_Odia_R1audio10K<n<100K1 likes339 downloads2y agoHugging Face12SPRINGLab /IndicTTS_Assamese Assamese Indic TTS Dataset This dataset is derived from the Indic TTS Database project, specifically using the Assamese monolingual recordings from both male and female speakers. The dataset contains high-quality speech recordings with corresponding text transcriptions, making it suitable for text-to-speech (TTS) research and development. Dataset Details Language: Assamese Total Duration: ~27.4 hours (Male: 5.16 hours, Female: 5.18 hours) Audio Format: WAV Sampling Rate:… See the full description on the dataset page: https://huggingface.co/datasets/SPRINGLab/IndicTTS_Assamese.audiotext-to-speech10K<n<100K0 likes315 downloads2y agoHugging Face13SPRINGLab /SPRING_INX_Bengali_R1audio100K<n<1M0 likes297 downloads2y agoHugging Face14SPRINGLab /SPRING_INX_Bengali_R2audio100K<n<1M1 likes296 downloads2y agoHugging Face15SPRINGLab /SPRING_INX_Punjabi_R2audio100K<n<1M0 likes283 downloads2y agoHugging Face16SPRINGLab /IndicTTS_Marathi Marathi Indic TTS Dataset This dataset is derived from the Indic TTS Database project, specifically using the Marathi monolingual recordings from both male and female speakers. The dataset contains high-quality speech recordings with corresponding text transcriptions, making it suitable for text-to-speech (TTS) research and development. Dataset Details Language: Marathi Total Duration: ~10.33 hours (Male: 5.16 hours, Female: 5.18 hours) Audio Format: WAV Sampling Rate:… See the full description on the dataset page: https://huggingface.co/datasets/SPRINGLab/IndicTTS_Marathi.audiotext-to-speech10K<n<100K1 likes266 downloads2y agoHugging Face17speechlab /SPRING_INX_R1tabular1K<n<10K0 likes258 downloads2y agoHugging Face18SPRINGLab /IndicTTS_Kannada Kannada Indic TTS Dataset This dataset is derived from the Indic TTS Database project, specifically using the Kannada monolingual recordings from both male and female speakers. The dataset contains high-quality speech recordings with corresponding text transcriptions, making it suitable for text-to-speech (TTS) research and development. Dataset Details Language: Kannada Total Duration: ~7.35 hours (Male: 3.4 hours, Female: 3.95 hours) Audio Format: WAV Sampling Rate:… See the full description on the dataset page: https://huggingface.co/datasets/SPRINGLab/IndicTTS_Kannada.audiotext-to-speech1K<n<10K4 likes246 downloads2y agoHugging Face19SPRINGLab /IndicVoices-R_Tamilaudio10K<n<100K3 likes234 downloads2y agoHugging Face20SPRINGLab /SPRING_INX_Tamil_R2audio10K<n<100K0 likes186 downloads2y agoHugging Face21SPRINGLab /SPRING_INX_Gujarati_R2audio10K<n<100K0 likes169 downloads2y agoHugging Face22SPRINGLab /IndicVoices-R_Bengaliaudio10K<n<100K0 likes162 downloads2y agoHugging Face23SPRINGLab /SPRING_INX_Assamese_R1audio10K<n<100K0 likes154 downloads2y agoHugging Face24SPRINGLab /SPRING_INX_Marathi_R2audio10K<n<100K1 likes152 downloads2y agoHugging Face25SPRINGLab /IndicTTS_Malayalam Malayalam Indic TTS Dataset This dataset is derived from the Indic TTS Database project, specifically using the Malayalam monolingual recordings from both male and female speakers. The dataset contains high-quality speech recordings with corresponding text transcriptions, making it suitable for text-to-speech (TTS) research and development. Dataset Details Language: Malayalam Total Duration: ~17.89 hours (Male: 9.7 hours, Female: 8.19 hours) Audio Format: WAV Sampling… See the full description on the dataset page: https://huggingface.co/datasets/SPRINGLab/IndicTTS_Malayalam.audiotext-to-speech10K<n<100K2 likes142 downloads2y agoHugging Face26SPRINGLab /IndicTTS_Manipuri Manipuri Indic TTS Dataset This dataset is derived from the Indic TTS Database project, specifically using the Manipuri monolingual recordings from both male and female speakers. The dataset contains high-quality speech recordings with corresponding text transcriptions, making it suitable for text-to-speech (TTS) research and development. Dataset Details Language: Manipuri Total Duration: ~20.75 hours (Male: 10.61 hours, Female: 10.14 hours) Audio Format: WAV Sampling… See the full description on the dataset page: https://huggingface.co/datasets/SPRINGLab/IndicTTS_Manipuri.audiotext-to-speech10K<n<100K2 likes133 downloads1y agoHugging Face27springofwindslabs /mcp-agent-trajectory-benchmark ⚡ Model Context Protocol (MCP) & Advanced Tool‑Use Alignment Tiers 15-second demo: strict JSONL trajectories + 7-point rubric validation (schema stability 100%). Schema Validation Summary Programmatic validation of this exact trial file - reproducible from data.jsonl. What this trial verifies — use these 50 rows to confirm, on your own stack: Schema integrity (strict JSONL, matches the published schema) Multi-turn / tool-use structural consistency… See the full description on the dataset page: https://huggingface.co/datasets/springofwindslabs/mcp-agent-trajectory-benchmark.texttext-generationn<1K0 likes126 downloads10h agoHugging Face28SPRINGLab /IndicTTS_Odia Odia Indic TTS Dataset This dataset is derived from the Indic TTS Database project, specifically using the Odia monolingual recordings from both male and female speakers. The dataset contains high-quality speech recordings with corresponding text transcriptions, making it suitable for text-to-speech (TTS) research and development. Dataset Details Language: Odia Total Duration: ~8.74 hours (Male: 4.47 hours, Female: 4.27 hours) Audio Format: WAV Sampling Rate: 48000Hz… See the full description on the dataset page: https://huggingface.co/datasets/SPRINGLab/IndicTTS_Odia.audiotext-to-speech10K<n<100K1 likes125 downloads2y agoHugging Face29SPRINGLab /LibriSpeech-100audio10K<n<100K2 likes117 downloads1y agoHugging Face30SPRINGLab /LibriSpeech-Testaudio1K<n<10K0 likes114 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.