datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
microsoft-speech-corpus-indian
Microsoft Speech Corpus – Indian Languages
Dataset Description
This dataset is a redistribution of the Microsoft Speech Corpus (Indian Languages) containing conversational and phrasal speech training and test data for Telugu, Tamil, and Gujarati languages. Each entry includes an audio recording and its corresponding transcript.
Attribution required: "Data provided by Microsoft and SpeechOcean.com"
⚠️ License: This data is provided for research purposes only. Commercial… See the full description on the dataset page: https://huggingface.co/datasets/deepdml/microsoft-speech-corpus-indian.Indian-Emotional-Speech-Corpus
Indian Emotional Speech Corpus
Dataset Description
This dataset comprises high-quality audio recordings of Indian speakers reading a standardized 50-word paragraph in four distinct emotional tones — happy, sad, surprised, and angry.
Each recording is approximately 20–25 seconds long and includes the full paragraph with tone shifts at specific points.
Text spoken by all participants:
(happy tone) Last Monday was perfect—I got the job I’d been dreaming of! I screamed… See the full description on the dataset page: https://huggingface.co/datasets/humyn-labs/Indian-Emotional-Speech-Corpus.KAI-indian-emotional-speech-corpus
Indian Emotional Speech Corpus
Dataset Description
This dataset comprises high-quality audio recordings of Indian speakers reading a standardized 50-word paragraph in four distinct emotional tones — happy, sad, surprised, and angry.
Each recording is approximately 20–25 seconds long and includes the full paragraph with tone shifts at specific points.
Text spoken by all participants:
(happy tone) Last Monday was perfect—I got the job I’d been dreaming of! I screamed… See the full description on the dataset page: https://huggingface.co/datasets/sudhanshu12/KAI-indian-emotional-speech-corpus.indian-speech-audio-extendedindian_speech_audio
