datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
microsoft-speech-corpus-indian
Microsoft Speech Corpus – Indian Languages
Dataset Description
This dataset is a redistribution of the Microsoft Speech Corpus (Indian Languages) containing conversational and phrasal speech training and test data for Telugu, Tamil, and Gujarati languages. Each entry includes an audio recording and its corresponding transcript.
Attribution required: "Data provided by Microsoft and SpeechOcean.com"
⚠️ License: This data is provided for research purposes only. Commercial… See the full description on the dataset page: https://huggingface.co/datasets/deepdml/microsoft-speech-corpus-indian.indian-language-deepfake-speech
Multilingual Deepfake Speech Dataset (Indian Languages)
Overview
This dataset contains real and synthetic speech across four Indian languages:
Telugu
Tamil
Malayalam
Konkani
It is designed for deepfake speech detection and cross-language generalization research.
Composition
Real speech: OpenSLR datasets
Synthetic speech:
MMS-TTS (neural TTS)
SIGVC (signal-based transformations)
RVC (voice conversion)
Total samples: ~15,000+
Average duration: ~5–7… See the full description on the dataset page: https://huggingface.co/datasets/satwc-reddy/indian-language-deepfake-speech.Indian-Emotional-Speech-Corpus
Indian Emotional Speech Corpus
Dataset Description
This dataset comprises high-quality audio recordings of Indian speakers reading a standardized 50-word paragraph in four distinct emotional tones — happy, sad, surprised, and angry.
Each recording is approximately 20–25 seconds long and includes the full paragraph with tone shifts at specific points.
Text spoken by all participants:
(happy tone) Last Monday was perfect—I got the job I’d been dreaming of! I screamed… See the full description on the dataset page: https://huggingface.co/datasets/humyn-labs/Indian-Emotional-Speech-Corpus.KAI-indian-emotional-speech-corpus
Indian Emotional Speech Corpus
Dataset Description
This dataset comprises high-quality audio recordings of Indian speakers reading a standardized 50-word paragraph in four distinct emotional tones — happy, sad, surprised, and angry.
Each recording is approximately 20–25 seconds long and includes the full paragraph with tone shifts at specific points.
Text spoken by all participants:
(happy tone) Last Monday was perfect—I got the job I’d been dreaming of! I screamed… See the full description on the dataset page: https://huggingface.co/datasets/sudhanshu12/KAI-indian-emotional-speech-corpus.indian-speech-audio-extendedindian_speech_audioIndian_English_Speech_Recognition_Corpus_Conversations
ID
King-ASR-631
Language
English
Duration
200 hours
Speakers
200 People
Parameters
16kHz, 16bits
Recording Device
Mobile
URL
https://dataoceanai.com/datasets/asr/indian-english-speech-recognition-corpus-conversations-mobile/
Indian_english_Example_speech
