datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hindi_audio_dataset_testIndicTTS-Hindi
Hindi Indic TTS Dataset
This dataset is derived from the Indic TTS Database project, specifically using the Hindi monolingual recordings from both male and female speakers. The dataset contains high-quality speech recordings with corresponding text transcriptions, making it suitable for text-to-speech (TTS) research and development.
Dataset Details
Language: Hindi
Total Duration: ~10.33 hours (Male: 5.16 hours, Female: 5.18 hours)
Audio Format: WAV
Sampling Rate: 48000Hz… See the full description on the dataset page: https://huggingface.co/datasets/SPRINGLab/IndicTTS-Hindi.syspin_hindi_mergedHindi-1482HrsIndicVoices-R_Hindihindi_karya_mergedS2R_Shrutilipi_hindi
Paytmlabs/S2R_Shrutilipi_hindi
Hindi speech dataset prepared from ai4bharat/Shrutilipi for Ultravox training.
Viewing samples on Hugging Face
The hindi config stores audio inside Parquet. The website dataset viewer often cannot decode that and shows no rows.
To inspect examples in the browser, open the Subset (config) drop-down and choose hindi_text_samples — text and continuation only (~2000 rows).
Ultravox training should keep using subset hindi (full audio).… See the full description on the dataset page: https://huggingface.co/datasets/Paytmlabs/S2R_Shrutilipi_hindi.hindi-asr-wdsAll_Hindi_ASR_v1.1All_Hindi_ASR_v1.2hindi-english-bilingual
Hindi/English/Hinglish Bilingual TTS Dataset
Synthetic TTS dataset generated by Rani voice (ai4bharat/indic-parler-tts) for training a lightweight bilingual student TTS model.
Designed for natural-sounding Hindi, English, and Hinglish (code-switched) speech synthesis.
Dataset Summary
Property
Value
Total utterances
23,277
Total audio
~4.7GB (24kHz WAV)
Languages
Hindi (hi), English (en), Hinglish (bi)
Sample rate
24kHz
Voice
Rani —… See the full description on the dataset page: https://huggingface.co/datasets/nameissakthi/hindi-english-bilingual.hindi_indic_voice_rbharatvani-hindi-speech-corpus
BharatVani Hindi Speech Corpus (150-Hour Studio Dataset)
Proprietary Speech Asset • TheCreatorOS • BharatVani AI
1. Overview
The BharatVani Hindi Speech Corpus is an enterprise-grade, high-fidelity Indian speech dataset engineered specifically for training sovereign neural Text-to-Speech (TTS) models, voice cloning engines, and speech foundation models in Devanagari Hindi.
Audio Clips: 103,784 Verified Studio Audio Clips (24,000 Hz, 16-bit Mono… See the full description on the dataset page: https://huggingface.co/datasets/Sheeba2026/bharatvani-hindi-speech-corpus.All_Hindi_ASR_v1.1processed_seamless_align_hindi_chunk_3processed_seamless_align_hindi_chunk_1hindi-audio-stories-20-30s
Hindi Audio Stories — 20–30 s clips (Qwen3-ASR, denoised)
Paired (audio, text) Hindi speech dataset. Each ~20–30 s denoised clip has
its transcript in two scripts, stored as separate rows (script column):
devanagari (Hindi) and latin (Hinglish romanization, uroman).
⚠️ Adult (NSFW) content. Research / non-commercial.
Stats
~12.6k clips × 2 scripts ≈ 25k rows · ~94 h · mean 26.8 s (97% in 20–30 s)
Audio: 24 kHz mono FLAC, UVR vocal-isolated (Mel-Band RoFormer —… See the full description on the dataset page: https://huggingface.co/datasets/backpropSukuna/hindi-audio-stories-20-30s.All_Hindi_ASR_Male_v1.1HindiDownloadedData2processed_seamless_align_hindi_chunk_5hindi_speech_10hprocessed_seamless_align_hindi_chunk_2S2R_Kathbhat_hindi
Paytmlabs/S2R_Kathbhat_hindi
Hindi speech dataset prepared from ai4bharat/Kathbath for Ultravox training.
Schema
Column
Type
Description
audio
Audio
Speech audio
text
string
Verbatim transcript
continuation
string
LLM-generated continuation (≤50 words)
Progress
Train chunks: 19/19
Validation: done
processed_seamless_align_hindi_chunk_4merged-hindi-audio-datasetorpheus-tts-hindihindi_dataset_stats_catagorical_description_audioprocessed_seamless_align_hindi_chunk_6Shrutilipi_Hindi_resampled_44100_merged_10hindi-whisper-chunks
Hindi Whisper Chunks
Preprocessed, feature-extracted audio chunks and labels used to fine-tune ArchCoder/whisper-small-hindi-lora, a LoRA adaptation of Whisper-small for Hindi speech recognition.
Dataset Summary
Raw Hindi audio recordings (approximately 12 minutes each) were segmented into short, Whisper-compatible chunks and converted into model-ready features. This dataset is the output of that preprocessing pipeline: Whisper-format log-mel filterbank features… See the full description on the dataset page: https://huggingface.co/datasets/ArchCoder/hindi-whisper-chunks.
