datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
bharatvani-hindi-speech-corpus
BharatVani Hindi Speech Corpus (150-Hour Studio Dataset)
Proprietary Speech Asset • TheCreatorOS • BharatVani AI
1. Overview
The BharatVani Hindi Speech Corpus is an enterprise-grade, high-fidelity Indian speech dataset engineered specifically for training sovereign neural Text-to-Speech (TTS) models, voice cloning engines, and speech foundation models in Devanagari Hindi.
Audio Clips: 103,784 Verified Studio Audio Clips (24,000 Hz, 16-bit Mono… See the full description on the dataset page: https://huggingface.co/datasets/Sheeba2026/bharatvani-hindi-speech-corpus.bharatvani-hindi-showcase
BharatVani Hindi Speech Corpus • Public Interactive Showcase
150-Hour Enterprise Devanagari Hindi Speech Corpus & Precomputed Latents
Curated & Mastered by BharatVani AI • TheCreatorOS
1. Interactive Dataset Preview
This repository is the official public evaluation showcase for the 150-Hour BharatVani Hindi Speech Corpus (103,784 Studio Clips).
Use the Dataset Viewer above to play real audio clips, inspect the word-level timestamp alignments, and… See the full description on the dataset page: https://huggingface.co/datasets/Sheeba2026/bharatvani-hindi-showcase.hindi-speech-instruct
Hindi Conversational Speech Dataset
Multi-turn Hindi conversational dataset for training speech language models. Each conversation has user audio turns paired with assistant text responses.
Dataset Statistics
Metric
Value
Total conversations
10
Total user turns (audio files)
25
Total assistant turns
25
Min turns per conversation
1
Max turns per conversation
4
Avg user text length
32.4 chars
Avg assistant text length
94.0 chars… See the full description on the dataset page: https://huggingface.co/datasets/somu9/hindi-speech-instruct.
