datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
bharatvani-hindi-showcase
BharatVani Hindi Speech Corpus • Public Interactive Showcase
150-Hour Enterprise Devanagari Hindi Speech Corpus & Precomputed Latents
Curated & Mastered by BharatVani AI • TheCreatorOS
1. Interactive Dataset Preview
This repository is the official public evaluation showcase for the 150-Hour BharatVani Hindi Speech Corpus (103,784 Studio Clips).
Use the Dataset Viewer above to play real audio clips, inspect the word-level timestamp alignments, and… See the full description on the dataset page: https://huggingface.co/datasets/Sheeba2026/bharatvani-hindi-showcase.parakeet-hindi-asr
Parakeet Hindi-English Bilingual ASR
Fine-tuning NVIDIA Parakeet TDT 0.6B for bilingual Hindi-English automatic speech recognition.
Quick Start
# Download
pip install huggingface_hub
huggingface-cli download ketav/parakeet-hindi-asr --repo-type dataset --local-dir ./parakeet-hindi-asr
# Install dependencies
pip install nemo_toolkit[asr] bitsandbytes sentencepiece
# Train (after updating paths in config)
cd parakeet-hindi-asr/scripts
python ft_0.6B_hi_v3.py… See the full description on the dataset page: https://huggingface.co/datasets/ketav/parakeet-hindi-asr.hindi-english-codeswitch-dataset
Hindi-English Code-Switch ASR Transcripts
Text transcripts and metadata for a large bilingual Hindi-English code-switch ASR training corpus, used to train Abhisingh-18/hindi-english-codeswitch-asr.
This release contains transcripts and metadata only — no audio files. Audio was sourced from multiple corpora and institutions and is not redistributed here.
Credits
Speech data collection and curation credit: SPRING Lab, IIT Madras.
Contents
File… See the full description on the dataset page: https://huggingface.co/datasets/Abhisingh-18/hindi-english-codeswitch-dataset.
