datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
telugu-indicf5-evaluationIndicST
IndicST: Indian Multilingual Translation Corpus For Evaluating Speech Large Language Models
Introduction
IndicST, a new dataset tailored for training and evaluating Speech LLMs for AST tasks (including ASR and TTS), featuring meticulously curated, automatically, and manually verified synthetic data. The dataset offers 10.8k hrs of training data and 1.13k hrs of evaluation data.
Use-Cases
ASR (Speech-to-Text)
Transcribing Indic languages
Handling… See the full description on the dataset page: https://huggingface.co/datasets/krutrim-ai-labs/IndicST.Indic_New_dataset_TTS
Indic TTS Dataset Hub (Mozilla)
Validated audio–text pairs for multiple Indic languages from Mozilla Common Voice.
Select the language from the Subset dropdown in the Dataset Viewer.
Columns
audio: WAV audio clip (16kHz)
text: transcription
duration: length in seconds
speaking_rate: characters per second
indicvc-datasetindic-fusion-hi-mr
Indic-FUSION Hindi-Marathi Pseudo-Parallel Speech
A research pseudo-parallel speech corpus for Indic-FUSION.
Construction
Source speech is sampled from ai4bharat/indicvoices_r.
For each source utterance:
The source dataset transcript is used directly.
NLLB generates a direct Indic-to-Indic target transcript.
ai4bharat/IndicF5 synthesizes target-language speech using the source utterance as the reference prompt.
The target audio is therefore synthetic, not… See the full description on the dataset page: https://huggingface.co/datasets/Harisri/indic-fusion-hi-mr.
