datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TamilVoiceCorpus
Tamil Conversational ASR Dataset
This is a dataset for Automatic Speech Recognition (ASR) focused on conversational Tamil, collected from various public sources on the web. Each sample is a short audio clip (averaging 10 seconds) paired with its corresponding transcription.
Dataset Summary
Language: Tamil (ta)
Domain: Conversational speech
Average Duration per Clip: ~10 seconds
Format: Audio (.wav) + text
Sample Rate: 16kHz recommended
Total Examples:
PureVox:
Train: 12… See the full description on the dataset page: https://huggingface.co/datasets/ragunath-ravi/TamilVoiceCorpus.embedded_world_2026_rag_tts
