ragunath-ravi/TamilVoiceCorpus
Tamil Conversational ASR Dataset This is a dataset for Automatic Speech Recognition (ASR) focused on conversational Tamil, collected from various public sources on the web. Each sample is a short audio clip (averaging 10 seconds) paired with its corresponding transcription. Dataset Summary Language: Tamil (ta) Domain: Conversational speech Average Duration per Clip: ~10 seconds Format: Audio (.wav) + text Sample Rate: 16kHz recommended Total Examples: PureVox:… See the full description on the dataset page: https://huggingface.co/datasets/ragunath-ravi/TamilVoiceCorpus.
532
No card is published for this repository, or it could not be fetched from Hugging Face right now.
