Prajwal-143/ASR-Tamil-cleaned
Dataset Card for Dataset Name Dataset Details Dataset Description This dataset is a combination of the Common Voice 16.0 and Open SLR datasets which is of 534 hours. It has been meticulously curated, normalized to a 16kHz sampling rate, and cleaned for better usability. This dataset aims to provide a comprehensive collection of speech data for various applications, including speech recognition, natural language processing, and machine learning… See the full description on the dataset page: https://huggingface.co/datasets/Prajwal-143/ASR-Tamil-cleaned.
Dataset Card for Dataset Name
Dataset Details
Dataset Description
<!-- Provide a longer summary of what this dataset is. --> This dataset is a combination of the Common Voice 16.0 and Open SLR datasets which is of 534 hours. It has been meticulously curated, normalized to a 16kHz sampling rate, and cleaned for better usability. This dataset aims to provide a comprehensive collection of speech data for various applications, including speech recognition, natural language processing, and machine learning research.
- Curated by: Prajwal N. Pharande
- Language: Tamil
Dataset Sources
<!-- Provide the basic links for the dataset. -->
- Repository: https://commonvoice.mozilla.org/ta/datasets and https://www.openslr.org/127/
Uses
This dataset can be used for a wide range of applications, including:
- Speech recognition system training and evaluation
- Natural language processing tasks involving spoken language
- Machine learning research on speech-related problems
- Voice synthesis and voice cloning experiments
Dataset Structure
- ``
path : Name of audio file which is converted into array.`` - ``
audio : Dictionary contaning path, array and sampling rate of an audio file.`` - ``
sentence : Transcription of an audio file in Tamil language.``
Dataset Creation
Data Collection and Processing
- Audio converted to arrays : All audio samples have been normalized to a 16kHz sampling rate, ensuring consistency and high quality across the dataset.
- Diverse Sources : The dataset combines data from Common Voice and Open SLR, providing a diverse range of voices, accents, and languages.
- Cleaned Data : Extensive efforts have been made to clean the data, removing noise, punctuations, duplicates, and irrelevant metadata, to enhance usability and accuracy.
Who are the source data producers?
<!-- This section describes the people or systems who originally created the data. It should also include self-reported demographic or identity information for the source data creators if this information is available. -->
- ``
Mozilla`` : A large-scale, publicly available dataset of speech data collected by Mozilla, contributed by volunteers worldwide. - ``
Open SLR`` : Various open speech and language resources collected and shared by the open-source community through Open Speech and Language Resources.
Dataset Card Authors
Prajwal N. Pharande
Dataset Card Contact
pharandeprajwal@gmail.com
