CoolFace
Datasetpublic

Prajwal-143/ASR-Tamil-cleaned

Dataset Card for Dataset Name Dataset Details Dataset Description This dataset is a combination of the Common Voice 16.0 and Open SLR datasets which is of 534 hours. It has been meticulously curated, normalized to a 16kHz sampling rate, and cleaned for better usability. This dataset aims to provide a comprehensive collection of speech data for various applications, including speech recognition, natural language processing, and machine learning… See the full description on the dataset page: https://huggingface.co/datasets/Prajwal-143/ASR-Tamil-cleaned.

sourceHugging Faceupdated 2y agoView on Hugging Face
3likes181downloads
Dataset Card

Dataset Card for Dataset Name

Dataset Details

Dataset Description

<!-- Provide a longer summary of what this dataset is. --> This dataset is a combination of the Common Voice 16.0 and Open SLR datasets which is of 534 hours. It has been meticulously curated, normalized to a 16kHz sampling rate, and cleaned for better usability. This dataset aims to provide a comprehensive collection of speech data for various applications, including speech recognition, natural language processing, and machine learning research.

  • Curated by: Prajwal N. Pharande
  • Language: Tamil

Dataset Sources

<!-- Provide the basic links for the dataset. -->

  • Repository: https://commonvoice.mozilla.org/ta/datasets and https://www.openslr.org/127/

Uses

This dataset can be used for a wide range of applications, including:

  • Speech recognition system training and evaluation
  • Natural language processing tasks involving spoken language
  • Machine learning research on speech-related problems
  • Voice synthesis and voice cloning experiments

Dataset Structure

  • ``path : Name of audio file which is converted into array.``
  • ``audio : Dictionary contaning path, array and sampling rate of an audio file.``
  • ``sentence : Transcription of an audio file in Tamil language. ``

Dataset Creation

Data Collection and Processing
  • Audio converted to arrays : All audio samples have been normalized to a 16kHz sampling rate, ensuring consistency and high quality across the dataset.
  • Diverse Sources : The dataset combines data from Common Voice and Open SLR, providing a diverse range of voices, accents, and languages.
  • Cleaned Data : Extensive efforts have been made to clean the data, removing noise, punctuations, duplicates, and irrelevant metadata, to enhance usability and accuracy.
Who are the source data producers?

<!-- This section describes the people or systems who originally created the data. It should also include self-reported demographic or identity information for the source data creators if this information is available. -->

  • ``Mozilla`` : A large-scale, publicly available dataset of speech data collected by Mozilla, contributed by volunteers worldwide.
  • ``Open SLR`` : Various open speech and language resources collected and shared by the open-source community through Open Speech and Language Resources.

Dataset Card Authors

Prajwal N. Pharande

Dataset Card Contact

pharandeprajwal@gmail.com