vrclc/openslr63
SLR63: Crowdsourced high-quality Malayalam multi-speaker speech data set This data set contains transcribed high-quality audio of Malayalam sentences recorded by volunteers. The data set consists of wave files, and a TSV file (line_index.tsv). The file line_index.tsv contains a anonymized FileID and the transcription of audio in the file. The data set has been manually quality checked, but there might still be errors. Please report any issues in the following issue tracker on… See the full description on the dataset page: https://huggingface.co/datasets/vrclc/openslr63.
336
1---2license: cc-by-4.03task_categories:4- automatic-speech-recognition5- text-to-speech6language:7- ml8pretty_name: OPEN SLR 639size_categories:10- 1K<n<10K11---12## SLR63: Crowdsourced high-quality Malayalam multi-speaker speech data set13 14This data set contains transcribed high-quality audio of Malayalam sentences recorded by volunteers. The data set consists of wave files, and a TSV file (line_index.tsv). The file line_index.tsv contains a anonymized FileID and the transcription of audio in the file.15 16The data set has been manually quality checked, but there might still be errors.17 18Please report any issues in the following issue tracker on GitHub. https://github.com/googlei18n/language-resources/issues19 20The dataset is distributed under Creative Commons Attribution-ShareAlike 4.0 International Public License. See LICENSE file and https://github.com/google/language-resources#license for license information.21 22Copyright 2018, 2019 Google, Inc. 23 24### Train Test Split created to ensure no speaker overlap