datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
librispeech_asr
Dataset Card for librispeech_asr
Dataset Summary
LibriSpeech is a corpus of approximately 1000 hours of 16kHz read English speech, prepared by Vassil Panayotov with the assistance of Daniel Povey. The data is derived from read audiobooks from the LibriVox project, and has been carefully segmented and aligned.
Supported Tasks and Leaderboards
automatic-speech-recognition, audio-speaker-identification: The dataset can be used to train a model for Automatic… See the full description on the dataset page: https://huggingface.co/datasets/openslr/librispeech_asr.openslr_quranic_asropenslr-140-hq-Kazakh
Kazakh Speech Dataset (KSD)
Identifier: SLR140
Source: https://www.openslr.org/140/
Summary: High-quality open source Kazakh speech corpus developed by the Department of Artificial Intelligence and Big Data of Al-Farabi Kazakh National University (554 hours)
Category: Speech
License: Attribution-ShareAlike 3.0 Unported (CC BY-SA 3.0 US)
About this resource:
High-quality open source Kazakh speech corpus.
The corpus contains about 554 hours of transcribed audio recordings… See the full description on the dataset page: https://huggingface.co/datasets/voice-biomarkers/openslr-140-hq-Kazakh.burmese-speech-refined-openslr-80
Burmese Speech Refined OpenSLR-80
Summary
This dataset is a speech dataset developed based on the original OpenSLR Dataset (SLR80), with the text and audio data carefully reviewed and further refined for Burmese language applications.
In the original OpenSLR Dataset, the Burmese text was transcribed based on how the words were pronounced in the corresponding audio recordings. In this dataset, the original audio and text data were used as a reference, and the text… See the full description on the dataset page: https://huggingface.co/datasets/thantzinphyo/burmese-speech-refined-openslr-80.sinhala-openslr-111hopenslr-sinhala-asr-normopenslr-sinhala-asrOpenSLR-126
Dataset Card for "OpenSLR-126"
More Information needed
openslr-sinhala-asr-depricated-versionopenslr-librispeech_asr-clean100-00000-19999-km-translatemyanmar-speech-dataset-openslr-80Please visit to the GitHub repository for other Myanmar Langauge datasets.
Myanmar Speech Dataset (OpenSLR-80)
This dataset consists exclusively of Myanmar speech recordings, extracted from the larger multilingual OpenSLR dataset.
For the complete multilingual dataset and additional information, please visit the original dataset repository
of OpenSLR HuggingFace page.
Original Source
OpenSLR is a site devoted to hosting speech and language resources, such as training… See the full description on the dataset page: https://huggingface.co/datasets/chuuhtetnaing/myanmar-speech-dataset-openslr-80.openslr65-tamil
OpenSLR-65 – Tamil Transcribed Speech
Source: https://www.openslr.org/65/
This dataset contains transcribed high-quality audio of Tamil sentences recorded
by volunteers. It is part of the OpenSLR collection of
free speech resources for low-resource languages.
The data was collected via the
Appen (formerly Figure Eight / CrowdFlower) crowdsourcing
platform and is intended for use in training automatic speech recognition (ASR)
and text-to-speech (TTS) systems.
Data… See the full description on the dataset page: https://huggingface.co/datasets/deepdml/openslr65-tamil.openslr-147-hq-Nahuatl
Veracruz Orizaba Nahuatl Endangered Language
Identifier: SLR147
Summary: Audio corpus of Orizaba (Veracruz) Nahuatl speech (Glottocode: oriz1235; ISO 639-3: nlv)
Category: Speech
License: Attribution-ShareAlike 3.0 Unported (CC BY-SA 3.0)
About this resource:
The substantive material of this deposit was gathered over a 13-month period from February 2022 to March 2023.
It comprised 657 files totaling approximately 119 hours, 26 minutes, 59 seconds of material. All but 81… See the full description on the dataset page: https://huggingface.co/datasets/voice-biomarkers/openslr-147-hq-Nahuatl.openslr_asropenslr-32-hq-SA-languages-Afrikaans
High quality TTS data for four South African languages - Afrikaans
Source - https://openslr.org/32/
Identifier: SLR32
Summary: Multi-speaker TTS data for four South African languages - Afrikaans
License: Attribution-ShareAlike 4.0 International (CC BY-SA 4.0)
About this resource:
This data set contains multi-speaker high quality transcribed audio data for four languages of South Africa. The data set consists of wave files, and a TSV file transcribing the audio.… See the full description on the dataset page: https://huggingface.co/datasets/voice-biomarkers/openslr-32-hq-SA-languages-Afrikaans.openslr-slr69-ca-trimmed-denoised
Dataset Card for openslr-slr69-ca-denoised
This is a post-processed version of the Catalan subset belonging to the Open Speech and Language Resources (OpenSLR) speech dataset.
Specifically the subset OpenSLR-69.
The original HF🤗 SLR-69 dataset is located here.
Same license is maintained: Attribution-ShareAlike 4.0 International.
Dataset Details
Dataset Description
We processed the data of the Catalan OpenSLR with the following recipe:
Trimming: Long… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/openslr-slr69-ca-trimmed-denoised.OpenSLR-79
Dataset Card for "OpenSLR-79"
More Information needed
fleurs-km-kh-openslr-SLR42
Combined openslr/openslr-SLR42 and google/fleurs-km-kh.
Each audio is shorter than or equal to 30 seconds.
Sampling rate is 16,000
Each transcription is already normalized by tha.normalize.processor (pip install tha)
openslr76
Dataset Card for "openslr76"
More Information needed
open_slr_lang_resourcefleurs_openslr42_mpwtNOTE: If your colab crashes, please use pip install --upgrade --quiet datasets[audio]==3.6.0 to install datasets[audio] version 3.6.0.
This dataset combined google/fleurs, openslr/openslr42, and cleaned seanghay/khmer_mpwt_speech.
Severals processes are executed:
clean up seanghay/khmer_mpwt_speech: manually correct wrong transcriptions over 2058 rows
normalize transcription: remove invisible white space; process ៗ, numbers, currencies, date into khmer text; and separate each word by space… See the full description on the dataset page: https://huggingface.co/datasets/KrorngAI/fleurs_openslr42_mpwt.OpenSLR54-Nepali-ASRopenslr42-khmer-male
OpenSLR SLR42 Khmer Male Speech
This dataset is a processed version of the OpenSLR SLR42 Khmer speech dataset.
Dataset Description
This dataset contains approximately 2,906 Khmer speech recordings with corresponding Khmer transcriptions.
Each example contains:
audio: Khmer speech recording
text: Khmer transcription
Dataset Structure
Column
Type
Description
audio
Audio
Khmer speech recording
text
String
Khmer transcription… See the full description on the dataset page: https://huggingface.co/datasets/phonsobon/openslr42-khmer-male.openslr36-sundanese-asropenslr80-burmese
OpenSLR-80 – Brumese Transcribed Speech
Source: https://www.openslr.org/80/
This dataset contains transcribed high-quality audio of Burmese sentences recorded
by female volunteers. It is part of the OpenSLR collection of
free speech resources for low-resource languages.
The data was collected via the
Appen (formerly Figure Eight / CrowdFlower) crowdsourcing
platform and is intended for use in training automatic speech recognition (ASR)
and text-to-speech (TTS) systems.… See the full description on the dataset page: https://huggingface.co/datasets/deepdml/openslr80-burmese.gujarati-f-openslr
Gujarati OpenSLR Female
Interspeech data downloaded from https://www.openslr.org/resources/78/gu_in_female.zip
Dataset Details
Gujarati Data (Most of the entries are <30 seconds and hence Whisper Models can be used for accurate timestamp prediction)
Also, the audio seems to have been spoken by a single female.
openslr-32-hq-SA-languages
SLR32 – High Quality TTS Data for Four South African Languages
Identifier: SLR32License: CC BY-SA 4.0Source: https://www.openslr.org/32/
This dataset contains multi-speaker high quality transcribed audio data for four
languages of South Africa: Afrikaans (af_za), Sesotho (st_za),
Setswana (tn_za) and isiXhosa (xh_za).
The dataset consists of WAV files and a TSV file transcribing the audio.
In each folder the file line_index.tsv contains a FileID (which in turn
encodes the UserID)… See the full description on the dataset page: https://huggingface.co/datasets/deepdml/openslr-32-hq-SA-languages.openslr61-es-ar-full
openslr61-es-ar-full
OpenSLR 61 (Crowdsourced high-quality Argentinian Spanish) consolidado COMPLETO con linaje.
Incluye male + female + weather messages argentinos.
Linaje (trazabilidad por sample)
source: siempre "openslr61"
subset: "main" (frases generales) o "weather" (mensajes de clima)
gender: "m" / "f"
speaker_id: ID anonimizado original del speaker
file_id: ID original del archivo OpenSLR
license: CC-BY-SA-4.0
Schema
campo
tipo… See the full description on the dataset page: https://huggingface.co/datasets/Kukedlc/openslr61-es-ar-full.telugu_OpenSLROpenSLR54-Nepali-ASR-parquet
OpenSLR 54: Large Nepali ASR training data set (unmodified parquet repackaging)
This is an unofficial repackaging of the official OpenSLR 54 release
(SLR54, https://www.openslr.org/54/), converted to parquet so it can be streamed with 🤗 datasets.
It is not affiliated with or endorsed by OpenSLR or the original authors.
All credit for the data belongs to the original creators (see Citation).
What's inside
157,905 utterances, 16 shards: one per original zip… See the full description on the dataset page: https://huggingface.co/datasets/JeevanDai/OpenSLR54-Nepali-ASR-parquet.
