datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
L2Arctic
L2-ARCTIC: a non-native English speech corpus
L2-ARCTIC contains English speech from 24 non-native speakers
of Vietnamese, Korean, Mandarin, Spanish, Hindi, and Arabic backgrounds.
It contains phonemic annotations using the sounds supported by ARPABet.
It was compiled by researchers at Texas A&M University and Iowa State University.
Read more on their official website.
This Processed Version
We have processed the dataset into an easily consumable Hugging Face dataset… See the full description on the dataset page: https://huggingface.co/datasets/KoelLabs/L2Arctic.l2-arctic-manual-v5.0-16k
l2-arctic-manual-v5.0-16k
This dataset is a prepared derivative of L2-ARCTIC v5.0 that keeps only
the manually annotated material and converts the audio to 16 kHz mono FLAC.
It is designed to plug into the current peacock-asr training code, which
can consume a Hugging Face dataset with audio plus phonemes.
Included splits
train: 1800 rows, 1.84 hours
validation: 899 rows, 0.94 hours
test: 900 rows, 0.88 hours
suitcase: 22 rows, 0.44 hours
The scripted subset uses the… See the full description on the dataset page: https://huggingface.co/datasets/chikingsley/l2-arctic-manual-v5.0-16k.l2arcticperceived-pr
L2-ARCTIC (Perceived)
[!NOTE]
This dataset is licensed under CC BY-NC 4.0.
If you use this dataset, please remember to cite:
@inproceedings{zhao2018l2arctic, % curated the dataset
title={L2-ARCTIC: A Non-native English Speech Corpus},
author={Zhao, Guanlong and Sonsaat, Sinem and Silpachai, Alif and Lucic, Ivana and Chukharev-Hudilainen, Evgeny and Levis, John and Gutierrez-Osuna, Ricardo},
booktitle={Proc. Interspeech 2018},
pages={2783--2787},
year={2018}
}… See the full description on the dataset page: https://huggingface.co/datasets/changelinglab/l2arcticperceived-pr.speechocean-l2eval
speechocean762: A non-native English corpus for pronunciation scoring task
Dataset Summary
speechocean762 is an open-source non-native English speech corpus designed for pronunciation assessment and L2 spoken proficiency modeling.
This Hugging Face version provides sentence-level audio and expert scores, organized into standard train / validation / test splits.
All speakers are Mandarin L1 learners of English, spanning both children and adults. Each utterance is evaluated… See the full description on the dataset page: https://huggingface.co/datasets/changelinglab/speechocean-l2eval.L2ArcticSpontaneousSplit
L2-ARCTIC Suitcase: a spontaneous non-native English speech corpus
L2-ARCTIC Suitcase Corpus contains English speech from 22 non-native speakers
of Vietnamese, Korean, Mandarin, Spanish, Hindi, and Arabic backgrounds.
It contains phonemic annotations using the sounds supported by ARPABet.
It was compiled by researchers at Texas A&M University and Iowa State University.
Read more on their official website.
This Processed Version
We have processed the dataset into an… See the full description on the dataset page: https://huggingface.co/datasets/KoelLabs/L2ArcticSpontaneousSplit.
