datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
L2-LibriTTSRL2Arctic
L2-ARCTIC: a non-native English speech corpus
L2-ARCTIC contains English speech from 24 non-native speakers
of Vietnamese, Korean, Mandarin, Spanish, Hindi, and Arabic backgrounds.
It contains phonemic annotations using the sounds supported by ARPABet.
It was compiled by researchers at Texas A&M University and Iowa State University.
Read more on their official website.
This Processed Version
We have processed the dataset into an easily consumable Hugging Face dataset… See the full description on the dataset page: https://huggingface.co/datasets/KoelLabs/L2Arctic.RPMC-L2
RPMC-L2
Paper | Code | Model | Zenodo
The Rock, Punk, Metal, and Core - Livehouse Lighting (RPMC-L2) dataset contains synchronized music and lighting data collected from professional live performance venues. This is the first stage lighting dataset designed to treat Automatic Stage Lighting Control (ASLC) as a generative task, introduced in the paper "Automatic Stage Lighting Control: Is it a Rule-Driven Process or Generative Task?".
Dataset Structure
Data… See the full description on the dataset page: https://huggingface.co/datasets/RS2002/RPMC-L2.l2-arctic-manual-v5.0-16k
l2-arctic-manual-v5.0-16k
This dataset is a prepared derivative of L2-ARCTIC v5.0 that keeps only
the manually annotated material and converts the audio to 16 kHz mono FLAC.
It is designed to plug into the current peacock-asr training code, which
can consume a Hugging Face dataset with audio plus phonemes.
Included splits
train: 1800 rows, 1.84 hours
validation: 899 rows, 0.94 hours
test: 900 rows, 0.88 hours
suitcase: 22 rows, 0.44 hours
The scripted subset uses the… See the full description on the dataset page: https://huggingface.co/datasets/chikingsley/l2-arctic-manual-v5.0-16k.L2ARCTIC-ComplementPhoneticQA-L2ARCTIC
PhoneticQA-L2ARCTIC v0.1
PhoneticQA-L2ARCTIC is a small, word-level multiple-choice AudioQA probe derived from L2-ARCTIC. It was created and released by Haopeng Geng to study the gap between phoneme-level error annotations and perceptually salient mispronunciations.
Benchmark design
120 full-utterance audio questions: 24 dev and 96 test.
Each item shows the canonical transcript and four candidate words.
The task is to select the word that sounds most clearly… See the full description on the dataset page: https://huggingface.co/datasets/Haopeng/PhoneticQA-L2ARCTIC.TTS_L2-regular-TTS_ls960-testlangsnap-l2-frenchl2-arctic-cleaned
L2 Arctic Cleaned
Cleaned version of the L2 Arctic dataset.It includes English L2 speech recordings split into train and test, with aligned metadata CSVs.
Test data contains the recording data of two Vietnamese people
(HQTV-Male, PNV-Female)
Dataset Structure
train/
train/audio/ — train audio files (.wav)
train/trainset.csv — train data with columns: Label,Canonical,Transcript,Error
train/metadata.csv — metadata with columns: file_name,Label
test/
test/audio/ —… See the full description on the dataset page: https://huggingface.co/datasets/vuihocrnd/l2-arctic-cleaned.TTS_L2-regular-ties_ls960-testTTS_L2-regular-dare_ls960-testl2-arctic-dataset-250l2arcticperceived-pr
L2-ARCTIC (Perceived)
[!NOTE]
This dataset is licensed under CC BY-NC 4.0.
If you use this dataset, please remember to cite:
@inproceedings{zhao2018l2arctic, % curated the dataset
title={L2-ARCTIC: A Non-native English Speech Corpus},
author={Zhao, Guanlong and Sonsaat, Sinem and Silpachai, Alif and Lucic, Ivana and Chukharev-Hudilainen, Evgeny and Levis, John and Gutierrez-Osuna, Ricardo},
booktitle={Proc. Interspeech 2018},
pages={2783--2787},
year={2018}
}… See the full description on the dataset page: https://huggingface.co/datasets/changelinglab/l2arcticperceived-pr.speechocean-l2eval
speechocean762: A non-native English corpus for pronunciation scoring task
Dataset Summary
speechocean762 is an open-source non-native English speech corpus designed for pronunciation assessment and L2 spoken proficiency modeling.
This Hugging Face version provides sentence-level audio and expert scores, organized into standard train / validation / test splits.
All speakers are Mandarin L1 learners of English, spanning both children and adults. Each utterance is evaluated… See the full description on the dataset page: https://huggingface.co/datasets/changelinglab/speechocean-l2eval.TTS_L2-regular-SQA_ls960-testTTS_L2-regular-linear_ls960-testTTS_L2-regular-SQA-14_ls960-testTTS_L2-regular-SQA-15_ls960-testaudio_L2-regular-dare_spoken-web-questionsl2-arctic-dataset-250audio_L2-regular_spoken-web-questionsL2EnglishScoring_speechocean762_fluencyaudio_L2-regular-14_spoken-web-questionsL2EnglishFluency_speechocean762-ScoringL2EnglishFluency_speechocean762-Rankingaudio_L2-regular-15_spoken-web-questionsaudio_L2-regular-linear_spoken-web-questionsaudio_L2-regular-ties_llama-questionsaudio_L2-regular-14_trivia_qa-audioL2EnglishScoring_speechocean762_prosodic
