Peacockery/common-voice-scripted-speech-26
Common Voice Scripted Speech A row-normalized multilingual ASR dataset built from Mozilla Data Collective Common Voice Scripted Speech. Each upstream archive is converted to appendable parquet shards under data/<upstream_split>/, one shard per source archive and split, with audio bytes embedded in an audio struct column. Status Manifest languages: 60 Languages uploaded: 18 Columns audio (bytes, path) sentence, locale, language, upstream_split… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/common-voice-scripted-speech-26.
Common Voice Scripted Speech
A row-normalized multilingual ASR dataset built from Mozilla Data Collective Common Voice Scripted Speech. Each upstream archive is converted to appendable parquet shards under data/<upstream_split>/, one shard per source archive and split, with audio bytes embedded in an audio struct column.
Status
- Manifest languages: 60
- Languages uploaded: 18
Columns
audio(bytes,path)sentence,locale,language,upstream_splitsource_dataset_id,source_archive,collection- upstream Common Voice metadata:
client_id,sentence_id,sentence_domain,up_votes,down_votes,age,gender,accents,variant,segment duration_msfromclip_durations.tsvwhen availablelicense,license_url
License
Mozilla Data Collective records these archives as Creative Commons Zero v1.0 Universal (CC0-1.0). See LICENSE and https://spdx.org/licenses/CC0-1.0.html.
