bookbot/common-voice-23-0-es-ipa-v1
Common Voice 23.0 Spanish — canonical IPA v1 This repository is the approved deterministic v1 publication cut of Spanish Common Voice 23.0 recipe cuts. It contains the accepted final Lhotse cuts only: split rows train 352,941 test 15,857 The data files have exactly three columns: id — the unique final Lhotse cut identifier (for example, common_voice_es_19696062-9131). audio — a Hugging Face Audio feature containing the original encoded source bytes and source… See the full description on the dataset page: https://huggingface.co/datasets/bookbot/common-voice-23-0-es-ipa-v1.
Common Voice 23.0 Spanish — canonical IPA v1
This repository is the approved deterministic v1 publication cut of Spanish Common Voice 23.0 recipe cuts. It contains the accepted final Lhotse cuts only:
The data files have exactly three columns:
id— the unique final Lhotse cut identifier (for example,common_voice_es_19696062-9131).audio— a Hugging FaceAudiofeature containing the original encoded source bytes and source basename.ipa_transcript— the exact whitespace-separated canonical v1 phone string fromsupervisions[0].textin the final cut.
The legacy source Parquet field phonemes_ipa (BabyGruut-era output) was not used. The IPA strings here are the v1 canonical recipe supervision, validated against the fixed atomic phone inventory in egs/bookbot_es/ASR/es-tokens.txt.
Source, license, and Mozilla notices
The source is Mozilla Foundation, [Common Voice Scripted Speech 23.0 — Spanish](https://mozilladatacollective.com/datasets/cmj8u3pe600fhnxxbqol7x8bg), obtained through Mozilla Data Collective. The source license is CC0 1.0. The pinned source manifest is common_voice_23_0_es/manifest.tsv, SHA-256 ca5e485fd427ca33802293a228b534e1ebb0a14e6f7187bf04f51dc56702d35c; the pinned source revision is 591a99e9c5c9e308231746c7baa03eadfe670483.
Mozilla’s Common Voice Legal Terms (effective 2025-10-31) state that Common Voice datasets are made available through Mozilla Data Collective under CC0, and ask users not to post, distribute, or mirror Common Voice datasets in whole or in part on other platforms or services. Users of this approved project publication must review and comply with the current Mozilla terms and Privacy Notice. CC0 does not require attribution, but the recommended attribution is: Mozilla Foundation, Common Voice Scripted Speech 23.0 — Spanish, Mozilla Data Collective, CC0 1.0.
This repository is a derived recipe cut, not an official Mozilla release. The source Parquet audio bytes are copied unchanged for every accepted full-recording cut; no resampling, volume, or speed augmentation is materialized. The final cut’s deterministic preparation metadata and file hashes are in `provenance/provenance.json`; the complete one-to-one source join audit is in `provenance/join_manifest.tsv.gz`.
Preparation and transformations
- Source rows were joined by the pinned
audio.pathbasename to the final cut supervision/recording identity. - Every accepted cut has
start = 0and covers its complete source recording. Therefore the embedded encoded source bytes are retained losslessly; there were zero trimmed cuts requiring extraction. idis the final cut ID, not a regenerated source ID.ipa_transcriptis copied exactly from the final canonical supervision and is not reconstructed from orthographic text or the legacyphonemes_ipafield.- The final cut metadata records the training recipe’s optional resample/volume/speed transforms, but those augmentations are intentionally not applied to this source-preserving publication.
The source manifest’s Charsiu raw phonemization is provenance only; downstream v1 normalization and rejection decisions are those recorded in the final cuts. The exact output shard hashes, row counts, cut manifest hashes, source revision, and transformation policy are machine-readable in provenance/provenance.json.
Fixed smoke sample
The deterministic CC0 smoke sample cut is common_voice_es_19696062-9131 in test, sourced from common_voice_es_19696062.mp3. Runtime packages should resample a copied waveform to 16 kHz; this dataset preserves the original encoded source sample rate and bytes.
Citation and attribution
Mozilla Foundation. Common Voice Scripted Speech 23.0 — Spanish.
Mozilla Data Collective. CC0 1.0.