CoolFace
Datasetpublic

bookbot/common-voice-23-0-es-ipa-v1

Common Voice 23.0 Spanish — canonical IPA v1 This repository is the approved deterministic v1 publication cut of Spanish Common Voice 23.0 recipe cuts. It contains the accepted final Lhotse cuts only: split rows train 352,941 test 15,857 The data files have exactly three columns: id — the unique final Lhotse cut identifier (for example, common_voice_es_19696062-9131). audio — a Hugging Face Audio feature containing the original encoded source bytes and source… See the full description on the dataset page: https://huggingface.co/datasets/bookbot/common-voice-23-0-es-ipa-v1.

sourceHugging Facecc0-1.0updated 3d agoView on Hugging Face
0likes90downloads
Dataset Card

Common Voice 23.0 Spanish — canonical IPA v1

This repository is the approved deterministic v1 publication cut of Spanish Common Voice 23.0 recipe cuts. It contains the accepted final Lhotse cuts only:

splitrows
train352,941
test15,857

The data files have exactly three columns:

  • —id — the unique final Lhotse cut identifier (for example, common_voice_es_19696062-9131).
  • —audio — a Hugging Face Audio feature containing the original encoded source bytes and source basename.
  • —ipa_transcript — the exact whitespace-separated canonical v1 phone string from supervisions[0].text in the final cut.

The legacy source Parquet field phonemes_ipa (BabyGruut-era output) was not used. The IPA strings here are the v1 canonical recipe supervision, validated against the fixed atomic phone inventory in egs/bookbot_es/ASR/es-tokens.txt.

Source, license, and Mozilla notices

The source is Mozilla Foundation, [Common Voice Scripted Speech 23.0 — Spanish](https://mozilladatacollective.com/datasets/cmj8u3pe600fhnxxbqol7x8bg), obtained through Mozilla Data Collective. The source license is CC0 1.0. The pinned source manifest is common_voice_23_0_es/manifest.tsv, SHA-256 ca5e485fd427ca33802293a228b534e1ebb0a14e6f7187bf04f51dc56702d35c; the pinned source revision is 591a99e9c5c9e308231746c7baa03eadfe670483.

Mozilla’s Common Voice Legal Terms (effective 2025-10-31) state that Common Voice datasets are made available through Mozilla Data Collective under CC0, and ask users not to post, distribute, or mirror Common Voice datasets in whole or in part on other platforms or services. Users of this approved project publication must review and comply with the current Mozilla terms and Privacy Notice. CC0 does not require attribution, but the recommended attribution is: Mozilla Foundation, Common Voice Scripted Speech 23.0 — Spanish, Mozilla Data Collective, CC0 1.0.

This repository is a derived recipe cut, not an official Mozilla release. The source Parquet audio bytes are copied unchanged for every accepted full-recording cut; no resampling, volume, or speed augmentation is materialized. The final cut’s deterministic preparation metadata and file hashes are in `provenance/provenance.json`; the complete one-to-one source join audit is in `provenance/join_manifest.tsv.gz`.

Preparation and transformations

  • —Source rows were joined by the pinned audio.path basename to the final cut supervision/recording identity.
  • —Every accepted cut has start = 0 and covers its complete source recording. Therefore the embedded encoded source bytes are retained losslessly; there were zero trimmed cuts requiring extraction.
  • —id is the final cut ID, not a regenerated source ID.
  • —ipa_transcript is copied exactly from the final canonical supervision and is not reconstructed from orthographic text or the legacy phonemes_ipa field.
  • —The final cut metadata records the training recipe’s optional resample/volume/speed transforms, but those augmentations are intentionally not applied to this source-preserving publication.

The source manifest’s Charsiu raw phonemization is provenance only; downstream v1 normalization and rejection decisions are those recorded in the final cuts. The exact output shard hashes, row counts, cut manifest hashes, source revision, and transformation policy are machine-readable in provenance/provenance.json.

Fixed smoke sample

The deterministic CC0 smoke sample cut is common_voice_es_19696062-9131 in test, sourced from common_voice_es_19696062.mp3. Runtime packages should resample a copied waveform to 16 kHz; this dataset preserves the original encoded source sample rate and bytes.

Citation and attribution

text
Mozilla Foundation. Common Voice Scripted Speech 23.0 — Spanish.
Mozilla Data Collective. CC0 1.0.