groxaxo/google-latam-spanish-uniform-vad
Google LATAM Spanish — Uniform VAD and 0.5 s Edges This public derivative contains 15,016 Latin American Spanish utterances from the Google crowdsourced TTS datasets packaged by ylacombe. Female and male audio were freshly exported from the same pinned upstream revisions and passed through exactly the same processing pipeline. Configurations Configuration Train Validation Total Hours including edge padding argentina-female 3,542 379 3,921 4.181… See the full description on the dataset page: https://huggingface.co/datasets/groxaxo/google-latam-spanish-uniform-vad.
Google LATAM Spanish — Uniform VAD and 0.5 s Edges
This public derivative contains 15,016 Latin American Spanish utterances from the Google crowdsourced TTS datasets packaged by ylacombe. Female and male audio were freshly exported from the same pinned upstream revisions and passed through exactly the same processing pipeline.
Configurations
Each configuration is independently speaker-disjoint between train and validation. Country and gender are never mixed within a configuration.
Uniform processing
Every source clip received the same treatment, without country- or gender-specific branches:
- Export the original upstream audio as 48 kHz, mono, 16-bit PCM.
- Detect the first and last speech regions with Silero VAD pinned to commit
60b7ffa243625ebdc1070275a29f18c87843786a. - Use probability threshold
0.5, minimum speech100 ms, minimum silence100 ms, and50 msprotective speech context. - Retain everything from the first detected region through the last detected region, including internal pauses.
- Remove only contiguous literal-zero samples at the retained edges.
- Add exactly 0.500 seconds (24,000 frames) of digital zero at both the start and end.
No fixed dBFS amplitude threshold was used. All 15,016 clips produced a VAD region and passed a full post-processing scan for format, exact edge padding, and manifest duration consistency.
Sources
The original recordings, transcripts, and speaker labels come from those datasets. Thanks to ylacombe for publishing and maintaining them.
Fields
audio: embedded WAV audioid: stable country/gender/speaker/source-index identifiertext: exact upstream transcriptspeaker_idandspeaker_key: speaker identifierscountryandgender: fixed values for the selected configurationsource_dataset,source_revision, andsource_index: provenanceduration: processed duration including both 0.5-second edgessample_rate: 48,000 Hzaudio_sha256: checksum of the processed WAVprocessing: pipeline identifier
License
CC BY-SA 4.0, matching the upstream dataset declarations. Preserve upstream attribution when redistributing or adapting this dataset.
