CoolFace
Datasetpublic

yigagilbert/synthetic-parallel-salt

Synthetic Parallel EN↔LG — salt Voice-controlled synthetic parallel speech dataset for Luganda-English speech-to-speech translation, generated by the Hibiki-Zero fine-tuning pipeline. Generation Component Model Translation Sunbird/translate-nllb-3.3b-salt TTS Sunbird/orpheus-3b-tts-multilingual English speakers: salt_eng_0001, salt_eng_0002, salt_eng_0003 Luganda speakers: salt_lug_0001, waxal_lug_0001, waxal_lug_0002, waxal_lug_0003… See the full description on the dataset page: https://huggingface.co/datasets/yigagilbert/synthetic-parallel-salt.

sourceHugging Facecc-by-4.0updated 4mo agoView on Hugging Face
0likes29downloads
Dataset Card

Synthetic Parallel EN↔LG — salt

Voice-controlled synthetic parallel speech dataset for Luganda-English speech-to-speech translation, generated by the Hibiki-Zero fine-tuning pipeline.

Generation

ComponentModel
TranslationSunbird/translate-nllb-3.3b-salt
TTSSunbird/orpheus-3b-tts-multilingual

English speakers: salt_eng_0001, salt_eng_0002, salt_eng_0003

Luganda speakers: salt_lug_0001, waxal_lug_0001, waxal_lug_0002, waxal_lug_0003, waxal_lug_0004, waxal_lug_0005, waxal_lug_0007

Speaker sampling: N ∈ {1 (70%), 2 (20%), 3 (10%)} draws per text pair.

Source datasets (transcripts only — source audio discarded)

  • Sunbird/tts (subset: eng, lang: eng)
  • Sunbird/tts (subset: lug, lang: lug)
  • Sunbird/speech (subset: eng_salt, lang: eng)
  • Sunbird/speech (subset: lug_salt, lang: lug)

Post-processing

  • Silero / energy VAD silence trim on every clip
  • Duration filter: Luganda 1.5–20.0 s, English 1.0–20.0 s
  • Duration ratio filter: 0.3–4.0
  • Speech-activity filter: ≥ 0.35
  • Sample rate: 24000 Hz (mono, float32 → PCM-16 WAV)

Schema

ColumnTypeDescription
audio_lugAudioLuganda synthesized clip
audio_engAudioEnglish synthesized clip
text_lugstringLuganda transcript
text_engstringEnglish transcript
speaker_lugstringOrpheus speaker ID used for Luganda
speaker_engstringOrpheus speaker ID used for English
src_dur_sfloatLuganda clip duration after VAD trim (s)
tgt_dur_sfloatEnglish clip duration after VAD trim (s)
dur_ratiofloattgt / src duration ratio
src_speech_ratiofloatVAD speech fraction — Luganda
tgt_speech_ratiofloatVAD speech fraction — English
source_datasetstringOriginating HF dataset repo
source_subsetstringOriginating subset / config name

Failure manifest

Rows that failed TTS generation (after 3 retries) or post-processing are logged to: https://huggingface.co/datasets/yigagilbert/synthetic-parallel-salt/resolve/main/failure_manifest.jsonl