datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
synthetic-parallel-external
Synthetic Parallel EN↔LG — external
Voice-controlled synthetic parallel speech dataset for Luganda-English
speech-to-speech translation, generated by the Hibiki-Zero fine-tuning pipeline.
Generation
Component
Model
Translation
Sunbird/translate-nllb-3.3b-salt
TTS
Sunbird/orpheus-3b-tts-multilingual
English speakers: salt_eng_0001, salt_eng_0002, salt_eng_0003
Luganda speakers: salt_lug_0001, waxal_lug_0001, waxal_lug_0002, waxal_lug_0003, waxal_lug_0004… See the full description on the dataset page: https://huggingface.co/datasets/yigagilbert/synthetic-parallel-external.YodaLingua-Luxembourgish-Extended
YodaLingua-Luxembourgish-Extended
YodaLingua is a high-quality speech dataset designed for training text-to-speech (TTS) systems, ASR models, and any application requiring clean, well-aligned audio–text pairs.This release contains the Luxembourgish-Extended portion of the multilingual YodaLingua collection.
🧾 Dataset Overview
Property
Value
Total clips
21,687 audio–transcription pairs
Total duration
75 hours
Speakers
2,674 distinct speakers
Audio format… See the full description on the dataset page: https://huggingface.co/datasets/Thomcles/YodaLingua-Luxembourgish-Extended.tts_stage2_extra
Bengali TTS Stage 2 Extra
Final filtered (keep=True) Bengali TTS dataset — extra speakers.
Layout
<speaker>.parquet (single shard, audio <= ~1GB)
<speaker>_001.parquet, _002.parquet, ... (multiple shards, split by audio byte size)
Audio sourcing convention
audio bytes sourced by uuid from bengali-tts-stage1-extra/<speaker>/part-*.parquet
See stat.json for per-speaker row/duration counts.
