datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
kencorpus_sw_culture
KenCorpus Swahili Culture Subset
A filtered subset of Kencorpus/KenCorpus_audio,
containing only rows where language=Swahili and genre=Culture (37 clips).
Audio files are in audio/, indexed by kencorpus_sw_culture.jsonl with path and duration fields,
following the layout of kyutai/DailyTalkContiguous.
sample_sw_culture
Swahili Culture Conversational Dataset (Stereo)
Overview
This dataset contains conversational Swahili audio samples derived from the
Culture genre subset of KenCorpus_audio
(CC-BY-4.0), restructured into diarized, speaker-separated stereo audio
chunks suitable for fine-tuning speech-to-speech dialogue models such as
Moshi / Hibiki.
Processing notebook: the full pipeline (diarization, conversational
chunking, stereo construction, and upload) is available here:… See the full description on the dataset page: https://huggingface.co/datasets/rlabz/sample_sw_culture.
