datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fongbe-speech-zenodo
Fongbe Speech Dataset (Complete & Tone-Preserved)
Dataset Summary
This dataset is a unified, high-quality collection of Fongbe speech data, specifically curated to preserve the linguistic integrity of this tonal language. It acts as a complete, unsegmented, and tone-accurate assembly of the Fongbe Continuous Speech Recognition corpora, merging:
The foundational ALFFA Project data (Train/Test splits, 2016).
The expanded Zenodo release (Validation split, 2022).… See the full description on the dataset page: https://huggingface.co/datasets/Professor/fongbe-speech-zenodo.fongbe-hausa-asr-dataset
Fongbe-Hausa ASR Dataset (Semi-Supervised)
This dataset provides ~6,770 audio-transcription pairs for Fongbe (fon) and Hausa (hau). It was created using a semi-supervised pipeline to convert long-form video content into a training-ready format for Automatic Speech Recognition (ASR).
Dataset Details
Total Examples: 6,770
Audio Format: WAV (16kHz, Mono)
Languages: Fongbe (Benin), Hausa (Nigeria/West Africa)
Annotation: Semi-supervised (Machine-generated labels)
License:… See the full description on the dataset page: https://huggingface.co/datasets/Professor/fongbe-hausa-asr-dataset.fongbe-speechfongbe-whisperfongbe-speech-dataset-female
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: Fréjus LALEYE
Shared by [optional]: Fréjus LALEYE
Language(s) (NLP): Fongbe
License: [More Information Needed]
Dataset Sources [optional]
Repository: https://github.com/laleye/pyFongbe
Paper [optional]:… See the full description on the dataset page: https://huggingface.co/datasets/beethogedeon/fongbe-speech-dataset-female.fongbefongbe_datasetfongbe-speech-dataset-male
