fongbe
Datasets
All datasets matching “fongbe”fongbe-speech-zenodo
Fongbe Speech Dataset (Complete & Tone-Preserved)
Dataset Summary
This dataset is a unified, high-quality collection of Fongbe speech data, specifically curated to preserve the linguistic integrity of this tonal language. It acts as a complete, unsegmented, and tone-accurate assembly of the Fongbe Continuous Speech Recognition corpora, merging:
The foundational ALFFA Project data (Train/Test splits, 2016).
The expanded Zenodo release (Validation split, 2022).… See the full description on the dataset page: https://huggingface.co/datasets/Professor/fongbe-speech-zenodo.fongbe-hausa-asr-dataset
Fongbe-Hausa ASR Dataset (Semi-Supervised)
This dataset provides ~6,770 audio-transcription pairs for Fongbe (fon) and Hausa (hau). It was created using a semi-supervised pipeline to convert long-form video content into a training-ready format for Automatic Speech Recognition (ASR).
Dataset Details
Total Examples: 6,770
Audio Format: WAV (16kHz, Mono)
Languages: Fongbe (Benin), Hausa (Nigeria/West Africa)
Annotation: Semi-supervised (Machine-generated labels)
License:… See the full description on the dataset page: https://huggingface.co/datasets/Professor/fongbe-hausa-asr-dataset.fongbe-speechmece-fongbe-corpus
MƐCE — Corpus d'instruction Fongbe (Tune_Pigier)
Corpus d'instruction / conversation centré sur le Fongbe (fon), utilisé pour
fine-tuner l'assistant vocal MƐCE (mémoire de fin d'études, École PIGIER Bénin).
Contenu
Format : ChatML — chaque ligne JSON est un objet {"messages": [...]} avec des
rôles system / user / assistant.
Taille : ~140 000 exemples — train.jsonl (126 725) + eval.jsonl (14 089).
Langues : fon (principal, avec tons/diacritiques), fr, en.… See the full description on the dataset page: https://huggingface.co/datasets/CapitainVigs/mece-fongbe-corpus.fongbe-whisperfongbe-speech-dataset-female
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: Fréjus LALEYE
Shared by [optional]: Fréjus LALEYE
Language(s) (NLP): Fongbe
License: [More Information Needed]
Dataset Sources [optional]
Repository: https://github.com/laleye/pyFongbe
Paper [optional]:… See the full description on the dataset page: https://huggingface.co/datasets/beethogedeon/fongbe-speech-dataset-female.
