datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Bambara-Speech-Translation-Data
AfVoices-Translated (Bambara-English)
This is a Bambara speech translation dataset, which is built on the African Next Voices (AfVoices) Bambara ASR corpus. It provides English translations for the human-corrected subset of the original collection, creating a parallel corpus for Bambara-English machine translation and speech-to-text tasks.
Methodology
We machine-translated the human-validated transcriptions from AfVoices using the Oolel-translator repository.
Inference… See the full description on the dataset page: https://huggingface.co/datasets/soynade-research/Bambara-Speech-Translation-Data.merged-bambara-dioula-datasetmerged-bambara-dioula-datasetBambara_AudioSynthetique_42K_V3
Description
Ce corpus comprend 42 000 entrées audio synthétiques en langue Bambara (bm), totalisant environ 44,4 heures d'enregistrement. Cette version 3 a été convertie au format Parquet pour optimiser les performances de lecture et garantir une compatibilité totale avec le Dataset Viewer de Hugging Face.
Origine et Traitement des Données Textuelles
Le corpus de texte a été constitué par l'agrégation de plusieurs sources linguistiques afin de garantir un volume suffisant… See the full description on the dataset page: https://huggingface.co/datasets/OumarDicko/Bambara_AudioSynthetique_42K_V3.bambara-tts-waxal
bambara-tts-waxal
Bambara studio speech from the WAXAL corpus — 1,926 recordings, 16 hours, 8 speakers,
44.1 kHz mono.
Load
from datasets import load_dataset
ds = load_dataset("djelia/bambara-tts-waxal", "google_waxal", split="train")
Splits: train, validation, test.
Fields
Field
Description
audio
44.1 kHz mono
text
Transcript
speaker_id
Speaker identifier (8 distinct)
gender
Speaker gender
locale
Locale code
id
Record… See the full description on the dataset page: https://huggingface.co/datasets/djelia/bambara-tts-waxal.bambara-asr-benchmark
Bambara ASR Benchmark
The first standardized evaluation set for Automatic Speech Recognition in Bambara (Bamanankan). One hour of studio-quality constitutional text, transcribed and validated by linguists from Mali's Direction Nationale de l'Éducation Non Formelle et des Langues Nationales (DNENF-LN).
This benchmark accompanies the paper "Where Are We at with Automatic Speech Recognition for the Bambara Language?" and the public leaderboard at MALIBA-AI/bambara-asr-leaderboard.… See the full description on the dataset page: https://huggingface.co/datasets/MALIBA-AI/bambara-asr-benchmark.Bambara_AudioSynthetique_V1_LEGACY
⚠️ [OBSOLETE / INCOMPLET] Bambara Audio Dataset - Version Archivée
Attention : Cette version est obsolète et ne contient qu'une fraction des données disponibles.
La Version 3 de ce projet est désormais la référence. Elle contient l'intégralité du corpus (42 000 fichiers contre seulement une partie ici) et a été optimisée techniquement.
👉 Accéder au Corpus Complet V3 (42 000 audios - 44.4h)
Pourquoi passer absolument à la V3 ?
Volume : Accès à… See the full description on the dataset page: https://huggingface.co/datasets/OumarDicko/Bambara_AudioSynthetique_V1_LEGACY.bambara-speech-kis-clean-split
Bambara Speech Dataset — Clean & Split
Dataset de reconnaissance vocale en bambara, nettoyé et splitté pour le fine-tuning de modèles ASR (ex: Whisper).
La source principale des données brutes est RobotsMali/bam-asr-early, auquel un remerciement chaleureux lui est attribué mais aussi à d'autres personnes référencées ci-dessous dans la section citation.
Statistiques
Total : 35 342 échantillons
Train : 24 738
Validation : 3 535
Test : 7 069
Durée moyenne : 3.23s… See the full description on the dataset page: https://huggingface.co/datasets/kalilouisangare/bambara-speech-kis-clean-split.test-bambara-ttsbambara-asrbambara-audio
Djelia Bambara Audio Dataset
Dataset Description
The Djelia Bambara Audio Dataset is a comprehensive resource aimed at supporting research and development in Bambara language processing. This dataset consists of audio extracted from YouTube videos, denoised and diarized to ensure high-quality segments. Additionally, it features a semi-annotated subset with transcriptions generated using the Djelia Whisper v1 model.
Features
Audio: High-quality audio clips… See the full description on the dataset page: https://huggingface.co/datasets/djelia/bambara-audio.Bambara_AudioSynthetique_V2_LEGACY
⚠️ [OBSOLETE / INCOMPLET] Bambara Audio Dataset - Version Archivée
Attention : Cette version est obsolète et ne contient qu'une fraction des données disponibles.
La Version 3 de ce projet est désormais la référence. Elle contient l'intégralité du corpus (42 000 fichiers contre seulement une partie ici) et a été optimisée techniquement.
👉 Accéder au Corpus Complet V3 (42 000 audios - 44.4h)
Pourquoi passer absolument à la V3 ?
Volume : Accès à… See the full description on the dataset page: https://huggingface.co/datasets/OumarDicko/Bambara_AudioSynthetique_V2_LEGACY.bambara-tts
Overview
Project
This dataset is part of a larger initiative aimed at empowering Bambara speakers to access global knowledge without language barriers.
Our goal is to eliminate the need for Bambara speakers to learn a secondary language before they can acquire new information or skills.
By providing a robust dataset for Text-to-Speech (TTS) applications, we aim to support the creation of tools for bambara language, thus democratizing access to knowledge.… See the full description on the dataset page: https://huggingface.co/datasets/oza75/bambara-tts.bambara-asr-v2
bambara-asr-v2
Multi-corpus Bambara speech — 185,708 examples, ~366 hours, 53.7 GB of Parquet. Seven configs,
each a train / dev / test triple of 16 kHz audio paired with a text target. Every config
draws on a single upstream corpus, so you can mix and weight them yourself.
Access is gated with manual approval — request it on the dataset page and authenticate
(hf auth login or HF_TOKEN) before loading.
Load
from datasets import load_dataset
jeli =… See the full description on the dataset page: https://huggingface.co/datasets/djelia/bambara-asr-v2.bambara-asr-dataset-y
bambara-asr-dataset-y
Bambara speech paired with the French source line it renders. 58,447 rows, 36.66 hours,
48.77 GB of Parquet. The only text column is fr — there is no Bambara text here.
Load
The config is default and the splits are not_combined and combined — there is no train
split, so a bare load_dataset returns a DatasetDict keyed by those two names.
from datasets import load_dataset
short = load_dataset("djelia/bambara-asr-dataset-y"… See the full description on the dataset page: https://huggingface.co/datasets/djelia/bambara-asr-dataset-y.bible-bambara-audio
Bible Bambara Audio Dataset
Overview
This dataset contains audio recordings of Bible passages in Bambara language along with their transcriptions. The dataset consists of approximately 42.7 hours of audio content, making it a valuable resource for speech processing tasks in Bambara language.
Project
This dataset is part of a larger initiative to preserve and digitize audio & texts in Bambara language, making them accessible in both text and audio… See the full description on the dataset page: https://huggingface.co/datasets/oza75/bible-bambara-audio.open-bambara-asr-datasetbambara-speech-recognition-benchmarkbambara-synthetic-audio
bambara-synthetic-audio
102,310 utterances of synthetic Bambara speech, generated by a text-to-speech model over
Bambara text. Every clip is machine-generated; no human voice is recorded here.
Load
from datasets import load_dataset
# Enhanced set, with per-clip quality scores
semi = load_dataset("djelia/bambara-synthetic-audio", "semi-clean", split="train")
# Larger generation set, no quality scores
v2 = load_dataset("djelia/bambara-synthetic-audio", "tts_v2"… See the full description on the dataset page: https://huggingface.co/datasets/djelia/bambara-synthetic-audio.maliba_bambarabambara-audio-b
bambara-audio-b
Bambara speech derived from scripture recordings, published in four processing stages: raw
segments, a length-filtered version, a speaker-diarized long-form cut, and a CTC
forced-alignment cut. 30.55 GB of Parquet.
Access is gated with manual approval — request it on the dataset page and authenticate
(hf auth login or HF_TOKEN) before loading.
Load
from datasets import load_dataset
ds = load_dataset("djelia/bambara-audio-b", "short-filtered"… See the full description on the dataset page: https://huggingface.co/datasets/djelia/bambara-audio-b.bambara-audio-y
bambara-audio-y
Bambara speech paired with the French source line it renders, a written Bambara translation of
that line, and a machine transcription of the audio. 58,447 rows, 36.66 hours, 48.77 GB of
Parquet.
Load
The config is default and the splits are not_combined and combined — there is no train
split, so a bare load_dataset returns a DatasetDict keyed by those two names.
from datasets import load_dataset
short = load_dataset("djelia/bambara-audio-y"… See the full description on the dataset page: https://huggingface.co/datasets/djelia/bambara-audio-y.bambara-asr
bambara-asr
Multi-task Bambara speech: transcription, speech-to-text translation into French and English,
and a multilingual training mix. 16 kHz audio in Parquet across nine configs.
Access is gated with manual approval — request it on the dataset page and authenticate
(hf auth login or HF_TOKEN) before loading.
Load
from datasets import load_dataset
ds = load_dataset("djelia/bambara-asr", "bm-to-bm", split="train")
Every config has train and test splits.… See the full description on the dataset page: https://huggingface.co/datasets/djelia/bambara-asr.bambara-numbersbambara-audiobambara-asr-evaluation
bambara-asr-evaluation
A Bambara ASR benchmark: 1,295 utterances, 2.04 hours of 16 kHz audio with reference
transcripts. Monolingual Bambara transcription — audio in, transcript out, WER out.
Load
from datasets import load_dataset
ds = load_dataset("djelia/bambara-asr-evaluation", split="test")
print(ds[0]["text"], ds[0]["source_dataset"])
One config and one split, so no config argument is needed.
Config
Split
Rows
Audio
default
test
1,295
2.043 h… See the full description on the dataset page: https://huggingface.co/datasets/djelia/bambara-asr-evaluation.bambara-ext-eval-dataBambara-Keyword-Spotting-AugBambara-Keyword-Spotting
