datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
afvoices
📘 African Next Voices – Bambara (AfVoices)
The AfVoices dataset is the largest open corpus of spontaneous Bambara speech at its release in late 2025. It contains 423 hours of segmented audio and 612 hours of original raw recordings collected across southern Mali. Speech was recorded in natural, conversational settings and annotated using a semi-automated transcription pipeline combining ASR pre-labels and human corrections. We release all the data processing code on GitHub.… See the full description on the dataset page: https://huggingface.co/datasets/RobotsMali/afvoices.bam-asr-early
All Bambara ASR Dataset
This is the dataset that fueled our early ASR experiments that gave as results the V0 models. It is primarily composed of the Jeli-ASR dataset (available at RobotsMali/jeli-asr), along with the Mali-Pense data curated and published by Aboubacar Ouattara (available at oza75/bambara-tts). Additionally, it includes 1 hour of audio recently collected by the RobotsMali AI4D Lab, featuring children's voices reading some of RobotsMali GAIFE books. This dataset is… See the full description on the dataset page: https://huggingface.co/datasets/RobotsMali/bam-asr-early.jeli-asr
Jeli-ASR Dataset
This repository contains the Jeli-ASR dataset, which is primarily a reviewed version of Aboubacar Ouattara's Bambara-ASR dataset (drawn from jeli-asr and available at oza75/bambara-asr) combined with the best data retained from the former version: jeli-data-manifest. This dataset features improved data quality for automatic speech recognition (ASR) and translation tasks, with variable length Bambara audio samples, Bambara transcriptions and French translations.… See the full description on the dataset page: https://huggingface.co/datasets/RobotsMali/jeli-asr.kunkado
Kunnafonidilaw ka cadeau 🇲🇱
A messy‑real Bambara ASR corpus for developing modern speech models & code‑switch studies
Quick Facts
value
Total duration
161.15 h
Reviewed subset
39.3 h (≈ 25 %)
Total segments
118 925
Languages
Bambara (majority) • French (code‑switch) • misc. Arabic (translit)
LICENSE
CC‑BY‑SA 4.0
kunkado aims to mirror how Malians speak bambara today: fast, informal, and full of French code‑switching. We hope it fuels robust… See the full description on the dataset page: https://huggingface.co/datasets/RobotsMali/kunkado.an-be-kalan-bench
Bambara Educational Speech Dataset
This dataset is a collection of READ Bambara text based on educational children's books from RobotsMali's GAIFE project. It is designed to support the training and benchmarking of Automatic Speech Recognition (ASR) models, with a particular focus on child speech, regional acoustics, and repetitive text structures (inherent to the domain).
The dataset is structured into two separate subsets to support specialized training and evaluation… See the full description on the dataset page: https://huggingface.co/datasets/RobotsMali/an-be-kalan-bench.bayelemabagaThe Bayelemabaga dataset is a collection of 44160 aligned machine translation ready Bambara-French lines,
originating from Corpus Bambara de Reference. The dataset is constitued of text extracted from 231 source files,
varing from periodicals, books, short stories, blog posts, part of the Bible and the Quran.finBamSpeech
FinBamSpeech
FinBamSpeech is a small, higher-quality, single-speaker corpus of 800 read Bambara sentences about finance, banking, and financial technology. RobotsMali created it to test domain adaptation of its first Bambara VITS checkpoints.
The dataset's narrow domain and single voice make it useful for controlled exploratory fine-tuning, but 800 utterances are not enough to claim broad language, speaker, or topic coverage.
Quick facts
Item
Value… See the full description on the dataset page: https://huggingface.co/datasets/RobotsMali/finBamSpeech.afvoices-notag
AfVoices Top-20 Speakers without Tags
RobotsMali/afvoices-notag is a small experimental TTS-oriented selection derived from RobotsMali/afvoices, the African Next Voices Bambara speech corpus. It contains the 20 participants with the highest utterance counts and excludes transcripts containing semantic/acoustic annotation tags.
This is the dataset used for RobotsMali's first Bambara VITS experiments. It is not a high-quality studio TTS corpus: the source is spontaneous speech… See the full description on the dataset page: https://huggingface.co/datasets/RobotsMali/afvoices-notag.Bam_ASR_Eval_500
Bam_ASR_Eval_500 Dataset
Dataset Description
Bam_ASR_Eval_500 is a curated evaluation dataset for Automatic Speech Recognition (ASR) models in Bambara (Bamanakan), a major language spoken in Mali and West Africa. This dataset comprises 500 audio recordings totaling approximately 36.69 minutes of annotated speech, designed specifically for benchmarking ASR systems. It focuses on real-world challenges in low-resource languages like Bambara, including spontaneous speech… See the full description on the dataset page: https://huggingface.co/datasets/RobotsMali/Bam_ASR_Eval_500.nyana-eval
Nyana-Eval Dataset
Dataset Description
Nyana-Eval is a compact, stratified evaluation subset for benchmarking Automatic Speech Recognition (ASR) models in Bambara. It consists of 45 audio recordings totaling approximately 3.03 minutes, carefully selected to represent real-world linguistic and acoustic challenges in low-resource Bambara speech. This dataset is derived from the larger RobotsMali/Bam_ASR_Eval_500 corpus and is optimized for quick, reproducible human… See the full description on the dataset page: https://huggingface.co/datasets/RobotsMali/nyana-eval.lau-eval
LAU eval dataset
This dataset was created while evaluating and comparing the models trained with Listen Attend Understand regularization and our E2E-ST model.
The audio is from jeli-asr test set; the regularization loss weight lambda in the paper is represented by the character "k" in the fields of this dataset, each field represent a model with a specific decoding strategy (CTC or TDT)
Citation
@misc{diarra2026listenattendunderstandregularization… See the full description on the dataset page: https://huggingface.co/datasets/RobotsMali/lau-eval.transcription-scorer
Transcription Scorer Dataset
The Transcription Scorer dataset was created to support research in reference-free evaluation of Automatic Speech Recognition (ASR) systems using human feedback. Unlike traditional evaluation metrics such as WER and its derivatives, this dataset reflects judgments of ASR outputs by human raters across multiple criteria, simulating the way a teacher grades students.
⚙️ What’s Inside
This dataset contains 1200 audio samples (from diverse sources… See the full description on the dataset page: https://huggingface.co/datasets/RobotsMali/transcription-scorer.
