robotsmali
afvoices
📘 African Next Voices – Bambara (AfVoices)
The AfVoices dataset is the largest open corpus of spontaneous Bambara speech at its release in late 2025. It contains 423 hours of segmented audio and 612 hours of original raw recordings collected across southern Mali. Speech was recorded in natural, conversational settings and annotated using a semi-automated transcription pipeline combining ASR pre-labels and human corrections. We release all the data processing code on GitHub.… See the full description on the dataset page: https://huggingface.co/datasets/RobotsMali/afvoices.bam-asr-early
All Bambara ASR Dataset
This is the dataset that fueled our early ASR experiments that gave as results the V0 models. It is primarily composed of the Jeli-ASR dataset (available at RobotsMali/jeli-asr), along with the Mali-Pense data curated and published by Aboubacar Ouattara (available at oza75/bambara-tts). Additionally, it includes 1 hour of audio recently collected by the RobotsMali AI4D Lab, featuring children's voices reading some of RobotsMali GAIFE books. This dataset is… See the full description on the dataset page: https://huggingface.co/datasets/RobotsMali/bam-asr-early.jeli-asr
Jeli-ASR Dataset
This repository contains the Jeli-ASR dataset, which is primarily a reviewed version of Aboubacar Ouattara's Bambara-ASR dataset (drawn from jeli-asr and available at oza75/bambara-asr) combined with the best data retained from the former version: jeli-data-manifest. This dataset features improved data quality for automatic speech recognition (ASR) and translation tasks, with variable length Bambara audio samples, Bambara transcriptions and French translations.… See the full description on the dataset page: https://huggingface.co/datasets/RobotsMali/jeli-asr.kunkado
Kunnafonidilaw ka cadeau 🇲🇱
A messy‑real Bambara ASR corpus for developing modern speech models & code‑switch studies
Quick Facts
value
Total duration
161.15 h
Reviewed subset
39.3 h (≈ 25 %)
Total segments
118 925
Languages
Bambara (majority) • French (code‑switch) • misc. Arabic (translit)
LICENSE
CC‑BY‑SA 4.0
kunkado aims to mirror how Malians speak bambara today: fast, informal, and full of French code‑switching. We hope it fuels robust… See the full description on the dataset page: https://huggingface.co/datasets/RobotsMali/kunkado.an-be-kalan-bench
Bambara Educational Speech Dataset
This dataset is a collection of READ Bambara text based on educational children's books from RobotsMali's GAIFE project. It is designed to support the training and benchmarking of Automatic Speech Recognition (ASR) models, with a particular focus on child speech, regional acoustics, and repetitive text structures (inherent to the domain).
The dataset is structured into two separate subsets to support specialized training and evaluation… See the full description on the dataset page: https://huggingface.co/datasets/RobotsMali/an-be-kalan-bench.bayelemabagaThe Bayelemabaga dataset is a collection of 44160 aligned machine translation ready Bambara-French lines,
originating from Corpus Bambara de Reference. The dataset is constitued of text extracted from 231 source files,
varing from periodicals, books, short stories, blog posts, part of the Bible and the Quran.
