datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
afvoices
📘 African Next Voices – Bambara (AfVoices)
The AfVoices dataset is the largest open corpus of spontaneous Bambara speech at its release in late 2025. It contains 423 hours of segmented audio and 612 hours of original raw recordings collected across southern Mali. Speech was recorded in natural, conversational settings and annotated using a semi-automated transcription pipeline combining ASR pre-labels and human corrections. We release all the data processing code on GitHub.… See the full description on the dataset page: https://huggingface.co/datasets/RobotsMali/afvoices.bam-asr-early
All Bambara ASR Dataset
This is the dataset that fueled our early ASR experiments that gave as results the V0 models. It is primarily composed of the Jeli-ASR dataset (available at RobotsMali/jeli-asr), along with the Mali-Pense data curated and published by Aboubacar Ouattara (available at oza75/bambara-tts). Additionally, it includes 1 hour of audio recently collected by the RobotsMali AI4D Lab, featuring children's voices reading some of RobotsMali GAIFE books. This dataset is… See the full description on the dataset page: https://huggingface.co/datasets/RobotsMali/bam-asr-early.jeli-asr
Jeli-ASR Dataset
This repository contains the Jeli-ASR dataset, which is primarily a reviewed version of Aboubacar Ouattara's Bambara-ASR dataset (drawn from jeli-asr and available at oza75/bambara-asr) combined with the best data retained from the former version: jeli-data-manifest. This dataset features improved data quality for automatic speech recognition (ASR) and translation tasks, with variable length Bambara audio samples, Bambara transcriptions and French translations.… See the full description on the dataset page: https://huggingface.co/datasets/RobotsMali/jeli-asr.kunkado
Kunnafonidilaw ka cadeau 🇲🇱
A messy‑real Bambara ASR corpus for developing modern speech models & code‑switch studies
Quick Facts
value
Total duration
161.15 h
Reviewed subset
39.3 h (≈ 25 %)
Total segments
118 925
Languages
Bambara (majority) • French (code‑switch) • misc. Arabic (translit)
LICENSE
CC‑BY‑SA 4.0
kunkado aims to mirror how Malians speak bambara today: fast, informal, and full of French code‑switching. We hope it fuels robust… See the full description on the dataset page: https://huggingface.co/datasets/RobotsMali/kunkado.an-be-kalan-bench
Bambara Educational Speech Dataset
This dataset is a collection of READ Bambara text based on educational children's books from RobotsMali's GAIFE project. It is designed to support the training and benchmarking of Automatic Speech Recognition (ASR) models, with a particular focus on child speech, regional acoustics, and repetitive text structures (inherent to the domain).
The dataset is structured into two separate subsets to support specialized training and evaluation… See the full description on the dataset page: https://huggingface.co/datasets/RobotsMali/an-be-kalan-bench.afvoices-notag
AfVoices Top-20 Speakers without Tags
RobotsMali/afvoices-notag is a small experimental TTS-oriented selection derived from RobotsMali/afvoices, the African Next Voices Bambara speech corpus. It contains the 20 participants with the highest utterance counts and excludes transcripts containing semantic/acoustic annotation tags.
This is the dataset used for RobotsMali's first Bambara VITS experiments. It is not a high-quality studio TTS corpus: the source is spontaneous speech… See the full description on the dataset page: https://huggingface.co/datasets/RobotsMali/afvoices-notag.human-robot-conversation-russian
Human-Robot Dataset
The dataset comprises 660+ hours of Russian speech across 20,000+ audio files featuring human-robot interactions between AI and humans. It is designed for research in conversational agents, focusing on various speech recognition methods, primarily aimed at advancing language models and machine learning applications.
By utilizing this dataset, researchers and developers can advance their understanding and capabilities in speech recognition, natural language… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/human-robot-conversation-russian.human-robot-conversation-korean
Human-Robot Dataset
The dataset comprises 660+ hours of audio recordings across 20,000+ files for human-robot interactions in the Korean language. It captures authentic dialogues between humans and artificial conversational agents, specifically designed for training language models and advancing speech recognition systems.
By utilizing this dataset, researchers and developers can advance their understanding and capabilities in robotic systems and conversational AI technologies.… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/human-robot-conversation-korean.Bam_ASR_Eval_500
Bam_ASR_Eval_500 Dataset
Dataset Description
Bam_ASR_Eval_500 is a curated evaluation dataset for Automatic Speech Recognition (ASR) models in Bambara (Bamanakan), a major language spoken in Mali and West Africa. This dataset comprises 500 audio recordings totaling approximately 36.69 minutes of annotated speech, designed specifically for benchmarking ASR systems. It focuses on real-world challenges in low-resource languages like Bambara, including spontaneous speech… See the full description on the dataset page: https://huggingface.co/datasets/RobotsMali/Bam_ASR_Eval_500.nyana-eval
Nyana-Eval Dataset
Dataset Description
Nyana-Eval is a compact, stratified evaluation subset for benchmarking Automatic Speech Recognition (ASR) models in Bambara. It consists of 45 audio recordings totaling approximately 3.03 minutes, carefully selected to represent real-world linguistic and acoustic challenges in low-resource Bambara speech. This dataset is derived from the larger RobotsMali/Bam_ASR_Eval_500 corpus and is optimized for quick, reproducible human… See the full description on the dataset page: https://huggingface.co/datasets/RobotsMali/nyana-eval.human-robot-conversation-korean
Human-Robot Conversation Dataset (Korean) - 660+ Hours
Dataset (Korean) contains 660+ hours of audio featuring dialogues between AI and a human in German across 20,000 recordings. The dataset supports conversational AI, speech recognition, and human-robot interaction research, with short M4A audio files (up to 2 minutes) and structured metadata for model training. - Get the data
Dataset characteristics:
Characteristic
Data
Description
Audio of dialogues between AI… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/human-robot-conversation-korean.human-robot-conversation-german
Human-Robot Dataset
The dataset comprises 660+ hours of audio recordings across 20,000+ files for human-robot interactions in the German language. It captures authentic dialogues between humans and artificial conversational agents, specifically designed for training language models and advancing speech recognition systems.
By utilizing this dataset, researchers and developers can advance their understanding and capabilities in speech recognition, natural language processing, and… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/human-robot-conversation-german.human-robot-conversation-russian
Human-Robot Conversation Dataset (Russian) - 660+ Hours
Dataset (Russian) contains 660+ hours of audio featuring dialogues between AI and a human in German across 20,000 recordings. The dataset supports conversational AI, speech recognition, and human-robot interaction research, with short M4A audio files (up to 2 minutes) and structured metadata for model training. - Get the data
Dataset characteristics:
Characteristic
Data
Description
Audio of dialogues between… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/human-robot-conversation-russian.human-robot-conversation-english
Human-Robot Dataset
The dataset comprises 660+ hours of English speech across 20,000+ audio files featuring human-robot interactions between AI and humans. It is designed for research in conversational agents, focusing on various speech recognition methods, primarily aimed at advancing language models and machine learning applications.
By utilizing this dataset, researchers and developers can advance their understanding and capabilities in robotic systems and conversational AI… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/human-robot-conversation-english.human-robot-conversation-english
Human-Robot Conversation Dataset (English) - 660+ Hours
Dataset (English) contains 660+ hours of audio featuring dialogues between AI and a human in English across 20,000 recordings. The dataset supports conversational AI, speech recognition, and human-robot interaction research, with short M4A audio files (up to 2 minutes) and structured metadata for model training. - Get the data
Dataset characteristics:
Characteristic
Data
Description
Audio of dialogues between… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/human-robot-conversation-english.transcription-scorer
Transcription Scorer Dataset
The Transcription Scorer dataset was created to support research in reference-free evaluation of Automatic Speech Recognition (ASR) systems using human feedback. Unlike traditional evaluation metrics such as WER and its derivatives, this dataset reflects judgments of ASR outputs by human raters across multiple criteria, simulating the way a teacher grades students.
⚙️ What’s Inside
This dataset contains 1200 audio samples (from diverse sources… See the full description on the dataset page: https://huggingface.co/datasets/RobotsMali/transcription-scorer.human-robot-conversation-german
Human-Robot Conversation Dataset (German) - 660+ Hours
Dataset (German) contains 660+ hours of audio featuring dialogues between AI and a human in German across 20,000 recordings. The dataset supports conversational AI, speech recognition, and human-robot interaction research, with short M4A audio files (up to 2 minutes) and structured metadata for model training. - Get the data
Dataset characteristics:
Characteristic
Data
Description
Audio of dialogues between AI… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/human-robot-conversation-german.
