datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Bambara-Speech-Translation-Data
AfVoices-Translated (Bambara-English)
This is a Bambara speech translation dataset, which is built on the African Next Voices (AfVoices) Bambara ASR corpus. It provides English translations for the human-corrected subset of the original collection, creating a parallel corpus for Bambara-English machine translation and speech-to-text tasks.
Methodology
We machine-translated the human-validated transcriptions from AfVoices using the Oolel-translator repository.
Inference… See the full description on the dataset page: https://huggingface.co/datasets/soynade-research/Bambara-Speech-Translation-Data.bambara-whisper-featuresmerged-bambara-dioula-datasetmerged-bambara-dioula-datasetBambara_AudioSynthetique_42K_V3
Description
Ce corpus comprend 42 000 entrées audio synthétiques en langue Bambara (bm), totalisant environ 44,4 heures d'enregistrement. Cette version 3 a été convertie au format Parquet pour optimiser les performances de lecture et garantir une compatibilité totale avec le Dataset Viewer de Hugging Face.
Origine et Traitement des Données Textuelles
Le corpus de texte a été constitué par l'agrégation de plusieurs sources linguistiques afin de garantir un volume suffisant… See the full description on the dataset page: https://huggingface.co/datasets/OumarDicko/Bambara_AudioSynthetique_42K_V3.Bambara_Translation_Corpus_FR-BM
🌍 Bambara Translation Corpus: Direct & Instruction-Tuned (FR-BM)
🚀 Overview & Vision
The Bambara Translation Corpus is a comprehensive bilingual dataset designed to bridge French and Bamanankan across two distinct paradigms: Direct Neural Machine Translation (NMT) and Prompt-Based Instruction Tuning.
This dual-mode architecture caters to both traditional seq2seq translation models and modern instruction-tuned Large Language Models (LLMs).
📊 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Makan09/Bambara_Translation_Corpus_FR-BM.Bambara-dataset_conversation
license: apache-2.0
language:
- fr
- bm
tags:
- bambara
- bamanankan
- instruction-tuning
- llm-alignment
- african-languages
- low-resource-nlp
- conversational-ai
task_categories:
- text-generation
- conditional-text-generation
size_categories:
- 10K<n<50K
pretty_name: Bambara Instruction Tuning Corpus (FR-BM)
🌍 Bambara Instruction Tuning Corpus (FR-BM)
🚀 Overview & Vision
Welcome to the Bambara Instruction Tuning Corpus… See the full description on the dataset page: https://huggingface.co/datasets/Makan09/Bambara-dataset_conversation.bambara-tts-waxal
bambara-tts-waxal
Bambara studio speech from the WAXAL corpus — 1,926 recordings, 16 hours, 8 speakers,
44.1 kHz mono.
Load
from datasets import load_dataset
ds = load_dataset("djelia/bambara-tts-waxal", "google_waxal", split="train")
Splits: train, validation, test.
Fields
Field
Description
audio
44.1 kHz mono
text
Transcript
speaker_id
Speaker identifier (8 distinct)
gender
Speaker gender
locale
Locale code
id
Record… See the full description on the dataset page: https://huggingface.co/datasets/djelia/bambara-tts-waxal.Bambara_asr_test_wercalculationBambara_texts_raws_corpus
🌍 Bambara Massive Raw Text Corpus (1.7M+ Lines)
🚀 Overview & Vision
Welcome to the Bambara Massive Raw Text Corpus—a monumental milestone for African language technology. Featuring over 1.7 million lines of raw Bamanankan text, this repository represents an unprecedented scale of unstructured linguistic data for a low-resource West African language.
Pre-training foundational models from scratch or performing Continued Pre-Training (CPT) on existing open-source… See the full description on the dataset page: https://huggingface.co/datasets/Makan09/Bambara_texts_raws_corpus.bambara-asr-benchmark
Bambara ASR Benchmark
The first standardized evaluation set for Automatic Speech Recognition in Bambara (Bamanankan). One hour of studio-quality constitutional text, transcribed and validated by linguists from Mali's Direction Nationale de l'Éducation Non Formelle et des Langues Nationales (DNENF-LN).
This benchmark accompanies the paper "Where Are We at with Automatic Speech Recognition for the Bambara Language?" and the public leaderboard at MALIBA-AI/bambara-asr-leaderboard.… See the full description on the dataset page: https://huggingface.co/datasets/MALIBA-AI/bambara-asr-benchmark.Bambara_AudioSynthetique_V1_LEGACY
⚠️ [OBSOLETE / INCOMPLET] Bambara Audio Dataset - Version Archivée
Attention : Cette version est obsolète et ne contient qu'une fraction des données disponibles.
La Version 3 de ce projet est désormais la référence. Elle contient l'intégralité du corpus (42 000 fichiers contre seulement une partie ici) et a été optimisée techniquement.
👉 Accéder au Corpus Complet V3 (42 000 audios - 44.4h)
Pourquoi passer absolument à la V3 ?
Volume : Accès à… See the full description on the dataset page: https://huggingface.co/datasets/OumarDicko/Bambara_AudioSynthetique_V1_LEGACY.bambara-orthography
Bambara orthography and text normalization
Writing conventions for Bambara (Bamanankan) in the standard Latin alphabet, a normalization table from common ASCII spellings to the standard, and a reference normalizer in Python. Maintained by Kooma.
Why this exists: Bambara is written in many ways in the wild — with or without ɛ/ɔ/ɲ/ŋ, with ny/ng digraphs, with or without tone marks, with French spellings for loanwords. Any comparison between two Bambara texts (a transcription and… See the full description on the dataset page: https://huggingface.co/datasets/kooma-ai/bambara-orthography.Dictionnaire_francais-wolof_et_francais-bambara
[!NOTE]
Dataset origin: https://books.google.fr/books?id=8xkOAAAAIAAJ&printsec=frontcover#v=onepage&q&f=false
bambara-mt-dataset
Bambara MT Dataset
Overview
The Bambara Machine Translation (MT) Dataset is a comprehensive collection of parallel text designed to advance natural language processing (NLP) for Bambara, a low-resource language spoken primarily in Mali. This dataset consolidates multiple sources to create the largest known Bambara MT dataset, supporting translation tasks and research to enhance language accessibility.
Languages
The dataset includes three language… See the full description on the dataset page: https://huggingface.co/datasets/MALIBA-AI/bambara-mt-dataset.bambara-english_sentence-pairsbambara-speech-kis-clean-split
Bambara Speech Dataset — Clean & Split
Dataset de reconnaissance vocale en bambara, nettoyé et splitté pour le fine-tuning de modèles ASR (ex: Whisper).
La source principale des données brutes est RobotsMali/bam-asr-early, auquel un remerciement chaleureux lui est attribué mais aussi à d'autres personnes référencées ci-dessous dans la section citation.
Statistiques
Total : 35 342 échantillons
Train : 24 738
Validation : 3 535
Test : 7 069
Durée moyenne : 3.23s… See the full description on the dataset page: https://huggingface.co/datasets/kalilouisangare/bambara-speech-kis-clean-split.alpaca-bambara-cleanedThis repository contains the dataset used for the TaCo paper.
Please refer to the paper for more details: OpenReview
If you have used our dataset, please cite it as follows:
Citation
@inproceedings{upadhayay2024taco,
title={TaCo: Enhancing Cross-Lingual Transfer for Low-Resource Languages in {LLM}s through Translation-Assisted Chain-of-Thought Processes},
author={Bibek Upadhayay and Vahid Behzadan},
booktitle={5th Workshop on practical ML for limited/low resource settings, ICLR},
year={2024}… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca-bambara-cleaned.bambara-lm-qaalpaca_bambara_tacoThis repository contains the dataset used for the TaCo paper.
The dataset follows the style outlined in the TaCo paper, as follows:
{
"instruction": "instruction in xx",
"input": "input in xx",
"output": "Instruction in English: instruction in en ,
Response in English: response in en ,
Response in xx: response in xx "
}
Please refer to the paper for more details: OpenReview
If you have used our dataset, please cite it as follows:
Citation… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca_bambara_taco.bambara-stt-test2bambara-french
[!NOTE]
Dataset origin: https://www.kaggle.com/datasets/ozaresearch1/bambara-french-parallel-dataset
Introduction
Bambara, also called Bamanankan or Bamana, is a language widely used as a vehicular and commercial language in West Africa and one of the national languages of Mali. Being member of the Mande language family, it is part of the main group in number of speakers, namely the Mandingo language group. This group includes, in addition to Bambara, Dioula in Côte d’Ivoire and… See the full description on the dataset page: https://huggingface.co/datasets/FrancophonIA/bambara-french.Code-170k-bambara
Dataset Description
Code-170k-bambara is a groundbreaking dataset containing 176,999 programming conversations, originally sourced from glaiveai/glaive-code-assistant-v2 and translated into Bambara, making coding education accessible to Bambara speakers.
🌟 Key Features
176,999 high-quality conversations about programming and coding
Pure Bambara language - democratizing coding education
Multi-turn dialogues covering various programming concepts
Diverse topics: algorithms… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/Code-170k-bambara.bambara-english-emotions-corpus
Bambara-english Emotion Analysis Corpus
Dataset Description
This dataset contains emotion-labeled text data in Bambara-english for emotion classification (joy, sadness, anger, fear, surprise, disgust, neutral). Emotions were extracted and processed from the English meanings of the sentences using the model j-hartmann/emotion-english-distilroberta-base. The dataset is part of a larger collection of African language emotion analysis resources.
Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/bambara-english-emotions-corpus.bambara-fon_sentence-pairs
Bambara-Fon_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Bambara-Fon_Sentence-Pairs
Number of Rows: 25525
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/bambara-fon_sentence-pairs.bambara-asrtest-bambara-ttsbambara-mt-dataset
Multilingual Parallel Dataset: Bambara-French-English
This dataset contains parallel text in three languages: Bambara (Bamanankan), French, and English. It combines content from the EGAFE educational books project by RobotMali and "La Guerre des Griots de Kita 1985" by Barbara G. Hoffman.
Dataset Overview
EGAFE Project
EGAFE (AI for Education) is an innovative project transforming education in Mali through advanced technology. The project focuses on… See the full description on the dataset page: https://huggingface.co/datasets/djelia/bambara-mt-dataset.bambara-french_sentence-pairs
Bambara-French_Sentence-Pairs Dataset
This dataset can be used for machine translation, sentence alignment, or other natural language processing tasks.
It is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Bambara-French_Sentence-Pairs
File Size: 39537941 bytes
Languages: Bambara, French
Dataset Description
The dataset contains sentence pairs in… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/bambara-french_sentence-pairs.bambara-mt-v2
bambara-mt-v2
An aggregated Bambara (Bamanankan, bm / bam_Latn) machine-translation corpus pairing
Bambara with French and English, assembled from eight upstream sources.
Load
from datasets import load_dataset
# aligned table with provenance
mt = load_dataset("djelia/bambara-mt-v2", "default", split="train")
# directional training pairs
pairs = load_dataset("djelia/bambara-mt-v2", "source_target_style", split="train")
Configs
Config
Rows… See the full description on the dataset page: https://huggingface.co/datasets/djelia/bambara-mt-v2.
