datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Bambara_Translation_Corpus_FR-BM
🌍 Bambara Translation Corpus: Direct & Instruction-Tuned (FR-BM)
🚀 Overview & Vision
The Bambara Translation Corpus is a comprehensive bilingual dataset designed to bridge French and Bamanankan across two distinct paradigms: Direct Neural Machine Translation (NMT) and Prompt-Based Instruction Tuning.
This dual-mode architecture caters to both traditional seq2seq translation models and modern instruction-tuned Large Language Models (LLMs).
📊 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Makan09/Bambara_Translation_Corpus_FR-BM.Bambara-dataset_conversation
license: apache-2.0
language:
- fr
- bm
tags:
- bambara
- bamanankan
- instruction-tuning
- llm-alignment
- african-languages
- low-resource-nlp
- conversational-ai
task_categories:
- text-generation
- conditional-text-generation
size_categories:
- 10K<n<50K
pretty_name: Bambara Instruction Tuning Corpus (FR-BM)
🌍 Bambara Instruction Tuning Corpus (FR-BM)
🚀 Overview & Vision
Welcome to the Bambara Instruction Tuning Corpus… See the full description on the dataset page: https://huggingface.co/datasets/Makan09/Bambara-dataset_conversation.Bambara_texts_raws_corpus
🌍 Bambara Massive Raw Text Corpus (1.7M+ Lines)
🚀 Overview & Vision
Welcome to the Bambara Massive Raw Text Corpus—a monumental milestone for African language technology. Featuring over 1.7 million lines of raw Bamanankan text, this repository represents an unprecedented scale of unstructured linguistic data for a low-resource West African language.
Pre-training foundational models from scratch or performing Continued Pre-Training (CPT) on existing open-source… See the full description on the dataset page: https://huggingface.co/datasets/Makan09/Bambara_texts_raws_corpus.Code-170k-bambara
Dataset Description
Code-170k-bambara is a groundbreaking dataset containing 176,999 programming conversations, originally sourced from glaiveai/glaive-code-assistant-v2 and translated into Bambara, making coding education accessible to Bambara speakers.
🌟 Key Features
176,999 high-quality conversations about programming and coding
Pure Bambara language - democratizing coding education
Multi-turn dialogues covering various programming concepts
Diverse topics: algorithms… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/Code-170k-bambara.bambara-lm-qa
Bambara-LM-QA Dataset
Dataset Summary
The Bambara-LM-QA dataset is designed to support the fine-tuning of large language models (LLMs) for the Bambara language. It encompasses a variety of tasks to improve Bambara NLP applications, including:
Translation tasks:
French → Bambara
Bambara → French
English → Bambara
Bambara → English
Alpaca dataset (Bambara version):
Automatically translated using Google Translate from the original Alpaca dataset.
ASR Transcription… See the full description on the dataset page: https://huggingface.co/datasets/djelia/bambara-lm-qa.bambara-texts
Dataset Summary
The Bambara-Texts dataset is a collection of monolingual Bambara text designed for pretraining language models. It provides a diverse set of textual data to improve natural language processing (NLP) applications for the Bambara language.
This dataset can be used for:
Pretraining large language models (LLMs)
Building word embeddings for Bambara
Language modeling tasks such as masked language modeling (MLM) and autoregressive modeling
Corpus-based linguistic research… See the full description on the dataset page: https://huggingface.co/datasets/djelia/bambara-texts.
