CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01soynade-research /Bambara-Speech-Translation-Data AfVoices-Translated (Bambara-English) This is a Bambara speech translation dataset, which is built on the African Next Voices (AfVoices) Bambara ASR corpus. It provides English translations for the human-corrected subset of the original collection, creating a parallel corpus for Bambara-English machine translation and speech-to-text tasks. Methodology We machine-translated the human-validated transcriptions from AfVoices using the Oolel-translator repository. Inference… See the full description on the dataset page: https://huggingface.co/datasets/soynade-research/Bambara-Speech-Translation-Data.audioautomatic-speech-recognition100K<n<1M1 likes543 downloads7mo agoHugging Face02Diomande /bambara-whisper-featurestext100K<n<1M0 likes408 downloads5mo agoHugging Face03madoss /merged-bambara-dioula-datasetaudio10K<n<100K0 likes172 downloads2mo agoHugging Face04OBY632 /merged-bambara-dioula-datasetaudio10K<n<100K0 likes107 downloads6mo agoHugging Face05OumarDicko /Bambara_AudioSynthetique_42K_V3 Description Ce corpus comprend 42 000 entrées audio synthétiques en langue Bambara (bm), totalisant environ 44,4 heures d'enregistrement. Cette version 3 a été convertie au format Parquet pour optimiser les performances de lecture et garantir une compatibilité totale avec le Dataset Viewer de Hugging Face. Origine et Traitement des Données Textuelles Le corpus de texte a été constitué par l'agrégation de plusieurs sources linguistiques afin de garantir un volume suffisant… See the full description on the dataset page: https://huggingface.co/datasets/OumarDicko/Bambara_AudioSynthetique_42K_V3.audioautomatic-speech-recognition10K<n<100K2 likes98 downloads8mo agoHugging Face06Makan09 /Bambara_Translation_Corpus_FR-BM 🌍 Bambara Translation Corpus: Direct & Instruction-Tuned (FR-BM) 🚀 Overview & Vision The Bambara Translation Corpus is a comprehensive bilingual dataset designed to bridge French and Bamanankan across two distinct paradigms: Direct Neural Machine Translation (NMT) and Prompt-Based Instruction Tuning. This dual-mode architecture caters to both traditional seq2seq translation models and modern instruction-tuned Large Language Models (LLMs). 📊 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Makan09/Bambara_Translation_Corpus_FR-BM.translation1 likes88 downloads25d agoHugging Face07Makan09 /Bambara-dataset_conversation license: apache-2.0 language: - fr - bm tags: - bambara - bamanankan - instruction-tuning - llm-alignment - african-languages - low-resource-nlp - conversational-ai task_categories: - text-generation - conditional-text-generation size_categories: - 10K<n<50K pretty_name: Bambara Instruction Tuning Corpus (FR-BM) 🌍 Bambara Instruction Tuning Corpus (FR-BM) 🚀 Overview & Vision Welcome to the Bambara Instruction Tuning Corpus… See the full description on the dataset page: https://huggingface.co/datasets/Makan09/Bambara-dataset_conversation.texttext-generation10K<n<100K0 likes81 downloads25d agoHugging Face08djelia /bambara-tts-waxal bambara-tts-waxal Bambara studio speech from the WAXAL corpus — 1,926 recordings, 16 hours, 8 speakers, 44.1 kHz mono. Load from datasets import load_dataset ds = load_dataset("djelia/bambara-tts-waxal", "google_waxal", split="train") Splits: train, validation, test. Fields Field Description audio 44.1 kHz mono text Transcript speaker_id Speaker identifier (8 distinct) gender Speaker gender locale Locale code id Record… See the full description on the dataset page: https://huggingface.co/datasets/djelia/bambara-tts-waxal.audiotext-to-speech1K<n<10K0 likes66 downloads2mo agoHugging Face09Makan09 /Bambara_asr_test_wercalculation1K<n<10K0 likes66 downloads15d agoHugging Face10Makan09 /Bambara_texts_raws_corpus 🌍 Bambara Massive Raw Text Corpus (1.7M+ Lines) 🚀 Overview & Vision Welcome to the Bambara Massive Raw Text Corpus—a monumental milestone for African language technology. Featuring over 1.7 million lines of raw Bamanankan text, this repository represents an unprecedented scale of unstructured linguistic data for a low-resource West African language. Pre-training foundational models from scratch or performing Continued Pre-Training (CPT) on existing open-source… See the full description on the dataset page: https://huggingface.co/datasets/Makan09/Bambara_texts_raws_corpus.texttext-generation1M<n<10M0 likes60 downloads25d agoHugging Face11MALIBA-AI /bambara-asr-benchmark Bambara ASR Benchmark The first standardized evaluation set for Automatic Speech Recognition in Bambara (Bamanankan). One hour of studio-quality constitutional text, transcribed and validated by linguists from Mali's Direction Nationale de l'Éducation Non Formelle et des Langues Nationales (DNENF-LN). This benchmark accompanies the paper "Where Are We at with Automatic Speech Recognition for the Bambara Language?" and the public leaderboard at MALIBA-AI/bambara-asr-leaderboard.… See the full description on the dataset page: https://huggingface.co/datasets/MALIBA-AI/bambara-asr-benchmark.audioautomatic-speech-recognitionn<1K0 likes51 downloads7mo agoHugging Face12OumarDicko /Bambara_AudioSynthetique_V1_LEGACY ⚠️ [OBSOLETE / INCOMPLET] Bambara Audio Dataset - Version Archivée Attention : Cette version est obsolète et ne contient qu'une fraction des données disponibles. La Version 3 de ce projet est désormais la référence. Elle contient l'intégralité du corpus (42 000 fichiers contre seulement une partie ici) et a été optimisée techniquement. 👉 Accéder au Corpus Complet V3 (42 000 audios - 44.4h) Pourquoi passer absolument à la V3 ? Volume : Accès à… See the full description on the dataset page: https://huggingface.co/datasets/OumarDicko/Bambara_AudioSynthetique_V1_LEGACY.audion<1K0 likes48 downloads8mo agoHugging Face13kooma-ai /bambara-orthography Bambara orthography and text normalization Writing conventions for Bambara (Bamanankan) in the standard Latin alphabet, a normalization table from common ASCII spellings to the standard, and a reference normalizer in Python. Maintained by Kooma. Why this exists: Bambara is written in many ways in the wild — with or without ɛ/ɔ/ɲ/ŋ, with ny/ng digraphs, with or without tone marks, with French spellings for loanwords. Any comparison between two Bambara texts (a transcription and… See the full description on the dataset page: https://huggingface.co/datasets/kooma-ai/bambara-orthography.textn<1K1 likes48 downloads29d agoHugging Face14FrancophonIA /Dictionnaire_francais-wolof_et_francais-bambara [!NOTE] Dataset origin: https://books.google.fr/books?id=8xkOAAAAIAAJ&printsec=frontcover#v=onepage&q&f=false translation0 likes39 downloads1y agoHugging Face15MALIBA-AI /bambara-mt-dataset Bambara MT Dataset Overview The Bambara Machine Translation (MT) Dataset is a comprehensive collection of parallel text designed to advance natural language processing (NLP) for Bambara, a low-resource language spoken primarily in Mali. This dataset consolidates multiple sources to create the largest known Bambara MT dataset, supporting translation tasks and research to enhance language accessibility. Languages The dataset includes three language… See the full description on the dataset page: https://huggingface.co/datasets/MALIBA-AI/bambara-mt-dataset.text100K<n<1M2 likes36 downloads11mo agoHugging Face16michsethowusu /bambara-english_sentence-pairstext100K<n<1M0 likes35 downloads1y agoHugging Face17kalilouisangare /bambara-speech-kis-clean-split Bambara Speech Dataset — Clean & Split Dataset de reconnaissance vocale en bambara, nettoyé et splitté pour le fine-tuning de modèles ASR (ex: Whisper). La source principale des données brutes est RobotsMali/bam-asr-early, auquel un remerciement chaleureux lui est attribué mais aussi à d'autres personnes référencées ci-dessous dans la section citation. Statistiques Total : 35 342 échantillons Train : 24 738 Validation : 3 535 Test : 7 069 Durée moyenne : 3.23s… See the full description on the dataset page: https://huggingface.co/datasets/kalilouisangare/bambara-speech-kis-clean-split.audio10K<n<100K2 likes33 downloads7mo agoHugging Face18saillab /alpaca-bambara-cleanedThis repository contains the dataset used for the TaCo paper. Please refer to the paper for more details: OpenReview If you have used our dataset, please cite it as follows: Citation @inproceedings{upadhayay2024taco, title={TaCo: Enhancing Cross-Lingual Transfer for Low-Resource Languages in {LLM}s through Translation-Assisted Chain-of-Thought Processes}, author={Bibek Upadhayay and Vahid Behzadan}, booktitle={5th Workshop on practical ML for limited/low resource settings, ICLR}, year={2024}… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca-bambara-cleaned.text10K<n<100K1 likes28 downloads2y agoHugging Face19oza75 /bambara-lm-qatext100K<n<1M0 likes26 downloads2y agoHugging Face20saillab /alpaca_bambara_tacoThis repository contains the dataset used for the TaCo paper. The dataset follows the style outlined in the TaCo paper, as follows: { "instruction": "instruction in xx", "input": "input in xx", "output": "Instruction in English: instruction in en , Response in English: response in en , Response in xx: response in xx " } Please refer to the paper for more details: OpenReview If you have used our dataset, please cite it as follows: Citation… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca_bambara_taco.text10K<n<100K1 likes24 downloads2y agoHugging Face21Danube /bambara-stt-test21K<n<10K0 likes24 downloads2y agoHugging Face22FrancophonIA /bambara-french [!NOTE] Dataset origin: https://www.kaggle.com/datasets/ozaresearch1/bambara-french-parallel-dataset Introduction Bambara, also called Bamanankan or Bamana, is a language widely used as a vehicular and commercial language in West Africa and one of the national languages of Mali. Being member of the Mande language family, it is part of the main group in number of speakers, namely the Mandingo language group. This group includes, in addition to Bambara, Dioula in Côte d’Ivoire and… See the full description on the dataset page: https://huggingface.co/datasets/FrancophonIA/bambara-french.texttranslation10K<n<100K0 likes24 downloads1y agoHugging Face23michsethowusu /Code-170k-bambara Dataset Description Code-170k-bambara is a groundbreaking dataset containing 176,999 programming conversations, originally sourced from glaiveai/glaive-code-assistant-v2 and translated into Bambara, making coding education accessible to Bambara speakers. 🌟 Key Features 176,999 high-quality conversations about programming and coding Pure Bambara language - democratizing coding education Multi-turn dialogues covering various programming concepts Diverse topics: algorithms… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/Code-170k-bambara.texttext-generation100K<n<1M0 likes24 downloads11mo agoHugging Face24michsethowusu /bambara-english-emotions-corpus Bambara-english Emotion Analysis Corpus Dataset Description This dataset contains emotion-labeled text data in Bambara-english for emotion classification (joy, sadness, anger, fear, surprise, disgust, neutral). Emotions were extracted and processed from the English meanings of the sentences using the model j-hartmann/emotion-english-distilroberta-base. The dataset is part of a larger collection of African language emotion analysis resources. Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/bambara-english-emotions-corpus.texttext-classification10K<n<100K0 likes23 downloads1y agoHugging Face25michsethowusu /bambara-fon_sentence-pairs Bambara-Fon_Sentence-Pairs Dataset This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks. This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1 Metadata File Name: Bambara-Fon_Sentence-Pairs Number of Rows: 25525 Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/bambara-fon_sentence-pairs.text10K<n<100K0 likes20 downloads1y agoHugging Face26oza75 /bambara-asrgatedaudio100K<n<1M3 likes18 downloads2y agoHugging Face27Danube /test-bambara-ttsaudio1K<n<10K0 likes17 downloads2y agoHugging Face28djelia /bambara-mt-datasetgated Multilingual Parallel Dataset: Bambara-French-English This dataset contains parallel text in three languages: Bambara (Bamanankan), French, and English. It combines content from the EGAFE educational books project by RobotMali and "La Guerre des Griots de Kita 1985" by Barbara G. Hoffman. Dataset Overview EGAFE Project EGAFE (AI for Education) is an innovative project transforming education in Mali through advanced technology. The project focuses on… See the full description on the dataset page: https://huggingface.co/datasets/djelia/bambara-mt-dataset.text10K<n<100K1 likes17 downloads2y agoHugging Face29michsethowusu /bambara-french_sentence-pairs Bambara-French_Sentence-Pairs Dataset This dataset can be used for machine translation, sentence alignment, or other natural language processing tasks. It is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1 Metadata File Name: Bambara-French_Sentence-Pairs File Size: 39537941 bytes Languages: Bambara, French Dataset Description The dataset contains sentence pairs in… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/bambara-french_sentence-pairs.text100K<n<1M0 likes17 downloads1y agoHugging Face30djelia /bambara-mt-v2gated bambara-mt-v2 An aggregated Bambara (Bamanankan, bm / bam_Latn) machine-translation corpus pairing Bambara with French and English, assembled from eight upstream sources. Load from datasets import load_dataset # aligned table with provenance mt = load_dataset("djelia/bambara-mt-v2", "default", split="train") # directional training pairs pairs = load_dataset("djelia/bambara-mt-v2", "source_target_style", split="train") Configs Config Rows… See the full description on the dataset page: https://huggingface.co/datasets/djelia/bambara-mt-v2.texttranslation10K<n<100K1 likes17 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.