CoolFace
29 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01soynade-research /Bambara-Speech-Translation-Data AfVoices-Translated (Bambara-English) This is a Bambara speech translation dataset, which is built on the African Next Voices (AfVoices) Bambara ASR corpus. It provides English translations for the human-corrected subset of the original collection, creating a parallel corpus for Bambara-English machine translation and speech-to-text tasks. Methodology We machine-translated the human-validated transcriptions from AfVoices using the Oolel-translator repository. Inference… See the full description on the dataset page: https://huggingface.co/datasets/soynade-research/Bambara-Speech-Translation-Data.audioautomatic-speech-recognition100K<n<1M1 likes549 downloads7mo agoHugging Face02madoss /merged-bambara-dioula-datasetaudio10K<n<100K0 likes210 downloads2mo agoHugging Face03OBY632 /merged-bambara-dioula-datasetaudio10K<n<100K0 likes110 downloads6mo agoHugging Face04OumarDicko /Bambara_AudioSynthetique_42K_V3 Description Ce corpus comprend 42 000 entrées audio synthétiques en langue Bambara (bm), totalisant environ 44,4 heures d'enregistrement. Cette version 3 a été convertie au format Parquet pour optimiser les performances de lecture et garantir une compatibilité totale avec le Dataset Viewer de Hugging Face. Origine et Traitement des Données Textuelles Le corpus de texte a été constitué par l'agrégation de plusieurs sources linguistiques afin de garantir un volume suffisant… See the full description on the dataset page: https://huggingface.co/datasets/OumarDicko/Bambara_AudioSynthetique_42K_V3.audioautomatic-speech-recognition10K<n<100K2 likes100 downloads8mo agoHugging Face05djelia /bambara-tts-waxal bambara-tts-waxal Bambara studio speech from the WAXAL corpus — 1,926 recordings, 16 hours, 8 speakers, 44.1 kHz mono. Load from datasets import load_dataset ds = load_dataset("djelia/bambara-tts-waxal", "google_waxal", split="train") Splits: train, validation, test. Fields Field Description audio 44.1 kHz mono text Transcript speaker_id Speaker identifier (8 distinct) gender Speaker gender locale Locale code id Record… See the full description on the dataset page: https://huggingface.co/datasets/djelia/bambara-tts-waxal.audiotext-to-speech1K<n<10K0 likes69 downloads2mo agoHugging Face06MALIBA-AI /bambara-asr-benchmark Bambara ASR Benchmark The first standardized evaluation set for Automatic Speech Recognition in Bambara (Bamanankan). One hour of studio-quality constitutional text, transcribed and validated by linguists from Mali's Direction Nationale de l'Éducation Non Formelle et des Langues Nationales (DNENF-LN). This benchmark accompanies the paper "Where Are We at with Automatic Speech Recognition for the Bambara Language?" and the public leaderboard at MALIBA-AI/bambara-asr-leaderboard.… See the full description on the dataset page: https://huggingface.co/datasets/MALIBA-AI/bambara-asr-benchmark.audioautomatic-speech-recognitionn<1K0 likes49 downloads8mo agoHugging Face07OumarDicko /Bambara_AudioSynthetique_V1_LEGACY ⚠️ [OBSOLETE / INCOMPLET] Bambara Audio Dataset - Version Archivée Attention : Cette version est obsolète et ne contient qu'une fraction des données disponibles. La Version 3 de ce projet est désormais la référence. Elle contient l'intégralité du corpus (42 000 fichiers contre seulement une partie ici) et a été optimisée techniquement. 👉 Accéder au Corpus Complet V3 (42 000 audios - 44.4h) Pourquoi passer absolument à la V3 ? Volume : Accès à… See the full description on the dataset page: https://huggingface.co/datasets/OumarDicko/Bambara_AudioSynthetique_V1_LEGACY.audion<1K0 likes40 downloads9mo agoHugging Face08kalilouisangare /bambara-speech-kis-clean-split Bambara Speech Dataset — Clean & Split Dataset de reconnaissance vocale en bambara, nettoyé et splitté pour le fine-tuning de modèles ASR (ex: Whisper). La source principale des données brutes est RobotsMali/bam-asr-early, auquel un remerciement chaleureux lui est attribué mais aussi à d'autres personnes référencées ci-dessous dans la section citation. Statistiques Total : 35 342 échantillons Train : 24 738 Validation : 3 535 Test : 7 069 Durée moyenne : 3.23s… See the full description on the dataset page: https://huggingface.co/datasets/kalilouisangare/bambara-speech-kis-clean-split.audio10K<n<100K2 likes33 downloads7mo agoHugging Face09Danube /test-bambara-ttsaudio1K<n<10K0 likes19 downloads2y agoHugging Face10oza75 /bambara-asrgatedaudio100K<n<1M3 likes18 downloads2y agoHugging Face11djelia /bambara-audiogated Djelia Bambara Audio Dataset Dataset Description The Djelia Bambara Audio Dataset is a comprehensive resource aimed at supporting research and development in Bambara language processing. This dataset consists of audio extracted from YouTube videos, denoised and diarized to ensure high-quality segments. Additionally, it features a semi-annotated subset with transcriptions generated using the Djelia Whisper v1 model. Features Audio: High-quality audio clips… See the full description on the dataset page: https://huggingface.co/datasets/djelia/bambara-audio.audioautomatic-speech-recognition100K<n<1M1 likes17 downloads2y agoHugging Face12OumarDicko /Bambara_AudioSynthetique_V2_LEGACY ⚠️ [OBSOLETE / INCOMPLET] Bambara Audio Dataset - Version Archivée Attention : Cette version est obsolète et ne contient qu'une fraction des données disponibles. La Version 3 de ce projet est désormais la référence. Elle contient l'intégralité du corpus (42 000 fichiers contre seulement une partie ici) et a été optimisée techniquement. 👉 Accéder au Corpus Complet V3 (42 000 audios - 44.4h) Pourquoi passer absolument à la V3 ? Volume : Accès à… See the full description on the dataset page: https://huggingface.co/datasets/OumarDicko/Bambara_AudioSynthetique_V2_LEGACY.audion<1K0 likes16 downloads9mo agoHugging Face13oza75 /bambara-ttsgated Overview Project This dataset is part of a larger initiative aimed at empowering Bambara speakers to access global knowledge without language barriers. Our goal is to eliminate the need for Bambara speakers to learn a secondary language before they can acquire new information or skills. By providing a robust dataset for Text-to-Speech (TTS) applications, we aim to support the creation of tools for bambara language, thus democratizing access to knowledge.… See the full description on the dataset page: https://huggingface.co/datasets/oza75/bambara-tts.audiotext-to-speech100K<n<1M5 likes16 downloads2y agoHugging Face14djelia /bambara-asr-v2gated bambara-asr-v2 Multi-corpus Bambara speech — 185,708 examples, ~366 hours, 53.7 GB of Parquet. Seven configs, each a train / dev / test triple of 16 kHz audio paired with a text target. Every config draws on a single upstream corpus, so you can mix and weight them yourself. Access is gated with manual approval — request it on the dataset page and authenticate (hf auth login or HF_TOKEN) before loading. Load from datasets import load_dataset jeli =… See the full description on the dataset page: https://huggingface.co/datasets/djelia/bambara-asr-v2.audioautomatic-speech-recognition100K<n<1M0 likes16 downloads2mo agoHugging Face15djelia /bambara-asr-dataset-ygated bambara-asr-dataset-y Bambara speech paired with the French source line it renders. 58,447 rows, 36.66 hours, 48.77 GB of Parquet. The only text column is fr — there is no Bambara text here. Load The config is default and the splits are not_combined and combined — there is no train split, so a bare load_dataset returns a DatasetDict keyed by those two names. from datasets import load_dataset short = load_dataset("djelia/bambara-asr-dataset-y"… See the full description on the dataset page: https://huggingface.co/datasets/djelia/bambara-asr-dataset-y.audioautomatic-speech-recognition10K<n<100K0 likes15 downloads2mo agoHugging Face16oza75 /bible-bambara-audiogated Bible Bambara Audio Dataset Overview This dataset contains audio recordings of Bible passages in Bambara language along with their transcriptions. The dataset consists of approximately 42.7 hours of audio content, making it a valuable resource for speech processing tasks in Bambara language. Project This dataset is part of a larger initiative to preserve and digitize audio & texts in Bambara language, making them accessible in both text and audio… See the full description on the dataset page: https://huggingface.co/datasets/oza75/bible-bambara-audio.audio10K<n<100K0 likes14 downloads2y agoHugging Face17sudoping01 /open-bambara-asr-datasetgatedaudio10K<n<100K0 likes13 downloads2y agoHugging Face18sudoping01 /bambara-speech-recognition-benchmarkgatedaudio1K<n<10K0 likes10 downloads2y agoHugging Face19djelia /bambara-synthetic-audiogated bambara-synthetic-audio 102,310 utterances of synthetic Bambara speech, generated by a text-to-speech model over Bambara text. Every clip is machine-generated; no human voice is recorded here. Load from datasets import load_dataset # Enhanced set, with per-clip quality scores semi = load_dataset("djelia/bambara-synthetic-audio", "semi-clean", split="train") # Larger generation set, no quality scores v2 = load_dataset("djelia/bambara-synthetic-audio", "tts_v2"… See the full description on the dataset page: https://huggingface.co/datasets/djelia/bambara-synthetic-audio.audioautomatic-speech-recognition100K<n<1M1 likes10 downloads2mo agoHugging Face20sudoping01 /maliba_bambaragatedaudio100K<n<1M1 likes9 downloads3mo agoHugging Face21djelia /bambara-audio-bgated bambara-audio-b Bambara speech derived from scripture recordings, published in four processing stages: raw segments, a length-filtered version, a speaker-diarized long-form cut, and a CTC forced-alignment cut. 30.55 GB of Parquet. Access is gated with manual approval — request it on the dataset page and authenticate (hf auth login or HF_TOKEN) before loading. Load from datasets import load_dataset ds = load_dataset("djelia/bambara-audio-b", "short-filtered"… See the full description on the dataset page: https://huggingface.co/datasets/djelia/bambara-audio-b.audioautomatic-speech-recognition10K<n<100K1 likes7 downloads2mo agoHugging Face22djelia /bambara-audio-ygated bambara-audio-y Bambara speech paired with the French source line it renders, a written Bambara translation of that line, and a machine transcription of the audio. 58,447 rows, 36.66 hours, 48.77 GB of Parquet. Load The config is default and the splits are not_combined and combined — there is no train split, so a bare load_dataset returns a DatasetDict keyed by those two names. from datasets import load_dataset short = load_dataset("djelia/bambara-audio-y"… See the full description on the dataset page: https://huggingface.co/datasets/djelia/bambara-audio-y.audioautomatic-speech-recognition10K<n<100K0 likes7 downloads2mo agoHugging Face23djelia /bambara-asrgated bambara-asr Multi-task Bambara speech: transcription, speech-to-text translation into French and English, and a multilingual training mix. 16 kHz audio in Parquet across nine configs. Access is gated with manual approval — request it on the dataset page and authenticate (hf auth login or HF_TOKEN) before loading. Load from datasets import load_dataset ds = load_dataset("djelia/bambara-asr", "bm-to-bm", split="train") Every config has train and test splits.… See the full description on the dataset page: https://huggingface.co/datasets/djelia/bambara-asr.audioautomatic-speech-recognition100K<n<1M0 likes7 downloads2mo agoHugging Face24sudoping01 /bambara-numbersgatedaudio1M<n<10M0 likes7 downloads1y agoHugging Face25sudoping01 /bambara-audiogatedaudio100K<n<1M1 likes6 downloads2y agoHugging Face26djelia /bambara-asr-evaluationgated bambara-asr-evaluation A Bambara ASR benchmark: 1,295 utterances, 2.04 hours of 16 kHz audio with reference transcripts. Monolingual Bambara transcription — audio in, transcript out, WER out. Load from datasets import load_dataset ds = load_dataset("djelia/bambara-asr-evaluation", split="test") print(ds[0]["text"], ds[0]["source_dataset"]) One config and one split, so no config argument is needed. Config Split Rows Audio default test 1,295 2.043 h… See the full description on the dataset page: https://huggingface.co/datasets/djelia/bambara-asr-evaluation.audioautomatic-speech-recognition1K<n<10K0 likes6 downloads2mo agoHugging Face27Panga-Azazia /bambara-ext-eval-datagatedaudion<1K0 likes3 downloads1y agoHugging Face28Panga-Azazia /Bambara-Keyword-Spotting-Auggatedaudio1K<n<10K0 likes3 downloads1y agoHugging Face29Panga-Azazia /Bambara-Keyword-Spottinggatedaudion<1K0 likes2 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.