CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01IntelligenceResearchLab /Hausa Hausa Ajami OCR Dataset Ce dataset contient des paires image/transcription de manuscrits haoussa en écriture ajami (écriture arabe adaptée au haoussa). Contenu Chaque ligne du fichier data/train/metadata.jsonl correspond à une ligne de texte ajami segmentée, avec : file_name : nom du fichier image correspondant (image de la ligne, recadrée) transcript : translittération en écriture latine de la ligne source : identifiant du manuscrit d'origine (voir tableau… See the full description on the dataset page: https://huggingface.co/datasets/IntelligenceResearchLab/Hausa.imageimage-to-text1K<n<10K3 likes1.4k downloads19h agoHugging Face02MeghanaKap /hausa_dataset_encodedgatedtext100K<n<1M0 likes927 downloads18d agoHugging Face03vpetukhov /bible_tts_hausa Dataset Card for BibleTTS Hausa Dataset Summary BibleTTS is a large high-quality open Text-to-Speech dataset with up to 80 hours of single speaker, studio quality 48kHz recordings. This is a Hausa part of the dataset. Aligned hours: 86.6, aligned verses: 40,603. Languages Hausa Dataset Structure Data Fields audio: audio path sentence: transcription of the audio locale: always set to ha book: 3-char book encoding verse: verse id… See the full description on the dataset page: https://huggingface.co/datasets/vpetukhov/bible_tts_hausa.textautomatic-speech-recognition10K<n<100K7 likes556 downloads4y agoHugging Face04Africanvoice /African_voices_hausa 🇳🇬 WaZoBiaSpeech: 1,000+ Hour Hausa (hau) Corpus Version: 30 Nov 2025 NOTE: This dataset is subject to regular Updates, corrections, and expansions. Please check this repository regularly for the latest release. 🌍 Dataset Overview WaZoBiaSpeech is a large-scale, high-quality, fully transcribed speech dataset for Hausa (hau). This corpus is designed to accelerate the development of speech technology in African contexts, promoting linguistic diversity and… See the full description on the dataset page: https://huggingface.co/datasets/Africanvoice/African_voices_hausa.audio10K<n<100K0 likes437 downloads3d agoHugging Face05suleiman2003 /W_hausa_v1audio100K<n<1M0 likes398 downloads2mo agoHugging Face06abdoulkarim1 /hausa_ajami_ocrimagen<1K0 likes392 downloads3mo agoHugging Face07robello2 /afrispeech-hausa Dataset Card for "afrispeech-hausa" More Information needed text1K<n<10K0 likes382 downloads11mo agoHugging Face08suleiman2003 /W_hausa_v3 Cleaned Hausa Speech Dataset v3 A cleaned and processed Hausa speech dataset built from multiple open-source Hugging Face datasets. Dataset Description This dataset contains cleaned, normalized, and deduplicated Hausa speech audio with aligned transcriptions. All audio is: Sample rate: 16,000 Hz (mono) Format: FLAC (lossless, embedded in Parquet) Duration range: 1–30 seconds per clip Loudness normalized: -20 dBFS RMS VAD trimmed: Non-speech segments removed with… See the full description on the dataset page: https://huggingface.co/datasets/suleiman2003/W_hausa_v3.audioautomatic-speech-recognition100K<n<1M0 likes330 downloads2mo agoHugging Face09suleiman2003 /W_hausa_v6audio100K<n<1M0 likes235 downloads2mo agoHugging Face10voicedata /9jalingo-reviewed-hausa-batch-0audio1K<n<10K0 likes225 downloads1h agoHugging Face11suleiman2003 /W_hausa_v4audio100K<n<1M0 likes217 downloads2mo agoHugging Face12suleiman2003 /W_hausa_v7 Unified Hausa Speech Dataset v5 Dataset Description A large-scale, cleaned, deduplicated, and quality-filtered Hausa speech dataset compiled from multiple open-source collections. Designed for Text-to-Speech (TTS) and Automatic Speech Recognition (ASR) research. All audio is 16 kHz mono FLAC, silence-trimmed, loudness-normalized to -20 dBFS, and sorted by speaker_id so that all clips from the same speaker appear consecutively. Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/suleiman2003/W_hausa_v7.audiotext-to-speech100K<n<1M0 likes217 downloads10d agoHugging Face13suleiman2003 /W_hausa_v2audio100K<n<1M0 likes208 downloads2mo agoHugging Face14voicedata /hausa African Voices Dataset Multi-speaker voice dataset for African languages. Hausa Audio + transcript pairs organized by speaker, with age_group, domain, and gender metadata. audio10K<n<100K0 likes201 downloads19d agoHugging Face15vaghawan /hausa_response_gemma_dfttext1K<n<10K0 likes186 downloads1mo agoHugging Face16UdS-LSV /hausa_voa_topics Dataset Card for Hausa VOA News Topic Classification dataset (hausa_voa_topics) Dataset Summary A news headline topic classification dataset, similar to AG-news, for Hausa. The news headlines were collected from VOA Hausa. Supported Tasks and Leaderboards [More Information Needed] Languages Hausa (ISO 639-1: ha) Dataset Structure Data Instances An instance consists of a news title sentence and the corresponding topic label.… See the full description on the dataset page: https://huggingface.co/datasets/UdS-LSV/hausa_voa_topics.texttext-classification1K<n<10K0 likes175 downloads2y agoHugging Face17suleiman2003 /W_hausa_v5audio100K<n<1M0 likes175 downloads2mo agoHugging Face18Mawube /s2tt-hausa-englishaudio10K<n<100K0 likes174 downloads6mo agoHugging Face19MeghanaKap /hausa_dataset_encoded_repo20to30text100K<n<1M0 likes158 downloads23d agoHugging Face20Arnold /hausa_common_voiceThis dataset is from the common voice corpus 7.0 using the Hausa dataset tabular1K<n<10K2 likes149 downloads5y agoHugging Face21CLEAR-Global /Hausa-Synthetic-ASR-Dataset-XTTSgatedSynthetic Hausa ASR dataset generated using a fine-tuned version of the XTTS-v2 model. Sample rate: 24kHz. Total duration: 574 hours. audioautomatic-speech-recognition100K<n<1M1 likes147 downloads1y agoHugging Face22vaghawan /hausa-audio-resampledaudio100K<n<1M0 likes144 downloads8mo agoHugging Face23HausaNLP /NaijaSenti-TwitterNaijaSenti is the first large-scale human-annotated Twitter sentiment dataset for the four most widely spoken languages in Nigeria — Hausa, Igbo, Nigerian-Pidgin, and Yorùbá — consisting of around 30,000 annotated tweets per language, including a significant fraction of code-mixed tweets.texttext-classification10K<n<100K8 likes134 downloads3y agoHugging Face24EYEDOL /naija-voices-hausa-split_0-1audio10K<n<100K0 likes131 downloads1y agoHugging Face250xnu /hausa Hausa Dataset The vocabulary foundation is organized by linguistic categories (pronouns, verbs, nouns, adjectives) with over 200 core Hausa words. Key Features Core Vocabulary Categories: Pronouns with gender distinctions (kai/ke for masculine/feminine 'you') Verbs covering daily activities and essential actions Nouns spanning family, nature, time, and cultural concepts Adjectives with proper Hausa formations Numbers from basic counting to large values Time… See the full description on the dataset page: https://huggingface.co/datasets/0xnu/hausa.tabular1M<n<10M1 likes118 downloads1y agoHugging Face26vaghawan /hausa-audio-questionsaudio10K<n<100K0 likes116 downloads4mo agoHugging Face27chukypedro /clean_hausa_datasetaudio100K<n<1M0 likes107 downloads1y agoHugging Face28EYEDOL /naija-voices-hausa-split_0-6audio10K<n<100K0 likes96 downloads1y agoHugging Face29EYEDOL /naija-voices-hausa-split_2-4audio10K<n<100K0 likes95 downloads1y agoHugging Face30Professor /fongbe-hausa-asr-dataset Fongbe-Hausa ASR Dataset (Semi-Supervised) This dataset provides ~6,770 audio-transcription pairs for Fongbe (fon) and Hausa (hau). It was created using a semi-supervised pipeline to convert long-form video content into a training-ready format for Automatic Speech Recognition (ASR). Dataset Details Total Examples: 6,770 Audio Format: WAV (16kHz, Mono) Languages: Fongbe (Benin), Hausa (Nigeria/West Africa) Annotation: Semi-supervised (Machine-generated labels) License:… See the full description on the dataset page: https://huggingface.co/datasets/Professor/fongbe-hausa-asr-dataset.audioautomatic-speech-recognition1K<n<10K0 likes95 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.