datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Hausa
Hausa Ajami OCR Dataset
Ce dataset contient des paires image/transcription de manuscrits haoussa en écriture ajami (écriture arabe adaptée au haoussa).
Contenu
Chaque ligne du fichier data/train/metadata.jsonl correspond à une ligne de texte ajami segmentée, avec :
file_name : nom du fichier image correspondant (image de la ligne, recadrée)
transcript : translittération en écriture latine de la ligne
source : identifiant du manuscrit d'origine (voir tableau… See the full description on the dataset page: https://huggingface.co/datasets/IntelligenceResearchLab/Hausa.hausa_dataset_encodedbible_tts_hausa
Dataset Card for BibleTTS Hausa
Dataset Summary
BibleTTS is a large high-quality open Text-to-Speech dataset with up to 80 hours of single speaker, studio quality 48kHz recordings.
This is a Hausa part of the dataset. Aligned hours: 86.6, aligned verses: 40,603.
Languages
Hausa
Dataset Structure
Data Fields
audio: audio path
sentence: transcription of the audio
locale: always set to ha
book: 3-char book encoding
verse: verse id… See the full description on the dataset page: https://huggingface.co/datasets/vpetukhov/bible_tts_hausa.African_voices_hausa
🇳🇬 WaZoBiaSpeech: 1,000+ Hour Hausa (hau) Corpus
Version: 30 Nov 2025
NOTE: This dataset is subject to regular Updates, corrections, and expansions. Please check this repository regularly for the latest release.
🌍 Dataset Overview
WaZoBiaSpeech is a large-scale, high-quality, fully transcribed speech dataset for Hausa (hau). This corpus is designed to accelerate the development of speech technology in African contexts, promoting linguistic diversity and… See the full description on the dataset page: https://huggingface.co/datasets/Africanvoice/African_voices_hausa.W_hausa_v1hausa_ajami_ocrafrispeech-hausa
Dataset Card for "afrispeech-hausa"
More Information needed
W_hausa_v3
Cleaned Hausa Speech Dataset v3
A cleaned and processed Hausa speech dataset built from multiple open-source Hugging Face datasets.
Dataset Description
This dataset contains cleaned, normalized, and deduplicated Hausa speech audio with aligned transcriptions. All audio is:
Sample rate: 16,000 Hz (mono)
Format: FLAC (lossless, embedded in Parquet)
Duration range: 1–30 seconds per clip
Loudness normalized: -20 dBFS RMS
VAD trimmed: Non-speech segments removed with… See the full description on the dataset page: https://huggingface.co/datasets/suleiman2003/W_hausa_v3.W_hausa_v69jalingo-reviewed-hausa-batch-0W_hausa_v4W_hausa_v7
Unified Hausa Speech Dataset v5
Dataset Description
A large-scale, cleaned, deduplicated, and quality-filtered Hausa speech dataset compiled from multiple open-source collections. Designed for Text-to-Speech (TTS) and Automatic Speech Recognition (ASR) research.
All audio is 16 kHz mono FLAC, silence-trimmed, loudness-normalized to -20 dBFS, and sorted by speaker_id so that all clips from the same speaker appear consecutively.
Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/suleiman2003/W_hausa_v7.W_hausa_v2hausa
African Voices Dataset
Multi-speaker voice dataset for African languages.
Hausa
Audio + transcript pairs organized by speaker, with age_group, domain, and gender metadata.
hausa_response_gemma_dfthausa_voa_topics
Dataset Card for Hausa VOA News Topic Classification dataset (hausa_voa_topics)
Dataset Summary
A news headline topic classification dataset, similar to AG-news, for Hausa. The news headlines were collected from VOA Hausa.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
Hausa (ISO 639-1: ha)
Dataset Structure
Data Instances
An instance consists of a news title sentence and the corresponding topic label.… See the full description on the dataset page: https://huggingface.co/datasets/UdS-LSV/hausa_voa_topics.W_hausa_v5s2tt-hausa-englishhausa_dataset_encoded_repo20to30hausa_common_voiceThis dataset is from the common voice corpus 7.0 using the Hausa dataset
Hausa-Synthetic-ASR-Dataset-XTTSSynthetic Hausa ASR dataset generated using a fine-tuned version of the XTTS-v2 model.
Sample rate: 24kHz.
Total duration: 574 hours.
hausa-audio-resampledNaijaSenti-TwitterNaijaSenti is the first large-scale human-annotated Twitter sentiment dataset for the four most widely spoken languages in Nigeria — Hausa, Igbo, Nigerian-Pidgin, and Yorùbá — consisting of around 30,000 annotated tweets per language, including a significant fraction of code-mixed tweets.naija-voices-hausa-split_0-1hausa
Hausa Dataset
The vocabulary foundation is organized by linguistic categories (pronouns, verbs, nouns, adjectives) with over 200 core Hausa words.
Key Features
Core Vocabulary Categories:
Pronouns with gender distinctions (kai/ke for masculine/feminine 'you')
Verbs covering daily activities and essential actions
Nouns spanning family, nature, time, and cultural concepts
Adjectives with proper Hausa formations
Numbers from basic counting to large values
Time… See the full description on the dataset page: https://huggingface.co/datasets/0xnu/hausa.hausa-audio-questionsclean_hausa_datasetnaija-voices-hausa-split_0-6naija-voices-hausa-split_2-4fongbe-hausa-asr-dataset
Fongbe-Hausa ASR Dataset (Semi-Supervised)
This dataset provides ~6,770 audio-transcription pairs for Fongbe (fon) and Hausa (hau). It was created using a semi-supervised pipeline to convert long-form video content into a training-ready format for Automatic Speech Recognition (ASR).
Dataset Details
Total Examples: 6,770
Audio Format: WAV (16kHz, Mono)
Languages: Fongbe (Benin), Hausa (Nigeria/West Africa)
Annotation: Semi-supervised (Machine-generated labels)
License:… See the full description on the dataset page: https://huggingface.co/datasets/Professor/fongbe-hausa-asr-dataset.
