datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hausa_response_gemmaHausa
Hausa Ajami OCR Dataset
Ce dataset contient des paires image/transcription de manuscrits haoussa en écriture ajami (écriture arabe adaptée au haoussa).
Contenu
Chaque ligne du fichier data/train/metadata.jsonl correspond à une ligne de texte ajami segmentée, avec :
file_name : nom du fichier image correspondant (image de la ligne, recadrée)
transcript : translittération en écriture latine de la ligne
source : identifiant du manuscrit d'origine (voir tableau… See the full description on the dataset page: https://huggingface.co/datasets/IntelligenceResearchLab/Hausa.hausa_dataset_encodedW_hausa_v1bible_tts_hausa
Dataset Card for BibleTTS Hausa
Dataset Summary
BibleTTS is a large high-quality open Text-to-Speech dataset with up to 80 hours of single speaker, studio quality 48kHz recordings.
This is a Hausa part of the dataset. Aligned hours: 86.6, aligned verses: 40,603.
Languages
Hausa
Dataset Structure
Data Fields
audio: audio path
sentence: transcription of the audio
locale: always set to ha
book: 3-char book encoding
verse: verse id… See the full description on the dataset page: https://huggingface.co/datasets/vpetukhov/bible_tts_hausa.African_voices_hausa
🇳🇬 WaZoBiaSpeech: 1,000+ Hour Hausa (hau) Corpus
Version: 30 Nov 2025
NOTE: This dataset is subject to regular Updates, corrections, and expansions. Please check this repository regularly for the latest release.
🌍 Dataset Overview
WaZoBiaSpeech is a large-scale, high-quality, fully transcribed speech dataset for Hausa (hau). This corpus is designed to accelerate the development of speech technology in African contexts, promoting linguistic diversity and… See the full description on the dataset page: https://huggingface.co/datasets/Africanvoice/African_voices_hausa.hausa-cultural-backdoorhausa_ajami_ocrafrispeech-hausa
Dataset Card for "afrispeech-hausa"
More Information needed
W_hausa_v5W_hausa_v3
Cleaned Hausa Speech Dataset v3
A cleaned and processed Hausa speech dataset built from multiple open-source Hugging Face datasets.
Dataset Description
This dataset contains cleaned, normalized, and deduplicated Hausa speech audio with aligned transcriptions. All audio is:
Sample rate: 16,000 Hz (mono)
Format: FLAC (lossless, embedded in Parquet)
Duration range: 1–30 seconds per clip
Loudness normalized: -20 dBFS RMS
VAD trimmed: Non-speech segments removed with… See the full description on the dataset page: https://huggingface.co/datasets/suleiman2003/W_hausa_v3.W_hausa_v7
Unified Hausa Speech Dataset v5
Dataset Description
A large-scale, cleaned, deduplicated, and quality-filtered Hausa speech dataset compiled from multiple open-source collections. Designed for Text-to-Speech (TTS) and Automatic Speech Recognition (ASR) research.
All audio is 16 kHz mono FLAC, silence-trimmed, loudness-normalized to -20 dBFS, and sorted by speaker_id so that all clips from the same speaker appear consecutively.
Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/suleiman2003/W_hausa_v7.hausa_2_eng_2
Dataset Card for Common Voice Corpus 16
Dataset Summary
The Common Voice dataset consists of a unique MP3 and corresponding text file.
Many of the 30328 recorded hours in the dataset also include demographic metadata like age, sex, and accent
that can help improve the accuracy of speech recognition engines.
The dataset currently consists of 19673 validated hours in 120 languages, but more voices and languages are always added.
Take a look at the Languages page to… See the full description on the dataset page: https://huggingface.co/datasets/eldad-akhaumere/hausa_2_eng_2.W_hausa_v6hausa_response_gemma_dftW_hausa_v4W_hausa_v2hausa-audio-resampled9jalingo-reviewed-hausa-batch-0hausa
African Voices Dataset
Multi-speaker voice dataset for African languages.
Hausa
Audio + transcript pairs organized by speaker, with age_group, domain, and gender metadata.
unified-hausa-speech
Unified Hausa Speech Dataset v5
Dataset Description
A large-scale, cleaned, deduplicated, and quality-filtered Hausa speech dataset compiled from 6 open-source collections. Designed for Text-to-Speech (TTS) and Automatic Speech Recognition (ASR) research on one of Africa's most widely spoken languages.
Hausa (ISO 639-1: ha) is a Chadic language spoken by over 80 million people across West and Central Africa — primarily in Nigeria and Niger, and as a trade language… See the full description on the dataset page: https://huggingface.co/datasets/suleiman2003/unified-hausa-speech.hausa_voa_topics
Dataset Card for Hausa VOA News Topic Classification dataset (hausa_voa_topics)
Dataset Summary
A news headline topic classification dataset, similar to AG-news, for Hausa. The news headlines were collected from VOA Hausa.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
Hausa (ISO 639-1: ha)
Dataset Structure
Data Instances
An instance consists of a news title sentence and the corresponding topic label.… See the full description on the dataset page: https://huggingface.co/datasets/UdS-LSV/hausa_voa_topics.NaijaSenti-TwitterNaijaSenti is the first large-scale human-annotated Twitter sentiment dataset for the four most widely spoken languages in Nigeria — Hausa, Igbo, Nigerian-Pidgin, and Yorùbá — consisting of around 30,000 annotated tweets per language, including a significant fraction of code-mixed tweets.hausa_common_voiceThis dataset is from the common voice corpus 7.0 using the Hausa dataset
hausa_voa_nerThe Hausa VOA NER dataset is a labeled dataset for named entity recognition in Hausa. The texts were obtained from
Hausa Voice of America News articles https://www.voahausa.com/ . We concentrate on
four types of named entities: persons [PER], locations [LOC], organizations [ORG], and dates & time [DATE].
The Hausa VOA NER data files contain 2 columns separated by a tab ('\t'). Each word has been put on a separate line and
there is an empty line after each sentences i.e the CoNLL format. The first item on each line is a word, the second
is the named entity tag. The named entity tags have the format I-TYPE which means that the word is inside a phrase
of type TYPE. For every multi-word expression like 'New York', the first word gets a tag B-TYPE and the subsequent words
have tags I-TYPE, a word with tag O is not part of a phrase. The dataset is in the BIO tagging scheme.
For more details, see https://www.aclweb.org/anthology/2020.emnlp-main.204/hausa_dataset_encoded_repo20to30suda-hausa-tts
Suda Hausa TTS Dataset
High-quality Hausa speech dataset prepared for fine-tuning XTTS-v2 as part of the Suda project — an open-source Hausa voice AI.
Dataset Details
Property
Value
Language
Hausa (ha)
Clips
4,960
Duration
~10 hours
Sample rate
22,050 Hz mono WAV
Format
LJSpeech (metadata.csv + wavs/)
Speaker
Single speaker (male, Nigerian Hausa)
Source
BibleTTS Hausa
License
CC-BY-SA 4.0
Source
Derived from BibleTTS Hausa by… See the full description on the dataset page: https://huggingface.co/datasets/Alkamal01/suda-hausa-tts.s2tt-hausa-englishHausa-Synthetic-ASR-Dataset-XTTSSynthetic Hausa ASR dataset generated using a fine-tuned version of the XTTS-v2 model.
Sample rate: 24kHz.
Total duration: 574 hours.
9jalingo-reviewed-hausa
