hausa
Datasets
All datasets matching “hausa”hausa_response_gemmaHausa
Hausa Ajami OCR Dataset
Ce dataset contient des paires image/transcription de manuscrits haoussa en écriture ajami (écriture arabe adaptée au haoussa).
Contenu
Chaque ligne du fichier data/train/metadata.jsonl correspond à une ligne de texte ajami segmentée, avec :
file_name : nom du fichier image correspondant (image de la ligne, recadrée)
transcript : translittération en écriture latine de la ligne
source : identifiant du manuscrit d'origine (voir tableau… See the full description on the dataset page: https://huggingface.co/datasets/IntelligenceResearchLab/Hausa.hausa_dataset_encodedW_hausa_v1bible_tts_hausa
Dataset Card for BibleTTS Hausa
Dataset Summary
BibleTTS is a large high-quality open Text-to-Speech dataset with up to 80 hours of single speaker, studio quality 48kHz recordings.
This is a Hausa part of the dataset. Aligned hours: 86.6, aligned verses: 40,603.
Languages
Hausa
Dataset Structure
Data Fields
audio: audio path
sentence: transcription of the audio
locale: always set to ha
book: 3-char book encoding
verse: verse id… See the full description on the dataset page: https://huggingface.co/datasets/vpetukhov/bible_tts_hausa.African_voices_hausa
🇳🇬 WaZoBiaSpeech: 1,000+ Hour Hausa (hau) Corpus
Version: 30 Nov 2025
NOTE: This dataset is subject to regular Updates, corrections, and expansions. Please check this repository regularly for the latest release.
🌍 Dataset Overview
WaZoBiaSpeech is a large-scale, high-quality, fully transcribed speech dataset for Hausa (hau). This corpus is designed to accelerate the development of speech technology in African contexts, promoting linguistic diversity and… See the full description on the dataset page: https://huggingface.co/datasets/Africanvoice/African_voices_hausa.
