hau
Qwen3.8-27B-Uncensored-HauhauCS-Aggressive-MTP-GGUFGemma-4-E4B-Uncensored-HauhauCS-AggressiveQwen3.6-35B-A3B-Uncensored-HauhauCS-AggressiveQwen3.5-9B-Uncensored-HauhauCS-AggressiveGemma4-12B-QAT-Uncensored-HauhauCS-BalancedGemma4-26B-A4B-QAT-Uncensored-HauhauCS-Balanced-MTPQwen3.5-122B-A10B-Uncensored-HauhauCS-AggressiveQwen3.5-4B-Uncensored-HauhauCS-Aggressive
Datasets
All datasets matching “hau”Hausa
Hausa Ajami OCR Dataset
Ce dataset contient des paires image/transcription de manuscrits haoussa en écriture ajami (écriture arabe adaptée au haoussa).
Contenu
Chaque ligne du fichier data/train/metadata.jsonl correspond à une ligne de texte ajami segmentée, avec :
file_name : nom du fichier image correspondant (image de la ligne, recadrée)
transcript : translittération en écriture latine de la ligne
source : identifiant du manuscrit d'origine (voir tableau… See the full description on the dataset page: https://huggingface.co/datasets/IntelligenceResearchLab/Hausa.hausa_dataset_encodedrouting_analysis-marco_nano-hau_Latn-checkpoints
marco_nano_base hau_Latn files
This dataset preserves the original relative paths of every regular file under
the ten marco_nano_base directories whose basename contains the literal
hau_Latn, plus the two matching zero-byte lock files beside them.
The scope intentionally includes hidden work data, preserved backups,
failed_edquot data, and all other regular files found inside those selected
trees. No .gitignore file or checkpoint-name exclusion rule is consulted.
The source… See the full description on the dataset page: https://huggingface.co/datasets/lylybig/routing_analysis-marco_nano-hau_Latn-checkpoints.bible_tts_hausa
Dataset Card for BibleTTS Hausa
Dataset Summary
BibleTTS is a large high-quality open Text-to-Speech dataset with up to 80 hours of single speaker, studio quality 48kHz recordings.
This is a Hausa part of the dataset. Aligned hours: 86.6, aligned verses: 40,603.
Languages
Hausa
Dataset Structure
Data Fields
audio: audio path
sentence: transcription of the audio
locale: always set to ha
book: 3-char book encoding
verse: verse id… See the full description on the dataset page: https://huggingface.co/datasets/vpetukhov/bible_tts_hausa.African_voices_hausa
🇳🇬 WaZoBiaSpeech: 1,000+ Hour Hausa (hau) Corpus
Version: 30 Nov 2025
NOTE: This dataset is subject to regular Updates, corrections, and expansions. Please check this repository regularly for the latest release.
🌍 Dataset Overview
WaZoBiaSpeech is a large-scale, high-quality, fully transcribed speech dataset for Hausa (hau). This corpus is designed to accelerate the development of speech technology in African contexts, promoting linguistic diversity and… See the full description on the dataset page: https://huggingface.co/datasets/Africanvoice/African_voices_hausa.hausa_response_gemma
