datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hausa_common_voiceThis dataset is from the common voice corpus 7.0 using the Hausa dataset
hausa_aug_lex
title: Lexicon Dataset for the Hausa Language
Dataset with English translation
license: cc-by-nd-4.0
hausa-ajami-blindspot-evalhausa_newsclass_topicfulah-hausa_sentence-pairs
Fulah-Hausa_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Fulah-Hausa_Sentence-Pairs
Number of Rows: 269337
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/fulah-hausa_sentence-pairs.hausa-shona_sentence-pairs
Hausa-Shona_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Hausa-Shona_Sentence-Pairs
Number of Rows: 829704
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/hausa-shona_sentence-pairs.hausa-somali_sentence-pairs
Hausa-Somali_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Hausa-Somali_Sentence-Pairs
Number of Rows: 530060
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/hausa-somali_sentence-pairs.hausa-rundi_sentence-pairs
Hausa-Rundi_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Hausa-Rundi_Sentence-Pairs
Number of Rows: 261557
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/hausa-rundi_sentence-pairs.Roleplay-Hausa
RolePlay-Hausa
Roleplay-Hausa Dataset is a dataset for roleplaying in the Hausa language for Large Language Model.
The base dataset is GPTeacher role play dataset by teknium 1, which can be found under this link, released under MIT License. The dataset is then translated into respective languages. The translation process is powered by Google Translate, using cloud translation API.
For more information and other language datasets for roleplay, see this github repo.
For contact… See the full description on the dataset page: https://huggingface.co/datasets/jojo-ai-mst/Roleplay-Hausa.hausa-xhosa_sentence-pairs
Hausa-Xhosa_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Hausa-Xhosa_Sentence-Pairs
Number of Rows: 627718
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/hausa-xhosa_sentence-pairs.fon-hausa_sentence-pairs
Fon-Hausa_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Fon-Hausa_Sentence-Pairs
Number of Rows: 103601
Number of Columns:… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/fon-hausa_sentence-pairs.hausa-nuer_sentence-pairs
Hausa-Nuer_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Hausa-Nuer_Sentence-Pairs
Number of Rows: 47684
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/hausa-nuer_sentence-pairs.HausaHate
Evaluation Benchmark for Hausa Hate Speech Detection
We introduce the first expert annotated corpus of Facebook comments for Hausa hate speech detection.
The corpus titled HausaHate comprises 2,000 comments extracted from Western African Facebook pages and
manually annotated by three Hausa native speakers, who are also NLP experts.
The corpus was annotated using two different layers. We first labeled each comment according to a
binary classification: offensive versus… See the full description on the dataset page: https://huggingface.co/datasets/franciellevargas/HausaHate.hausa-swahili_sentence-pairs
Hausa-Swahili_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Hausa-Swahili_Sentence-Pairs
Number of Rows: 956078
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/hausa-swahili_sentence-pairs.hausa-kikuyu_sentence-pairs
Hausa-Kikuyu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Hausa-Kikuyu_Sentence-Pairs
Number of Rows: 114672
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/hausa-kikuyu_sentence-pairs.hausa-twi_sentence-pairs
Hausa-Twi_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Hausa-Twi_Sentence-Pairs
Number of Rows: 212497
Number of Columns:… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/hausa-twi_sentence-pairs.amharic-hausa_sentence-pairs
Amharic-Hausa_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Amharic-Hausa_Sentence-Pairs
Number of Rows: 751955
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/amharic-hausa_sentence-pairs.Hausa-Speech-Dataset
🎧 Hausa Speech Dataset
The Hausa Speech Dataset is a structured and high-quality speech audio dataset developed to support modern AI systems requiring diverse audio data and reliable voice data. It includes 174 hours of recordings distributed across 733 files, available in MP3 and WAV formats, with a total size of 362 MB. This carefully curated audio dataset ensures balanced speaker representation, with 52% female and 48% male speakers, and a broad age distribution from 18 to 50+… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/Hausa-Speech-Dataset.Hausa-EnglishCodeswitchbemba-hausa_sentence-pairs
Bemba-Hausa_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Bemba-Hausa_Sentence-Pairs
Number of Rows: 180646
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/bemba-hausa_sentence-pairs.hausa-wolof_sentence-pairs
Hausa-Wolof_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Hausa-Wolof_Sentence-Pairs
Number of Rows: 77051
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/hausa-wolof_sentence-pairs.hausa-tumbuka_sentence-pairs
Hausa-Tumbuka_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Hausa-Tumbuka_Sentence-Pairs
Number of Rows: 260765
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/hausa-tumbuka_sentence-pairs.hausa-oromo_sentence-pairs
Hausa-Oromo_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Hausa-Oromo_Sentence-Pairs
Number of Rows: 118255
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/hausa-oromo_sentence-pairs.hausa-kimbundu_sentence-pairs
Hausa-Kimbundu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Hausa-Kimbundu_Sentence-Pairs
Number of Rows: 95341
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/hausa-kimbundu_sentence-pairs.hausa-igbo_sentence-pairs
Hausa-Igbo_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Hausa-Igbo_Sentence-Pairs
Number of Rows: 713539
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/hausa-igbo_sentence-pairs.hausa-lingala_sentence-pairs
Hausa-Lingala_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Hausa-Lingala_Sentence-Pairs
Number of Rows: 172587
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/hausa-lingala_sentence-pairs.ewe-hausa_sentence-pairs
Ewe-Hausa_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Ewe-Hausa_Sentence-Pairs
Number of Rows: 204935
Number of Columns:… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/ewe-hausa_sentence-pairs.akan-hausa_sentence-pairs
Akan-Hausa_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Akan-Hausa_Sentence-Pairs
Number of Rows: 98881
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/akan-hausa_sentence-pairs.hausa-zulu_sentence-pairs
Hausa-Zulu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Hausa-Zulu_Sentence-Pairs
Number of Rows: 964410
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/hausa-zulu_sentence-pairs.hausa-yoruba_sentence-pairs
Hausa-Yoruba_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Hausa-Yoruba_Sentence-Pairs
Number of Rows: 795566
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/hausa-yoruba_sentence-pairs.
