datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
somali-asr-synthetic-youtube
Somali ASR Synthetic YouTube Dataset
A Somali-language speech dataset derived from YouTube audio, intended for training and evaluating automatic speech recognition (ASR) and speech-to-text (STT) models. Transcriptions were generated synthetically (silver-standard) via ASR bootstrapping.
Dataset Summary
Split
Samples
train
~4,393
validation
200
test
100
Total
~4,693
Language: Somali (so)
Audio format: WAV, 16 kHz, mono, 16-bit PCM
Total… See the full description on the dataset page: https://huggingface.co/datasets/IbrahimDayax/somali-asr-synthetic-youtube.somali-news-datasetsomali-stem-dataset
Somali STEM Terminology Dataset (Sample)
Overview
A bilingual English–Somali terminology dataset covering 5 STEM domains.
This is a 100-row public sample. The full dataset contains 3,000+ terms.
Somali is spoken by 20+ million people but is critically under-represented in AI training data. This is the only known structured Somali STEM lexicon in machine-readable format.
Fields
Field
Description
English
Scientific term in English
Somali
Somali… See the full description on the dataset page: https://huggingface.co/datasets/planwise-data/somali-stem-dataset.hausa-somali_sentence-pairs
Hausa-Somali_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Hausa-Somali_Sentence-Pairs
Number of Rows: 530060
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/hausa-somali_sentence-pairs.oromo-somali_sentence-pairs
Oromo-Somali_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Oromo-Somali_Sentence-Pairs
Number of Rows: 88097
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/oromo-somali_sentence-pairs.somali-tswana_sentence-pairs
Somali-Tswana_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Somali-Tswana_Sentence-Pairs
Number of Rows: 197114
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/somali-tswana_sentence-pairs.rundi-somali_sentence-pairs
Rundi-Somali_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Rundi-Somali_Sentence-Pairs
Number of Rows: 210469
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/rundi-somali_sentence-pairs.somali-tumbuka_sentence-pairs
Somali-Tumbuka_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Somali-Tumbuka_Sentence-Pairs
Number of Rows: 179589
Number… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/somali-tumbuka_sentence-pairs.lingala-somali_sentence-pairs
Lingala-Somali_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Lingala-Somali_Sentence-Pairs
Number of Rows: 138569
Number… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/lingala-somali_sentence-pairs.somali-zulu_sentence-pairs
Somali-Zulu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Somali-Zulu_Sentence-Pairs
Number of Rows: 605556
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/somali-zulu_sentence-pairs.somali-tigrinya_sentence-pairs
Somali-Tigrinya_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Somali-Tigrinya_Sentence-Pairs
Number of Rows: 169620
Number… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/somali-tigrinya_sentence-pairs.somali-swahili_sentence-pairs
Somali-Swahili_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Somali-Swahili_Sentence-Pairs
Number of Rows: 630275
Number… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/somali-swahili_sentence-pairs.kongo-somali_sentence-pairs
Kongo-Somali_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Kongo-Somali_Sentence-Pairs
Number of Rows: 96580
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/kongo-somali_sentence-pairs.bambara-somali_sentence-pairs
Bambara-Somali_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Bambara-Somali_Sentence-Pairs
Number of Rows: 71330
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/bambara-somali_sentence-pairs.akan-somali_sentence-pairs
Akan-Somali_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Akan-Somali_Sentence-Pairs
Number of Rows: 78264
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/akan-somali_sentence-pairs.afrikaans-somali_sentence-pairs
Afrikaans-Somali_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Afrikaans-Somali_Sentence-Pairs
Number of Rows: 1432536… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/afrikaans-somali_sentence-pairs.somali-umbundu_sentence-pairs
Somali-Umbundu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Somali-Umbundu_Sentence-Pairs
Number of Rows: 96450
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/somali-umbundu_sentence-pairs.dyula-somali_sentence-pairs
Dyula-Somali_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Dyula-Somali_Sentence-Pairs
Number of Rows: 112864
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/dyula-somali_sentence-pairs.somali-yoruba_sentence-pairs
Somali-Yoruba_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Somali-Yoruba_Sentence-Pairs
Number of Rows: 378911
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/somali-yoruba_sentence-pairs.somali-tts
Somali Text-to-Speech Dataset
This dataset is maintained by the SomaliDatasets Organization.
Goal
Collect 1,000,000 high-quality Somali speech recordings.
Repository Structure
audio/
metadata.csv
README.md
Contributions are collected through the CaawiyeAI platform.
somali-tsonga_sentence-pairs
Somali-Tsonga_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Somali-Tsonga_Sentence-Pairs
Number of Rows: 196991
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/somali-tsonga_sentence-pairs.kikuyu-somali_sentence-pairs
Kikuyu-Somali_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Kikuyu-Somali_Sentence-Pairs
Number of Rows: 93562
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/kikuyu-somali_sentence-pairs.ewe-somali_sentence-pairs
Ewe-Somali_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Ewe-Somali_Sentence-Pairs
Number of Rows: 174329
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/ewe-somali_sentence-pairs.pedi-somali_sentence-pairs
Pedi-Somali_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Pedi-Somali_Sentence-Pairs
Number of Rows: 150443
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/pedi-somali_sentence-pairs.amharic-somali_sentence-pairs
Amharic-Somali_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Amharic-Somali_Sentence-Pairs
Number of Rows: 516996
Number… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/amharic-somali_sentence-pairs.kinyarwanda-somali_sentence-pairs
Kinyarwanda-Somali_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Kinyarwanda-Somali_Sentence-Pairs
Number of Rows: 268329… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/kinyarwanda-somali_sentence-pairs.kamba-somali_sentence-pairs
Kamba-Somali_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Kamba-Somali_Sentence-Pairs
Number of Rows: 87143
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/kamba-somali_sentence-pairs.ganda-somali_sentence-pairs
Ganda-Somali_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Ganda-Somali_Sentence-Pairs
Number of Rows: 139229
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/ganda-somali_sentence-pairs.dinka-somali_sentence-pairs
Dinka-Somali_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Dinka-Somali_Sentence-Pairs
Number of Rows: 47951
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/dinka-somali_sentence-pairs.bemba-somali_sentence-pairs
Bemba-Somali_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Bemba-Somali_Sentence-Pairs
Number of Rows: 125821
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/bemba-somali_sentence-pairs.
