datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
somali-asr-synthetic-youtube
Somali ASR Synthetic YouTube Dataset
A Somali-language speech dataset derived from YouTube audio, intended for training and evaluating automatic speech recognition (ASR) and speech-to-text (STT) models. Transcriptions were generated synthetically (silver-standard) via ASR bootstrapping.
Dataset Summary
Split
Samples
train
~4,393
validation
200
test
100
Total
~4,693
Language: Somali (so)
Audio format: WAV, 16 kHz, mono, 16-bit PCM
Total… See the full description on the dataset page: https://huggingface.co/datasets/IbrahimDayax/somali-asr-synthetic-youtube.somali-news-datasetSOMAJGYAAN
SomajGyaan (সমাজজ্ঞান) - Bangla MCQ Dataset
📊 Dataset Description
SomajGyaan (সমাজজ্ঞান) is a comprehensive Bangla multiple-choice question dataset featuring 4,234 questions across 7 academic categories with ~12,000 unique answer options.
Dataset Summary
Total Questions: 4,234
Unique Answer Options: ~12,000
Answer Diversity: 70.8%
Language: Bangla (Bengali)
Categories: 7 (History, Economics, Geography, Politics, Social Studies, Law… See the full description on the dataset page: https://huggingface.co/datasets/farihashifa/SOMAJGYAAN.somali-stem-dataset
Somali STEM Terminology Dataset (Sample)
Overview
A bilingual English–Somali terminology dataset covering 5 STEM domains.
This is a 100-row public sample. The full dataset contains 3,000+ terms.
Somali is spoken by 20+ million people but is critically under-represented in AI training data. This is the only known structured Somali STEM lexicon in machine-readable format.
Fields
Field
Description
English
Scientific term in English
Somali
Somali… See the full description on the dataset page: https://huggingface.co/datasets/planwise-data/somali-stem-dataset.hausa-somali_sentence-pairs
Hausa-Somali_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Hausa-Somali_Sentence-Pairs
Number of Rows: 530060
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/hausa-somali_sentence-pairs.somali-tswana_sentence-pairs
Somali-Tswana_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Somali-Tswana_Sentence-Pairs
Number of Rows: 197114
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/somali-tswana_sentence-pairs.rundi-somali_sentence-pairs
Rundi-Somali_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Rundi-Somali_Sentence-Pairs
Number of Rows: 210469
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/rundi-somali_sentence-pairs.lingala-somali_sentence-pairs
Lingala-Somali_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Lingala-Somali_Sentence-Pairs
Number of Rows: 138569
Number… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/lingala-somali_sentence-pairs.somali-zulu_sentence-pairs
Somali-Zulu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Somali-Zulu_Sentence-Pairs
Number of Rows: 605556
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/somali-zulu_sentence-pairs.somali-tigrinya_sentence-pairs
Somali-Tigrinya_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Somali-Tigrinya_Sentence-Pairs
Number of Rows: 169620
Number… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/somali-tigrinya_sentence-pairs.somali-swahili_sentence-pairs
Somali-Swahili_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Somali-Swahili_Sentence-Pairs
Number of Rows: 630275
Number… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/somali-swahili_sentence-pairs.oromo-somali_sentence-pairs
Oromo-Somali_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Oromo-Somali_Sentence-Pairs
Number of Rows: 88097
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/oromo-somali_sentence-pairs.somali-tumbuka_sentence-pairs
Somali-Tumbuka_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Somali-Tumbuka_Sentence-Pairs
Number of Rows: 179589
Number… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/somali-tumbuka_sentence-pairs.kongo-somali_sentence-pairs
Kongo-Somali_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Kongo-Somali_Sentence-Pairs
Number of Rows: 96580
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/kongo-somali_sentence-pairs.bambara-somali_sentence-pairs
Bambara-Somali_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Bambara-Somali_Sentence-Pairs
Number of Rows: 71330
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/bambara-somali_sentence-pairs.akan-somali_sentence-pairs
Akan-Somali_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Akan-Somali_Sentence-Pairs
Number of Rows: 78264
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/akan-somali_sentence-pairs.afrikaans-somali_sentence-pairs
Afrikaans-Somali_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Afrikaans-Somali_Sentence-Pairs
Number of Rows: 1432536… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/afrikaans-somali_sentence-pairs.somali-umbundu_sentence-pairs
Somali-Umbundu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Somali-Umbundu_Sentence-Pairs
Number of Rows: 96450
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/somali-umbundu_sentence-pairs.dyula-somali_sentence-pairs
Dyula-Somali_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Dyula-Somali_Sentence-Pairs
Number of Rows: 112864
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/dyula-somali_sentence-pairs.somali-yoruba_sentence-pairs
Somali-Yoruba_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Somali-Yoruba_Sentence-Pairs
Number of Rows: 378911
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/somali-yoruba_sentence-pairs.somali-tsonga_sentence-pairs
Somali-Tsonga_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Somali-Tsonga_Sentence-Pairs
Number of Rows: 196991
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/somali-tsonga_sentence-pairs.kikuyu-somali_sentence-pairs
Kikuyu-Somali_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Kikuyu-Somali_Sentence-Pairs
Number of Rows: 93562
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/kikuyu-somali_sentence-pairs.pedi-somali_sentence-pairs
Pedi-Somali_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Pedi-Somali_Sentence-Pairs
Number of Rows: 150443
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/pedi-somali_sentence-pairs.ewe-somali_sentence-pairs
Ewe-Somali_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Ewe-Somali_Sentence-Pairs
Number of Rows: 174329
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/ewe-somali_sentence-pairs.amharic-somali_sentence-pairs
Amharic-Somali_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Amharic-Somali_Sentence-Pairs
Number of Rows: 516996
Number… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/amharic-somali_sentence-pairs.Identification_of_Somatic_Driver_Mutations_in_Indian_Oral_OSCC
Indian OSCC Somatic Driver Mutation Dataset
This repository contains the processed datasets used in the study: "Identification of Somatic Driver Mutations in Indian Oral Squamous Cell Carcinoma Using XGBoost and Integrative Genomic Features".
Files
train_MAF.csv – Somatic mutation data used for training
gene_variants.csv – Curated cancer driver gene list
ROH.csv – Runs of homozygosity intervals
SBS_Signature_HNSC.csv – Gene-level SBS13 annotation
synthetic_mutations.csv… See the full description on the dataset page: https://huggingface.co/datasets/aarushidas/Identification_of_Somatic_Driver_Mutations_in_Indian_Oral_OSCC.kinyarwanda-somali_sentence-pairs
Kinyarwanda-Somali_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Kinyarwanda-Somali_Sentence-Pairs
Number of Rows: 268329… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/kinyarwanda-somali_sentence-pairs.kamba-somali_sentence-pairs
Kamba-Somali_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Kamba-Somali_Sentence-Pairs
Number of Rows: 87143
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/kamba-somali_sentence-pairs.ganda-somali_sentence-pairs
Ganda-Somali_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Ganda-Somali_Sentence-Pairs
Number of Rows: 139229
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/ganda-somali_sentence-pairs.dinka-somali_sentence-pairs
Dinka-Somali_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Dinka-Somali_Sentence-Pairs
Number of Rows: 47951
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/dinka-somali_sentence-pairs.
