datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
umbraThe SAR data are obtained from the UMBRA Open Data Program (https://umbra.space/open-data/).
The dataset has been created combining the WorldCover by ESA (https://esa-worldcover.org/en) with the UMBRA images.
BIOME information has been extracted from the RESOLVE biome dataset (https://ecoregions.appspot.com/).
Reverse geo-coding with OSM nominatim (https://nominatim.openstreetmap.org).
(more details asap)
Authors: Federico Ricciuti, Federico Serva, Alessandro Sebastianelli
License: Same as… See the full description on the dataset page: https://huggingface.co/datasets/fedric95/umbra.take_umbrella_out_of_umbrella_stand_25_08_03_parquetCCTV_Weapon_Detection_Rifles_vs_Umbrellas
Synthetic Weapon Detection: Rifles vs. Umbrellas (CCTV)
Overview
This is an open-source synthetic dataset designed to solve the single biggest problem in Weapon Detection AI: False Positives.
Standard models often confuse common handheld objects—like closed umbrellas, tripods, or tools—with firearms. This dataset focuses specifically on Hard Negative Mining, containing a balanced split of lethal weapons (Rifles) and visual lookalikes (Umbrellas) seen from realistic CCTV… See the full description on the dataset page: https://huggingface.co/datasets/Simuletic/CCTV_Weapon_Detection_Rifles_vs_Umbrellas.PubMedClaimNMTMD
NMTMD (NMT-Melinda-Dataset)
Official repository for the Opensource Text dataset for NMT for local languages in West Africa (EWE Corpus) and implement the Yodi model afterward.
Note: This repository will evolve into the official repository for the Yodi model, once the necessary data is gathered.
Objective
• Develop a Machine Translation Text and Speech Dataset NMT for local languages in West Africa (EWE Corpus)
Key Results
-> Develop &… See the full description on the dataset page: https://huggingface.co/datasets/Umbaji/NMTMD.NMTMD
NMTMD (NMT-Melinda-Dataset)
Official repository for the Opensource Text dataset for NMT for local languages in West Africa (EWE Corpus) and implement the Yodi model afterward.
Note: This repository will evolve into the official repository for the Yodi model, once the necessary data is gathered.
Objective
• Develop a Machine Translation Text and Speech Dataset NMT for local languages in West Africa (EWE Corpus)
Key Results
-> Develop &… See the full description on the dataset page: https://huggingface.co/datasets/Umbaji001/NMTMD.golden_sample_datasetumbundu-sentiments-corpus
Umbundu Sentiment Corpus
Dataset Description
This dataset contains sentiment-labeled text data in Umbundu for binary sentiment classification (Positive/Negative). Sentiments are extracted and processed from the English meanings of the sentences using DistilBERT for sentiment classification. The dataset is part of a larger collection of African language sentiment analysis resources.
Dataset Statistics
Total samples: 83,350
Positive sentiment: 48940 (58.7%)… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/umbundu-sentiments-corpus.umbundu-emotions-corpus
Umbundu Emotion Analysis Corpus
Dataset Description
This dataset contains emotion-labeled text data in Umbundu for emotion classification (joy, sadness, anger, fear, surprise, disgust, neutral). Emotions were extracted and processed from the English meanings of the sentences using the model j-hartmann/emotion-english-distilroberta-base. The dataset is part of a larger collection of African language emotion analysis resources.
Dataset Statistics
Total samples:… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/umbundu-emotions-corpus.31k_umber_kiteakan-umbundu_sentence-pairs
Akan-Umbundu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Akan-Umbundu_Sentence-Pairs
Number of Rows: 22651
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/akan-umbundu_sentence-pairs.wan_closing_umbrellaThis dataset contains videos generated using Wan 2.1 T2V 14B.
wan_opening_umbrellaThis dataset contains videos generated using Wan 2.1 T2V 14B.
swahili-umbundu_sentence-pairs
Swahili-Umbundu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Swahili-Umbundu_Sentence-Pairs
Number of Rows: 276674
Number… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/swahili-umbundu_sentence-pairs.english-umbundu_sentence-pairs_mt560
English-Umbundu Parallel Dataset
This dataset contains parallel sentences in English and Umbundu (Angola).
Dataset Information
Language Pair: English ↔ Umbundu
Language Code: umb
Country: Angola
Original Source: OPUS MT560 Dataset
Dataset Structure
The dataset contains parallel sentences that can be used for:
Machine translation training
Cross-lingual NLP tasks
Language model fine-tuning
Citation
If you use this dataset, please cite the citation… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/english-umbundu_sentence-pairs_mt560.diophantine-equations-dataset
Diophantine Equations Dataset
A curated dataset of 1,434 Diophantine equation problems with complete step-by-step solutions.
Dataset Statistics
Total examples: 1,434
Train: 1,218 | Validation: 144 | Test: 72
Usage
from datasets import load_dataset
dataset = load_dataset("Umbaji/diophantine-equations-dataset")
Features
Complete solutions with reasoningDiverse problem typesLaTeX notation preserved
Citation
Bibtex… See the full description on the dataset page: https://huggingface.co/datasets/Umbaji/diophantine-equations-dataset.umbrela-indo-irPubmedFact1klicense: apache-2.0
pubid: Directly inherited from the input.
claim: Reformatted from the original "question" field to express a clear claim statement.
context: Contains:
contexts: A subset of the original context information.
labels: Corresponding labels related to the context.
final_decision: Converted from a textual decision to a numerical value:
"yes" (case insensitive) → 1
"no" → 0
"maybe" → 2
shona-umbundu_sentence-pairs
Shona-Umbundu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Shona-Umbundu_Sentence-Pairs
Number of Rows: 153467
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/shona-umbundu_sentence-pairs.tsonga-umbundu_sentence-pairs
Tsonga-Umbundu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Tsonga-Umbundu_Sentence-Pairs
Number of Rows: 132324
Number… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/tsonga-umbundu_sentence-pairs.kongo-umbundu_sentence-pairs
Kongo-Umbundu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Kongo-Umbundu_Sentence-Pairs
Number of Rows: 58711
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/kongo-umbundu_sentence-pairs.scifact-open
Data Stats
206 claims
500k distractors
Data Structure
Test
claim
evidence: GT evidence
evidence_id: GT evidence id
label: GT label
evidences: list of all evidences
evidence_ids: list of all evidence ids
labels: list of all labels
Distractors
evidence
evidence_id
Process Code
import pandas as pd
from datasets import Dataset
claims = pd.read_csv("./scifact_open_retriever_test.csv")
claims.head()
docs =… See the full description on the dataset page: https://huggingface.co/datasets/umbc-scify/scifact-open.Casual-Autopsy__L3-Umbral-Mind-RP-v2.0-8B-details
Dataset Card for Evaluation run of Casual-Autopsy/L3-Umbral-Mind-RP-v2.0-8B
Dataset automatically created during the evaluation run of model Casual-Autopsy/L3-Umbral-Mind-RP-v2.0-8B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Casual-Autopsy__L3-Umbral-Mind-RP-v2.0-8B-details.oromo-umbundu_sentence-pairs
Oromo-Umbundu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Oromo-Umbundu_Sentence-Pairs
Number of Rows: 48783
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/oromo-umbundu_sentence-pairs.dinka-umbundu_sentence-pairs
Dinka-Umbundu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Dinka-Umbundu_Sentence-Pairs
Number of Rows: 11337
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/dinka-umbundu_sentence-pairs.french-umbundu_sentence-pairssomali-umbundu_sentence-pairs
Somali-Umbundu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Somali-Umbundu_Sentence-Pairs
Number of Rows: 96450
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/somali-umbundu_sentence-pairs.nuer-umbundu_sentence-pairs
Nuer-Umbundu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Nuer-Umbundu_Sentence-Pairs
Number of Rows: 9607
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/nuer-umbundu_sentence-pairs.kinyarwanda-umbundu_sentence-pairs
Kinyarwanda-Umbundu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Kinyarwanda-Umbundu_Sentence-Pairs
Number of Rows:… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/kinyarwanda-umbundu_sentence-pairs.kikuyu-umbundu_sentence-pairs
Kikuyu-Umbundu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Kikuyu-Umbundu_Sentence-Pairs
Number of Rows: 29453
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/kikuyu-umbundu_sentence-pairs.
