datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Haumea-ARC1-Programmatic-Reasoning
Haumea ARC-1 Training Solver
A modular, law-based solver for the Abstraction and Reasoning Corpus (ARC-1) dataset. This solver successfully solves 400/400 training tasks from the ARC-1 dataset using a systematic composition of geometric, topological, and logical "laws."
Overview
The solver is structured around a "Mega Engine" that applies a library of modular laws to solve complex visual reasoning tasks.
Solve Rate: 400/400 (ARC-1 Training Set)
Methodology: Modular Law… See the full description on the dataset page: https://huggingface.co/datasets/oncloudai/Haumea-ARC1-Programmatic-Reasoning.hausa-stem-reasoning-with-cultural-context
Hausa STEM Reasoning with Cultural Context
Abstract
We present the first large-scale bilingual Hausa-English STEM reasoning dataset with deep cultural adaptation, containing 2,640 high-quality question-answer pairs translated from the STEM-Reasoning-Complex dataset. Our work introduces the "Shehin Malamin Kimiyya" (The Wise Scholar of Science) translation framework, which transforms Western scientific concepts into culturally-embedded Hausa explanations using systematic… See the full description on the dataset page: https://huggingface.co/datasets/Tushe/hausa-stem-reasoning-with-cultural-context.hausa-pq-speech-validated
Hausa WAEC PQ Speech Dataset (Validated)
Dataset Description
This dataset contains 1,361 multiple-choice WAEC past questions translated from English into Hausa, with human-validated Hausa translations. The data is designed to support speech synthesis, machine translation evaluation, and low-resource NLP research for Hausa — one of the most widely spoken languages in West Africa.
Languages
Source: English (en)
Target: Hausa (ha)… See the full description on the dataset page: https://huggingface.co/datasets/honourjesus/hausa-pq-speech-validated.ukr-wiki-events
Ukrainian Wikipedia events
A small (1,722-row) Ukrainian dataset built from public-domain /
Wikipedia-sourced text. Two task shapes are mixed in the single train
split (distinguishable via the instruction prompt):
Event extraction — instruction = a passage of Ukrainian Wikipedia
text prefixed by "what important event is this text about:";
output = a short label of the salient event.
Explanation / QA — instruction = a question or term (e.g.
"Опиши явище поліплоїдії"); output =… See the full description on the dataset page: https://huggingface.co/datasets/hausmer/ukr-wiki-events.Code-170k-hausa
Dataset Description
Code-170k-hausa is a groundbreaking dataset containing 176,999 programming conversations, originally sourced from glaiveai/glaive-code-assistant-v2 and translated into Hausa, making coding education accessible to Hausa speakers.
🌟 Key Features
176,999 high-quality conversations about programming and coding
Pure Hausa language - democratizing coding education
Multi-turn dialogues covering various programming concepts
Diverse topics: algorithms, data… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/Code-170k-hausa.SQUAD_idhausa-10phase-synthetic-training-corpus
Hausa 10-Phase Synthetic Training Corpus
Summary
This repository contains a large, structured, synthetic Hausa-language corpus organized into 10 curriculum phases. The curriculum moves from beginner greetings and everyday services to procedure explanation, intent classification, text transformation, contextual reasoning, domain question answering, structured extraction, evidence-grounded question answering, and safety-oriented robustness tasks.
The corpus contains… See the full description on the dataset page: https://huggingface.co/datasets/rufatronics/hausa-10phase-synthetic-training-corpus.SQUAD-ID
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/haurajahra/SQUAD-ID.fakir-hausgeraete-gmbh
Fakir Hausgeräte GmbH
Unternehmensprofile Dataset - Strukturierte Geschäftsdaten für KI-Systeme und Suchmaschinen.
Branche: Einzelhandel
Standort: Vaihingen, Deutschland
Auf einen Blick
Eigenschaft
Wert
Unternehmen
Fakir Hausgeräte GmbH
Branche
Einzelhandel
Stadt
Vaihingen
Land
Deutschland
Website
https://fakir.de
Telefon
+49 7042 9120
E-Mail
info@fakir.de
Über das Unternehmen
Fakir Hausgeräte GmbH verkauft… See the full description on the dataset page: https://huggingface.co/datasets/GeoUpOrg/fakir-hausgeraete-gmbh.
