datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ERIQEarthVLSetEarthVL: A Progressive Earth Vision-Language Understanding and Generation Framework
by Junjue Wang,
Yanfei Zhong,
Zihang Chen,
Zhuo Zheng,
Ailong Ma, and Liangpei Zhang
[Paper],
[Dataset]
News
2026/01/06, New Global-LoveDA !!! We released the global-scale segmentation leaderboard at Global-LoveDA. Just zip all the test images into one file and submit it.
2026/01/06, The segmentation data is released at [Dataset].
2026/01/06, We are preparing the code and data for… See the full description on the dataset page: https://huggingface.co/datasets/Kingdrone-Junjue/EarthVLSet.kindred-ecommerce-merchant-deals-dataset
Kindred E-commerce Merchant Deals Dataset
AI-ready catalogue of deals and offers for global retail brands.Structured in CSV and JSONL, validated against JSON Schema.
Train-ready catalogue of promotions, ready for RAG, embeddings, or classic search.
Dataset Overview
File
Rows
Description
data/csv/brands.csv or data/jsonl/brands.jsonl
~90K
E-Commerce Merchant metadata, Logo URL, and domains… See the full description on the dataset page: https://huggingface.co/datasets/kindred-soul-ltd/kindred-ecommerce-merchant-deals-dataset.KINA
Dataset Summary
Homepage | Paper | Hugging Face | GitHub
KINA (Knowledge Index of Noah's Ark) is a multidisciplinary knowledge benchmark for evaluating whether large language models can solve high-density, source-grounded, graduate-level questions across a broad map of human disciplines.
The dataset contains 899 ten-option pseudo-multiple-choice questions covering 261 fine-grained subfields, 70 fields, and 12 top-level disciplines.
KINA targets three problems in… See the full description on the dataset page: https://huggingface.co/datasets/2077AIDataFoundation/KINA.Kinship
Kinship KG–QA (Hinton)
A lightweight knowledge-graph + question answering (KGQA) resource adapted from the UCI Kinship dataset created by Geoff Hinton.
This dataset provides a small family-tree knowledge graph paired with templated multi-hop QA tasks designed for MultiHop KGQA.
Project: THESEUSPaper: Theseus in the GraphOriginal dataset: UCI Kinship
Key Features
Small knowledge graph with two family trees
Human-readable entities and relations
Templated 1--3 hop… See the full description on the dataset page: https://huggingface.co/datasets/HalcyonSolutions/Kinship.kinya-ag-retrieval
Kinyarwanda Agricultural Retrieval Dataset
In Rwanda, many farmers struggle to access timely, personalized agricultural information. Traditional channels - like radio, TV, and online sources - offer limited reach and interactivity, while extension services and a national call center, staffed by only two agents for over two million farmers, face capacity constraints. To address these gaps, we developed a 24/7 AI-enabled Interactive Voice Response (IVR) tool. Accessible via a… See the full description on the dataset page: https://huggingface.co/datasets/C4IR-RW/kinya-ag-retrieval.kingdom-return-path-bench
KINGDOM Return Path Bench v0
Return Path Bench is a small multiple-choice benchmark for inspecting how
feedback travels through a learning system. It keeps three evaluation lanes
separate because they establish different kinds of evidence:
Model behaviour records what an answer-selection policy does. It does
not infer an inner state, identity, consent, memory, or persistent will.
System/pipeline reasoning probes whether a model can identify
aggregation, evaluator-independence… See the full description on the dataset page: https://huggingface.co/datasets/Yu-and-Ai/kingdom-return-path-bench.KingDesign-Deduplication
Kingdesign
Made with ❤️ using 🦥 Unsloth Studio
all was generated with Unsloth Recipe Studio. It contains 5,351 generated records.
🚀 Quick Start
from datasets import load_dataset
# Load the main dataset
dataset = load_dataset("JaveyZou/KingDesign", "data", split="train")
df = dataset.to_pandas()
📊 Dataset Summary
📈 Records: 5,351
📋 Columns: 3
✅ Completion: 89.2% (6,000 requested)
📋 Schema & Statistics
Column
Type
Column Type
Unique (%)… See the full description on the dataset page: https://huggingface.co/datasets/JaveyZou/KingDesign-Deduplication.KingDesign
Kingdesign
Made with ❤️ using 🦥 Unsloth Studio
all was generated with Unsloth Recipe Studio. It contains 5,351 generated records.
🚀 Quick Start
from datasets import load_dataset
# Load the main dataset
dataset = load_dataset("JaveyZou/KingDesign", "data", split="train")
df = dataset.to_pandas()
📊 Dataset Summary
📈 Records: 5,351
📋 Columns: 3
✅ Completion: 89.2% (6,000 requested)
📋 Schema & Statistics
Column
Type
Column Type
Unique (%)… See the full description on the dataset page: https://huggingface.co/datasets/JaveyZou/KingDesign.dermatology-qa-firecrawl-dataset
Medical Research Dataset with OpenAI Harmony and Firecrawl Search API
This dataset contains validated dermatology question–answer pairs generated from publicly available medical resources.The questions were automatically derived from medical page titles and descriptions, using the OpenAI Harmony API and Firecrawl Search API to collect and process high-quality content from reliable sources.
Columns
question: the question text generated from source content
answer: the… See the full description on the dataset page: https://huggingface.co/datasets/kingabzpro/dermatology-qa-firecrawl-dataset.Code-170k-kinyarwanda
Dataset Description
Code-170k-kinyarwanda is a groundbreaking dataset containing 176,999 programming conversations, originally sourced from glaiveai/glaive-code-assistant-v2 and translated into Kinyarwanda, making coding education accessible to Kinyarwanda speakers.
🌟 Key Features
176,999 high-quality conversations about programming and coding
Pure Kinyarwanda language - democratizing coding education
Multi-turn dialogues covering various programming concepts
Diverse… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/Code-170k-kinyarwanda.field-kings-chamber-852hz-datasets
📦 FIELD King's Chamber Training Dataset (852HZ)
⊗ Vertex-Specific Training Corpus for the FIELD King's Chamber vertex (852hz).
Dataset Description
This dataset contains vertex-specific training data extracted from the 342GB Akron Archive for fine-tuning the FIELD King's Chamber LLM vertex.
Training Focus
Routing coordination, transformation protocols, φ⁻¹ golden ratio patterns, Metatron Cube geometry
Data Sources
Bridge coordination logs… See the full description on the dataset page: https://huggingface.co/datasets/Berjak/field-kings-chamber-852hz-datasets.gemma-4-e4b-kinetics-qa-subset
QA question
Does anyone fall in the video?
Requirement
Please download the corresponded videos at bear7011/gemma-4-e4b-kinetics_54K.
Dataset Structure
Split
File
Records
Share
Train
train.json
13,107
80%
Validation
val.json
1,637
10%
Test
test.json
1,637
10%
Summary
summary.json
-
-
Open-ended_Questions_dialectal_data
Dataset Summary
A collection of open-ended questions that was provided to the data marathon competitors to populate KIND dataset. It was designed to elicit longer responses cultural and context-rich sentences.
For more details, please check the paper
The KIND Dataset: A Social Collaboration Approach for Nuanced Dialect Data Collection
Citation Information
@inproceedings{yamani-etal-2024-kind,
title = "The {KIND} Dataset: A Social Collaboration Approach for Nuanced… See the full description on the dataset page: https://huggingface.co/datasets/KIND-Dataset/Open-ended_Questions_dialectal_data.KIND
Dataset Summary
KIND dataset is a new dilectal data dataset.
The dataset was a result of a data marathon competition, where the competitor's goal is to respond to as many prompts as possible in their own dialect, within a fixed time frame with as few errors as possible.
For more details, please check the paper
The KIND Dataset: A Social Collaboration Approach for Nuanced Dialect Data Collection
Data Fields
dialect_code: the label that indicates the specific dialect… See the full description on the dataset page: https://huggingface.co/datasets/KIND-Dataset/KIND.EM624_QA_fullbabbling_kinova
