datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
damru-knowledge
🐕 Damru Knowledge
A continuously growing, self-collected question-answer knowledge base that powers Damru AI — a self-learning assistant built for exam preparation and general-purpose help, with a focus on Indian students.
The dataset is harvested and quality-filtered automatically, 24x7, from multiple open sources and a self-evaluating reasoning engine. New rows are appended every hour as parquet shards under data/.
📦 What's inside
Column
Type
Description… See the full description on the dataset page: https://huggingface.co/datasets/Damaru-ai/damru-knowledge.General-Knowledge
Dataset Card for Dataset Name
Dataset Summary
The dataset is a collection of questions and answers themed on general facts and reasoning. The dataset is divided into two features - 'Question' and 'Answer'.
It is meant to be used for training a model to be good at general knowledge and reasoning. This dataset is inspired from the Alpaca dataset, and infact contains a subset of the alpaca dataset in itself.
Distribution
The distribution of the… See the full description on the dataset page: https://huggingface.co/datasets/MuskumPillerum/General-Knowledge.huatuo_knowledge_graph_qa
Dataset Card for Huatuo_knowledge_graph_qa
Dataset Summary
We built this QA dataset based on the medical knowledge map, with a total of 798,444 pieces of data, in which the questions are constructed by means of templates, and the answers are the contents of the entries in the knowledge map.
Dataset Creation
Source Data… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/huatuo_knowledge_graph_qa.cancer-knowledge-base
Cancer Knowledge Base — the open, verified oncology KB for RAG & LLM evaluation
The only open CC-BY-4.0 oncology knowledge base that combines:
110/110 trials cited with PMID + NCT + PubMed/ClinicalTrials.gov URLs, and 32 prognosis
rows linked to verified SEER 2016–2022 references — no LLM-synthetic dataset has this.
A provable 152-question MCQ benchmark — every answer derives from this KB's own structured
data and carries a citation + golden docs, so it is open-book verifiable… See the full description on the dataset page: https://huggingface.co/datasets/ranjithraj/cancer-knowledge-base.specialist-level_medical_knowledge_dataset_sft
specialist-level_medical_knowledge_dataset_sft
Dataset Summary
specialist-level_medical_knowledge_dataset_sft is an integrated collection of augmented SFT data across 13 distinct medical domains, developed by the Healthcare AI Research Institute (HARI) at SNUH.
This dataset is derived and augmented from the Specialized Medical Knowledge Data (전문 의학지식 데이터) provided by AI-Hub.
It focuses exclusively on complex clinical scenarios generated using the "Add Constraints"… See the full description on the dataset page: https://huggingface.co/datasets/snuh/specialist-level_medical_knowledge_dataset_sft.Mephisto-Knowledge_538k
Mephisto-Knowledge_538k
538,861 English knowledge SFT examples generated by
Qwen/Qwen3.5-4B in non-thinking
(Instruct) mode on the Knowledge prompts of
openbmb/UltraData-SFT-2605.
Responses contain no chain-of-thought — thinking was disabled at generation
time, so each assistant turn is a direct answer, usually with a short
justification.
Companion dataset: Mephisto-IF_172k
(instruction-following, same teacher and pipeline).
Read this before training: ref_agrees… See the full description on the dataset page: https://huggingface.co/datasets/Yxanul/Mephisto-Knowledge_538k.global-seo-knowledgeauxiliary-views-knowledge-acquisition
Auxiliary Views Knowledge Acquisition
This repository contains the cleaned source documents and evaluation
probes used in Knowledge Acquisition During Pre-training? Large Language Models
Learn Better With Auxiliary Views (arXiv:2609.04180).
News
August 21, 2026: Our paper was accepted to Findings of EMNLP 2026.
Configurations
Configuration
Split
Rows
documents
train
30
factual_cloze
test
6,435
factual_mcqa_5shot
test
4,515… See the full description on the dataset page: https://huggingface.co/datasets/jiosephlee/auxiliary-views-knowledge-acquisition.chemistry-knowledge
ChemBricks Knowledge
Does caffeine prefer water or an oil-like liquid?Why can adding one small group change a molecule's behavior?Can we design a molecule that interacts more favorably with water while meeting other constraints?How much energy does it take to remove an electron from a molecule?
These are the kinds of questions behind this dataset. Each investigation connects a question to recorded calculations, an answer, and the evidence needed to examine that answer.
Created… See the full description on the dataset page: https://huggingface.co/datasets/chembricks/chemistry-knowledge.peka_persian_knowledge_assessment
PeKA (Persian Knowledge Assessment)
PeKA is a dataset introduced in the paper "Advancing Persian LLM Evaluation", accepted at NAACL 2025 findings. It was developed as part of a broader effort to evaluate and benchmark large language models (LLMs) for multiple Persian knowledge topics.
For comprehensive details regarding the dataset’s construction, scope, task, and intended use, please refer to the original paper.
This dataset is constructed so that answering these questions… See the full description on the dataset page: https://huggingface.co/datasets/MatinaAI/peka_persian_knowledge_assessment.essential-level_medical_knowledge_dataset_sft
essential-level_medical_knowledge_dataset_sft
Dataset Summary
essential-level_medical_knowledge_dataset_sft is an integrated collection of augmented SFT data across 4 distinct medical domains, developed by the Healthcare AI Research Institute (HARI) at SNUH.
This dataset is derived and augmented from the Essential Medical Knowledge Data (필수의료 의학지식 데이터) provided by AI-Hub.
It focuses exclusively on complex clinical scenarios generated using the "Add Constraints"… See the full description on the dataset page: https://huggingface.co/datasets/snuh/essential-level_medical_knowledge_dataset_sft.Dataset-For-Indian-legal-knowledge-base About This Dataset
This dataset is the knowledge backbone of LegalEagle — an AI-powered contract review platform for Indian startups and freelancers. It contains Indian statutes, contract templates, landmark case references, and clause examples, curated specifically for retrieval-augmented generation (RAG) in the Indian legal domain.
All government statutes included are in the public domain (Government of India publications).
Dataset Structure
dataset/
├── acts/… See the full description on the dataset page: https://huggingface.co/datasets/d-riti/Dataset-For-Indian-legal-knowledge-base.oran_spec_knowledge_graph
🌐 Knowledge Graph for Open Radio Access Network (O-RAN)
A large-scale, semantically grounded knowledge graph built from O-RAN Alliance specifications,designed to enhance LLM reasoning and retrieval for next-generation telecom systems.
Overview • Motivation • Dataset Details • Getting Started • Use Cases
Overview
O-RAN (Open Radio Access Network) is an industry-driven paradigm for designing mobile networks with open, interoperable interfaces and intelligent… See the full description on the dataset page: https://huggingface.co/datasets/GSMA/oran_spec_knowledge_graph.omnimcp_graphrag_knowledge_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_graphrag_knowledge_teaser.GPTKB_v1This is the GPTKB dataset from the ACL 2025 paper:
@InProceedings{GPTKB,
title={Enabling LLM Knowledge Analysis via Extensive Materialization},
author={Hu, Yujia and Nguyen, Tuan-Phong and Ghosh, Shrestha and Razniewski, Simon},
year={2025},
booktitle={ACL},
}
Preprint: https://arxiv.org/pdf/2411.04920
Web interface for browsing GPTKB: https://gptkb.org
task685_mmmlu_answer_generation_clinical_knowledge
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task685_mmmlu_answer_generation_clinical_knowledge
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task685_mmmlu_answer_generation_clinical_knowledge.SA-Knowledge
SA-Knowledge
This repository collects corpora and evaluation data for four South African
languages: isiZulu, isiXhosa, Sepedi and Sesotho. The resources were developed
for the doctoral thesis Injecting Commonsense Knowledge into Pretrained
Language Models for Low Resource Languages (University of Cape Town, 2026).
Each subset corresponds to a thesis chapter and can be used independently.
Point of contact: Sello Ralethe
Supervisor: Dr. Jan Buys, Department of Computer Science… See the full description on the dataset page: https://huggingface.co/datasets/sello-ralethe/SA-Knowledge.ref-annotation-benchmark
RenoBench: A Citation Parsing Benchmark
RenoBench (Reference Annotation Benchmark) is a standardized evaluation benchmark for citation parsing—the task of annotating plain-text bibliographic references with structured components following the JATS (Journal Article Tag Suite) standard.
Dataset Description
RenoBench contains 10,000 plain-text citations paired with their corresponding JATS XML annotations. The dataset was assembled by extracting plain-text references from… See the full description on the dataset page: https://huggingface.co/datasets/public-knowledge-project/ref-annotation-benchmark.pharos-knowledge-packs
Pharos Knowledge Pack Library
Zero-token domain expertise for open-weight language models.
108 packs | 5,700+ triples | 50 US states covered | Verified with source URLs
What Are Pharos Packs?
Walk-encoded knowledge graphs designed for injection into a model's KV cache at inference time. No fine-tuning, no retraining, no API calls. The model gains domain expertise in milliseconds, and the packs work across any open-weight architecture.
Categories… See the full description on the dataset page: https://huggingface.co/datasets/LiberationLabs/pharos-knowledge-packs.dev-knowledge-base
Dev Knowledge Base (Programming Documentation Dataset)
A large-scale, structured dataset of programming documentation collected from official sources across languages, frameworks, tools, and AI ecosystems.
Do Follow me on Github: https://github.com/nuhmanpk
Overview
This dataset contains cleaned and structured documentation content scraped from official developer docs across multiple domains such as:
Programming languages
Frameworks (frontend, backend)
DevOps &… See the full description on the dataset page: https://huggingface.co/datasets/nuhmanpk/dev-knowledge-base.Bharat-Knowledge-Probe-Benchmark
BKP-500 — Bharat Knowledge Probe
Does your model know where it is?
BKP-500 is a benchmark of things every Indian knows and frontier LLMs routinely fumble — lakh/crore
arithmetic, Indian digit grouping, state-specific land units (bigha, katha, guntha...), traditional
mass units, the Indian fiscal year, agricultural crop seasons, government schemes, and structural
identifiers (PAN, GSTIN, IFSC, PIN codes).
The evaluation harness that runs a model against this dataset and grades… See the full description on the dataset page: https://huggingface.co/datasets/sthanika-ai/Bharat-Knowledge-Probe-Benchmark.general_knowledge_data
General Knowledge Reproduction Data
This dataset repository contains the processed General Knowledge training data used for the final reproducibility path of Tuan Dang Nguyen's CS-552 General Knowledge individual model.
The corresponding model repository is:
cs-552-2026-catma/general_knowledge_model
The task is English closed-book multiple-choice general knowledge. Models are trained to answer with exactly one option letter inside a LaTeX boxed expression, for example:… See the full description on the dataset page: https://huggingface.co/datasets/cs-552-2026-catma/general_knowledge_data.african-history-knowledge-merged-sft-cleaned
African History Knowledge Merged SFT — Cleaned
A reproducible, format-cleaned version of MaatAI/african-history-knowledge-merged-sft, pinned to source commit 0a40eb041d85d59b86219641de0fd87786ee0f77.
Split
Rows
train
28,585
validation
1,589
test
1,589
Total
31,763
Cleaning performed
Quarantined 13 training records: 12 have no final answer after a closing thinking tag, and one has ambiguous repeated closing tags. Their original text and… See the full description on the dataset page: https://huggingface.co/datasets/MaatAI/african-history-knowledge-merged-sft-cleaned.car_knowledge
car_knowledge
This dataset contains car knowledge instruction-output pairs generated for LLM fine-tuning.
Dataset Description
Each record contains:
instruction: The input question or task about car knowledge.
gpt_output: The response generated by GPT-5.
gemini_output: The response generated by Gemini.
Dataset Statistics
Total records: 3027
Files: 4 parquet file(s) in data/, up to 1000 records each.
Usage
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/jackliu2006/car_knowledge.cooking-knowledge-basics
Comprehensive Cooking Knowledge Q&A Dataset
This dataset (cooking_knowledge.csv) contains a rich collection of synthetically generated Question-Answer (Q&A) pairs covering diverse aspects of cooking knowledge, with particular emphasis on food chemistry, flavor pairing, cooking techniques, dietary accommodations, and culinary traditions. The data was created using a large language model with advanced reasoning capabilities, prompted with various grounded contexts and real-world… See the full description on the dataset page: https://huggingface.co/datasets/ktiyab/cooking-knowledge-basics.GPTKB_v1.5This hosts the GPTKB v1.5 dataset. Visit https://gptkb.org to browse GPTKB and for further information.
Papers:
GPTKB methodology: https://arxiv.org/pdf/2411.04920
GPTKB v1.5: https://arxiv.org/pdf/2507.05740
Citations:
@InProceedings{GPTKB,
title={Enabling LLM Knowledge Analysis via Extensive Materialization},
author={Hu, Yujia and Nguyen, Tuan-Phong and Ghosh, Shrestha and Razniewski, Simon},
year={2025},
booktitle={ACL},
}
@article{GPTKB15,
title={GPTKB v1.5: A Massive… See the full description on the dataset page: https://huggingface.co/datasets/Knowledge-aware-AI/GPTKB_v1.5.openthoughts3_math
OpenThoughts3 Math
This dataset contains the math-only, complete-solution subset used for supervised fine-tuning in LLM-Fusion experiments. It was derived from open-thoughts/OpenThoughts3-1.2M.
Dataset summary
103,760 training rows
32,193 unique math questions
Up to four solutions per question, selected deterministically with seed 20260910
All rows have domain = "math" and source = "ai2-adapt-dev/openmath-2-math"
Solutions are retained only when the assistant… See the full description on the dataset page: https://huggingface.co/datasets/knowledge-distillation/openthoughts3_math.Tree-of-Web-KnowledgeInspired by Tree of Knowledge (ToK), now remade as Proof of Concept: Tree-of-Web-Knowledge aka ToWK.
Alpaca Dataset created using llama2, Code, Cleaned using score of llm-blender/PairRM and dedup.
Possible improvement: - custom Web search instead of JSON obj by VinciGit00/Scrapegraph-ai.
🔍
.hf-sanitized.hf-sanitized-UDgbtn3GgVkKb3cKXMTHL .img-lbl { position: relative; display: inline-block; cursor: pointer; }
.hf-sanitized.hf-sanitized-UDgbtn3GgVkKb3cKXMTHL .pv { width: 500px; height: auto;… See the full description on the dataset page: https://huggingface.co/datasets/Nekochu/Tree-of-Web-Knowledge.my-knowledge-base
Dataset Card for GTimothee/my-knowledge-base
This repository was created using the giskard library, an open-source Python framework designed to evaluate and test AI systems.
This dataset comprises a giskard's KnowledgeBase containing 310 documents. If embeddings were generated before the saving process, they are included and will be automatically loaded into a vector store when required.
Usage
You can load this knowledge base using the following code:
from… See the full description on the dataset page: https://huggingface.co/datasets/GTimothee/my-knowledge-base.HQ-knowledgedistills-1.2M-magpieThis dataset is.an exact mix of 900k general qwen conversation with general questions, math, code and another 300k of Gemma 2 27B generations, for creative writing.
The dataset was made for "healing" pruned LLM's, especially ones based off of qwen2.5 series, as some conversations include the models saying who they are.
Unlike the previous 900K version, we also mixed in Gemma generations, to add more creative writing examples.
Many thanks to the magpie project for making this possible, this… See the full description on the dataset page: https://huggingface.co/datasets/Pinkstackorg/HQ-knowledgedistills-1.2M-magpie.
