datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cancer-knowledge-base
Cancer Knowledge Base — the open, verified oncology KB for RAG & LLM evaluation
The only open CC-BY-4.0 oncology knowledge base that combines:
110/110 trials cited with PMID + NCT + PubMed/ClinicalTrials.gov URLs, and 32 prognosis
rows linked to verified SEER 2016–2022 references — no LLM-synthetic dataset has this.
A provable 152-question MCQ benchmark — every answer derives from this KB's own structured
data and carries a citation + golden docs, so it is open-book verifiable… See the full description on the dataset page: https://huggingface.co/datasets/ranjithraj/cancer-knowledge-base.chemistry-knowledge
ChemBricks Knowledge
Does caffeine prefer water or an oil-like liquid?Why can adding one small group change a molecule's behavior?Can we design a molecule that interacts more favorably with water while meeting other constraints?How much energy does it take to remove an electron from a molecule?
These are the kinds of questions behind this dataset. Each investigation connects a question to recorded calculations, an answer, and the evidence needed to examine that answer.
Created… See the full description on the dataset page: https://huggingface.co/datasets/chembricks/chemistry-knowledge.peka_persian_knowledge_assessment
PeKA (Persian Knowledge Assessment)
PeKA is a dataset introduced in the paper "Advancing Persian LLM Evaluation", accepted at NAACL 2025 findings. It was developed as part of a broader effort to evaluate and benchmark large language models (LLMs) for multiple Persian knowledge topics.
For comprehensive details regarding the dataset’s construction, scope, task, and intended use, please refer to the original paper.
This dataset is constructed so that answering these questions… See the full description on the dataset page: https://huggingface.co/datasets/MatinaAI/peka_persian_knowledge_assessment.SA-Knowledge
SA-Knowledge
This repository collects corpora and evaluation data for four South African
languages: isiZulu, isiXhosa, Sepedi and Sesotho. The resources were developed
for the doctoral thesis Injecting Commonsense Knowledge into Pretrained
Language Models for Low Resource Languages (University of Cape Town, 2026).
Each subset corresponds to a thesis chapter and can be used independently.
Point of contact: Sello Ralethe
Supervisor: Dr. Jan Buys, Department of Computer Science… See the full description on the dataset page: https://huggingface.co/datasets/sello-ralethe/SA-Knowledge.dev-knowledge-base
Dev Knowledge Base (Programming Documentation Dataset)
A large-scale, structured dataset of programming documentation collected from official sources across languages, frameworks, tools, and AI ecosystems.
Do Follow me on Github: https://github.com/nuhmanpk
Overview
This dataset contains cleaned and structured documentation content scraped from official developer docs across multiple domains such as:
Programming languages
Frameworks (frontend, backend)
DevOps &… See the full description on the dataset page: https://huggingface.co/datasets/nuhmanpk/dev-knowledge-base.car_knowledge
car_knowledge
This dataset contains car knowledge instruction-output pairs generated for LLM fine-tuning.
Dataset Description
Each record contains:
instruction: The input question or task about car knowledge.
gpt_output: The response generated by GPT-5.
gemini_output: The response generated by Gemini.
Dataset Statistics
Total records: 3027
Files: 4 parquet file(s) in data/, up to 1000 records each.
Usage
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/jackliu2006/car_knowledge.Creative-knowledge-for-Writing
Creative knowledge for Writing
This dataset was designed to enhance or enhance the use of high-engagement words and phrases unique to high-quality novels.
It contains long excerpts of narrative text (minimum 15 sentences, maximum 55 sentences), which include:
characters' emotions,
sudden events,
plot twists,
direct dialogues with descriptions of emotions and feelings,
descriptions of landscapes, people, and things,
descriptions of sensations and feelings
The columns of the dataset… See the full description on the dataset page: https://huggingface.co/datasets/Croc-Prog-HF/Creative-knowledge-for-Writing.enterprise-rag-internal-knowledge-search-benchmark-sample
Enterprise RAG and Internal Knowledge Search Benchmark Dataset -- Free Evaluation Sample
This dataset packages a synthetic internal company workspace and an evidence-linked benchmark table into one product for teams building enterprise RAG systems, internal search assistants, knowledge-base copilots, and deep-search evaluation pipelines.
The benchmark is designed around a realistic fictional company, AsteraOps Cloud, with multiple departments, renamed projects, stale roadmaps… See the full description on the dataset page: https://huggingface.co/datasets/Karmane/enterprise-rag-internal-knowledge-search-benchmark-sample.hydrocause-knowledge-corpus
HydroCause Knowledge Corpus
The domain knowledge base released alongside the HydroCause-RAG framework. It is the "raw knowledge source" used by every downstream component:
The retrieval layer (Part 2-3) grounds numerical claims against it.
The P7 audit gate (Part 3 §3.4, SGS Audit) scores semantic reasoning against it.
The teacher-LLM synthesis (Part 5) drafts Q&A pairs from its passages.
The domain-tuned model HydroCause-3B-DK is fine-tuned on those Q&A pairs.… See the full description on the dataset page: https://huggingface.co/datasets/DamiOresotu/hydrocause-knowledge-corpus.LLMpedia
LLMpedia
Encyclopedic articles generated entirely from the parametric memory of large
language models — no retrieval — released as a benchmark for studying LLM
factuality, unverifiability, and subject-choice behavior at scale.
This dataset accompanies the paper
"LLMpedia: A Transparent Framework to Materialize an LLM's Encyclopedic
Knowledge at Scale" (Saeed & Razniewski, 2026), arXiv:2603.24080.
Motivation
Benchmarks like MMLU suggest frontier models are near… See the full description on the dataset page: https://huggingface.co/datasets/Knowledge-aware-AI/LLMpedia.simson-unified-knowledge-graph
🧠 Simson Unified Knowledge Graph
173 Nodes × 348 Edges – der Klebstoff zwischen allen Simson-Datasets.
Was das ist
Ein maschinenlesbarer Graph, der alle 6 Datasets miteinander verknüpft:
Dataset
Status
Nodes
racing-planet-simson-traces
Diagnose-Traces
15
simson-forum-qa-pairs
Forum-Wissen
30
simson-repair-manual
Technische Daten
14
racing-planet-product-catalog
Teilekatalog
37
simson-youtube-tutorials
Video-Tutorials
20… See the full description on the dataset page: https://huggingface.co/datasets/jmp1987/simson-unified-knowledge-graph.enterprise-rag-internal-knowledge-search-benchmark
Enterprise RAG and Internal Knowledge Search Benchmark Dataset
This dataset packages a synthetic internal company workspace and an evidence-linked benchmark table into one product for teams building enterprise RAG systems, internal search assistants, knowledge-base copilots, and deep-search evaluation pipelines.
The benchmark is designed around a realistic fictional company, AsteraOps Cloud, with multiple departments, renamed projects, stale roadmaps, support escalations… See the full description on the dataset page: https://huggingface.co/datasets/Karmane/enterprise-rag-internal-knowledge-search-benchmark.palestinian-cultural-knowledge
Palestinian Cultural Knowledge Corpus
v0.2.0 — supersedes the earlier data/wikipedia_ar/ v0.1.0 partial upload
(484 Arabic Wikipedia documents only). This release expands to the full 5-source
corpus below and moves the data to data/full_corpus/.
A multi-source Arabic/English text corpus about Palestinian history, culture, and
heritage, built for the Palestinian Cultural Knowledge
Platform
— a RAG + knowledge-graph research project. 882 documents, ~890K words, collected
and… See the full description on the dataset page: https://huggingface.co/datasets/palestinian-kg/palestinian-cultural-knowledge.
