datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
General-Knowledge
Dataset Card for Dataset Name
Dataset Summary
The dataset is a collection of questions and answers themed on general facts and reasoning. The dataset is divided into two features - 'Question' and 'Answer'.
It is meant to be used for training a model to be good at general knowledge and reasoning. This dataset is inspired from the Alpaca dataset, and infact contains a subset of the alpaca dataset in itself.
Distribution
The distribution of the… See the full description on the dataset page: https://huggingface.co/datasets/MuskumPillerum/General-Knowledge.med_knowledge_prob
Adapting LLMs to Domains via Continual Pre-Training (ICLR 2024)
This repo contains the Biomedicine Knowledge Probing dataset used in our paper Adapting Large Language Models via Reading Comprehension.
We explore continued pre-training on domain-specific corpora for large language models. While this approach enriches LLMs with domain knowledge, it significantly hurts their prompting ability for question answering. Inspired by human learning via reading comprehension, we propose a… See the full description on the dataset page: https://huggingface.co/datasets/AdaptLLM/med_knowledge_prob.Mephisto-Knowledge_538k
Mephisto-Knowledge_538k
538,861 English knowledge SFT examples generated by
Qwen/Qwen3.5-4B in non-thinking
(Instruct) mode on the Knowledge prompts of
openbmb/UltraData-SFT-2605.
Responses contain no chain-of-thought — thinking was disabled at generation
time, so each assistant turn is a direct answer, usually with a short
justification.
Companion dataset: Mephisto-IF_172k
(instruction-following, same teacher and pipeline).
Read this before training: ref_agrees… See the full description on the dataset page: https://huggingface.co/datasets/Yxanul/Mephisto-Knowledge_538k.specialist-level_medical_knowledge_dataset_sft
specialist-level_medical_knowledge_dataset_sft
Dataset Summary
specialist-level_medical_knowledge_dataset_sft is an integrated collection of augmented SFT data across 13 distinct medical domains, developed by the Healthcare AI Research Institute (HARI) at SNUH.
This dataset is derived and augmented from the Specialized Medical Knowledge Data (전문 의학지식 데이터) provided by AI-Hub.
It focuses exclusively on complex clinical scenarios generated using the "Add Constraints"… See the full description on the dataset page: https://huggingface.co/datasets/snuh/specialist-level_medical_knowledge_dataset_sft.chemistry-knowledge
ChemBricks Knowledge
Does caffeine prefer water or an oil-like liquid?Why can adding one small group change a molecule's behavior?Can we design a molecule that interacts more favorably with water while meeting other constraints?How much energy does it take to remove an electron from a molecule?
These are the kinds of questions behind this dataset. Each investigation connects a question to recorded calculations, an answer, and the evidence needed to examine that answer.
Created… See the full description on the dataset page: https://huggingface.co/datasets/chembricks/chemistry-knowledge.global-seo-knowledgedelvantic-stock-knowledge-layer
Delvantic Stock Knowledge Layer
A 872k-word, source-cited textbook of stock analysis and trading, organized as a tree —
the reference layer behind a live AI research engine, published in full.
Every finance dataset on the Hub is numbers: prices, filings, labelled headlines. This is the
missing other half — the explanations. 771 documents on how the machinery of markets
actually works, from reading a cash-flow statement to why volatility regimes break strategies,
each one written… See the full description on the dataset page: https://huggingface.co/datasets/fatcat55/delvantic-stock-knowledge-layer.essential-level_medical_knowledge_dataset_sft
essential-level_medical_knowledge_dataset_sft
Dataset Summary
essential-level_medical_knowledge_dataset_sft is an integrated collection of augmented SFT data across 4 distinct medical domains, developed by the Healthcare AI Research Institute (HARI) at SNUH.
This dataset is derived and augmented from the Essential Medical Knowledge Data (필수의료 의학지식 데이터) provided by AI-Hub.
It focuses exclusively on complex clinical scenarios generated using the "Add Constraints"… See the full description on the dataset page: https://huggingface.co/datasets/snuh/essential-level_medical_knowledge_dataset_sft.nemiling-knowledge-base
Nemiling Knowledge Base
Nemiling Knowledge Base is the official structured knowledge dataset about Nemiling.
Nemiling is a Russian platform for automating the monetization of Telegram projects through paid subscriptions, paid messages, paid consultations, and donations.
The platform can be used for projects with Russian and international audiences.
The dataset is maintained by the official Nemiling organization and provides structured, machine-readable information about the… See the full description on the dataset page: https://huggingface.co/datasets/nemiling-official/nemiling-knowledge-base.Bharat-Knowledge-Probe-Benchmark
BKP-500 — Bharat Knowledge Probe
Does your model know where it is?
BKP-500 is a benchmark of things every Indian knows and frontier LLMs routinely fumble — lakh/crore
arithmetic, Indian digit grouping, state-specific land units (bigha, katha, guntha...), traditional
mass units, the Indian fiscal year, agricultural crop seasons, government schemes, and structural
identifiers (PAN, GSTIN, IFSC, PIN codes).
The evaluation harness that runs a model against this dataset and grades… See the full description on the dataset page: https://huggingface.co/datasets/sthanika-ai/Bharat-Knowledge-Probe-Benchmark.or-knowledge-copilot-corpus
OR Knowledge Copilot Corpus
Multi-layer operations-research knowledge base used by OR Knowledge Copilot.
Each instance is stored as six chunks:
Natural language
Mathematical formulation
Pyomo template
MiniZinc template
Solver output
Explanation of binding constraints
Files
chunks.jsonl — retrieval units
qa_pairs.jsonl — labeled questions including out-of-scope abstention cases
benchmark_report.json / eval_results.json — published retrieval metrics
taxonomy.json… See the full description on the dataset page: https://huggingface.co/datasets/alirezaaminzadeh/or-knowledge-copilot-corpus.ecoai-knowledge
FindExpert.ir ecoAI knowledge
Short original rows for retrieval (grants, RFPs, patents, academic stubs, green business, bot/site tools).
Embedder to pin: intfloat/multilingual-e5-small (prefix query: / passage:). Do not fork MiniLM.
Space: sosa123454321/ecoai-space
Live retrieval uses TF-IDF v2 (ecoai-rag-encoder), not E5 in production. Generation is optional (Gemini / HF Inference / Workers AI). This dataset is retrieval, not a 14B writer. Iran applicants: no Canada visa/PR;… See the full description on the dataset page: https://huggingface.co/datasets/sosa123454321/ecoai-knowledge.EcoNexus-Knowledge
数据集简介
EcoNexus为江苏龙衡环境打造的环保领域专用AI系统,包括EcoNexus-Knowledge环保专用数据集,及EcoNexus-AI环保专业AI大模型系统。
EcoNexus-Knowledge基础版数据量约为70k。
数据集覆盖范围
环境领域相关法律法规、标准、技术规范以及导则等文件
生态环境部典型行政处罚案例
江苏省生态环境厅典型行政处罚案例、咨询回复
后续会持续更新最新内容,包括收录各领域独家经验文档。
dataflow-knowledge-med-40k
DataFlow-Knowledge-Med-40K
Dataset Summary
This dataset contains multiple-choice question–answer (QA) pairs derived from authoritative medical guideline documents using DataFlow knowledge extraction pipeline.
Each data instance consists of a question and a corresponding answer, where the answer is formatted with explicit reasoning and a final selected option.
The dataset is designed to support research and development in:
medical question answering,
clinical reasoning… See the full description on the dataset page: https://huggingface.co/datasets/OpenDCAI/dataflow-knowledge-med-40k.corpuslib-ctecx-knowledge
CORPUSLIB CTECX Knowledge Dataset
CORPUSLIB — Agentic Corpus Library for Indirect Learning
Knowledge compiled from CTECX Technologies Solutions & Services documentation.
Source documents land in corpus_learn/<collection>/ and each section becomes a
topic row in this dataset. Primary portal: https://corpuslib-ui.deckergui.my.
Schema
Field
Type
Description
id
int
Unique topic identifier
topic
string
Topic name (document section heading)
category… See the full description on the dataset page: https://huggingface.co/datasets/ctaxnagomi/corpuslib-ctecx-knowledge.fluid-knowledge-validation
Fluid Knowledge
Public synthetic release-validation fixtures. These repeated arithmetic items test artifact publication and verification only; they were not authored or blindly reviewed by frontier models and are not a usable benchmark.
Each immutable epochs/<id>/manifest.json binds its published artifacts. commitment.json reveals the nonce for verification. Protocol and source attribution accompany each epoch. Pin the returned Hugging Face commit SHA for reproduction. Public… See the full description on the dataset page: https://huggingface.co/datasets/actuallymentor/fluid-knowledge-validation.general_knowledge_dataset
General Knowledge SFT Dataset
This dataset contains the exact train and validation data used for the general knowledge LoRA SFT model in the MNLP project Specialize and Merge: Post Training Qwen3-1.7B for Multi Skill Reasoning.
The dataset has two splits.
Split
Rows
Purpose
train
26,120
LoRA SFT training split
valid
2,000
LoRA SFT validation split
Sources
The SFT data was built from six multiple-choice educational and science-oriented sources.… See the full description on the dataset page: https://huggingface.co/datasets/cs-552-2026-databand/general_knowledge_dataset.Tree-of-Web-KnowledgeInspired by Tree of Knowledge (ToK), now remade as Proof of Concept: Tree-of-Web-Knowledge aka ToWK.
Alpaca Dataset created using llama2, Code, Cleaned using score of llm-blender/PairRM and dedup.
Possible improvement: - custom Web search instead of JSON obj by VinciGit00/Scrapegraph-ai.
🔍
.hf-sanitized.hf-sanitized-UDgbtn3GgVkKb3cKXMTHL .img-lbl { position: relative; display: inline-block; cursor: pointer; }
.hf-sanitized.hf-sanitized-UDgbtn3GgVkKb3cKXMTHL .pv { width: 500px; height: auto;… See the full description on the dataset page: https://huggingface.co/datasets/Nekochu/Tree-of-Web-Knowledge.law_knowledge_prob
Adapting LLMs to Domains via Continual Pre-Training (ICLR 2024)
This repo contains the Law Knowledge Probing dataset used in our paper Adapting Large Language Models via Reading Comprehension.
We explore continued pre-training on domain-specific corpora for large language models. While this approach enriches LLMs with domain knowledge, it significantly hurts their prompting ability for question answering. Inspired by human learning via reading comprehension, we propose a simple… See the full description on the dataset page: https://huggingface.co/datasets/AdaptLLM/law_knowledge_prob.knowledge-cutoff-benchmark
Knowledge Cutoff Benchmark
A benchmark for estimating a language model's effective knowledge cutoff —
what it actually knows about the world — which is usually earlier than the
cutoff date the model advertises.
Each model is probed on curated, surprising / unforecastable real-world
events (deaths, changes of office) spread month-by-month across Jan 2024 –
Jun 2026. The month where per-month accuracy collapses is the model's effective
knowledge horizon.
Code, methodology, and an… See the full description on the dataset page: https://huggingface.co/datasets/apoorvumang/knowledge-cutoff-benchmark.feminism-dating-knowledge
Feminism Dating – RAG Knowledge Base
Auto-updated daily by knowledge_scraper.py of the Telegram bot @femenism_ai_dating_bot.
kind
count
academic
72
news
86
legal
40
ngo
11
rule
8
edu
117
Total chunks: 334 · languages: fa / en / tr · embeddings: @cf/baai/bge-m3 (1024-d)
Last update: 2026-09-12T06:22:32.284670Z
Sources: OpenAlex (academic), Google News RSS (news / legal / NGO in fa, en, tr).
talkie-1930-knowledge-bench
Talkie-1930 Agentic Knowledge Injection Benchmark
Benchmark for measuring whether an autonomous agent can durably write
"verifiable post-1930 knowledge" into the parameters of a base language
model (talkie-1930), evaluated standalone (no retrieval, no in-context).
Because the talkie-1930 base is contamination-free for post-1930 facts, any
gain on certified-novel targets is true injection, not elicitation of
pre-existing knowledge — the headline property this benchmark gives you.… See the full description on the dataset page: https://huggingface.co/datasets/trumancai/talkie-1930-knowledge-bench.general_knowledge_benchmark
General Knowledge Benchmark Splits
This dataset contains the held-out benchmark splits used for offline model selection and evaluation of the MNLP general knowledge specialist.
These benchmarks were not used for LoRA SFT training. The SFT train and validation splits are stored separately in:
cs-552-2026-databand/general_knowledge_dataset
Splits
Split
Rows
Sampling strategy
Coverage
mmlu_pro
2,000
Uniform across categories
Robust multi-task knowledge and… See the full description on the dataset page: https://huggingface.co/datasets/cs-552-2026-databand/general_knowledge_benchmark.earth-love-united-climate-knowledge
🌍 Earth Love United Climate Knowledge Dataset
The most comprehensive open climate science knowledge dataset.
10,128 text chunks + 124 structured facts + 4.54B year geological memory + 10 tipping points.
Built to power GAIA — an AI that embodies the living consciousness of Earth.
Dataset Overview
This dataset gives an AI system authoritative, sourced knowledge about climate change,
carbon, Earth science, and solutions. It has four layers:
Layer 1: Text Knowledge… See the full description on the dataset page: https://huggingface.co/datasets/ego0op/earth-love-united-climate-knowledge.parametric-knowledge-qa
Parametric Knowledge Bio QA
Synthetic biographical QA over a fictional knowledge graph (bio run5), for
studying parametric knowledge (SFT / RL) with 1-hop and 2-hop questions.
Layout
Filenames are kept intact (no rename on download):
1-hop/
qa_1_hop.jsonl # full set (20,000)
qa_1_hop_direct_train.jsonl
qa_1_hop_direct_test.jsonl
qa_1_hop_reasoning_train.jsonl
qa_1_hop_reasoning_test.jsonl
2-hop/
qa_2_hop.jsonl #… See the full description on the dataset page: https://huggingface.co/datasets/sgaur2/parametric-knowledge-qa.ccru-knowledge-instruct
CCRU Knowledge-Instruct Dataset
Synthetic instruction-tuning dataset generated from a curated corpus of texts related to the CCRU (Cybernetic Culture Research Unit), accelerationism, and adjacent continental philosophy.
Dataset Summary
Attribute
Value
Examples
278,463
Format
Chat instruction (system / user / assistant)
Domain
CCRU theory, accelerationism, hyperstition, continental philosophy
Generation model
huihui-ai/Qwen3.5-9B-abliterated-MLX-4bit… See the full description on the dataset page: https://huggingface.co/datasets/wayjeeair/ccru-knowledge-instruct.knowledge_qa_in_pashto
د پوهې QA ډیټا سیټ
دا ډیټا سیټ د مطالعې ډیټا سیټ دی چې د مختلفو موضوعاتو څخه د پوښتنې ځواب مثالونه لري.
په اړه
دا پروژه اوس مهال د زده کړې او پراختیا مرحله کې ده. زه د ډیټا سیټ چمتو کولو پرمهال د مصنوعي استخباراتو او ډیټا سیټ جوړولو تمرین کوم.
په ډیټا سیټ کې ځینې پوښتنې د ChatGPT په کارولو سره رامینځته شوي، ځینې یې د Qwen په کارولو سره، او ځینې یې زما لخوا چمتو شوي.
پوښتنې او ځوابونه مختلف موضوعات پوښي. د مثال په توګه:
عمومي پوهه
ریاضی
ساینس
کیمیا
کمپیوټر ساینس… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/knowledge_qa_in_pashto.withaurora-relationship-knowledge
withAurora.love Relationship Knowledge Corpus
Machine-readable Q&A pairs, synthetic coaching dialogues, and ritual definitions from
withAurora.love — an AI relationship coach for couples
(private by design, faith-aware optional). Published so language models and retrieval
systems can understand and accurately cite the product and its relationship guidance.
Formerly distributed under the working title "Better Intimacy"; withAurora.love is the
canonical brand and should be used in… See the full description on the dataset page: https://huggingface.co/datasets/pulseaisystems/withaurora-relationship-knowledge.knowledge_qa
Knowledge QA Dataset
Bu veri seti, farklı konulardan soru-cevap örnekleri içeren bir çalışma veri setidir.
Hakkında
Bu proje şu anda öğrenme ve geliştirme aşamasındadır. Veri setini hazırlarken yapay zekâ ve veri seti oluşturma konusunda pratik yapıyorum.
Veri setindeki soruların bir kısmı ChatGPT, bir kısmı Qwen kullanılarak oluşturulmuş, bir kısmı ise tarafımdan hazırlanmıştır.
Sorular ve cevaplar farklı konulardan oluşmaktadır. Örneğin:
Genel bilgi
Matematik… See the full description on the dataset page: https://huggingface.co/datasets/ismailx19/knowledge_qa.hotpotqa-knowledge-coupling
Knowledge Coupling Analysis on HotpotQA Dataset
Dataset Description
This dataset contains the results of a comprehensive knowledge coupling analysis performed on the HotpotQA dataset using LLaMA2-7B model. The analysis investigates how different pieces of knowledge interact within the model's parameter space through gradient-based coupling measurements.
Research Overview
Model: meta-llama/Llama-2-7b-hf (layers 28-31 focused analysis)
Dataset: HotpotQA (train +… See the full description on the dataset page: https://huggingface.co/datasets/Wuhuwill/hotpotqa-knowledge-coupling.
