datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ASearcher-Local-KnowledgeKnowledge-QA-SingleTurn-Dataset
Knowledge QA Single-turn Dataset(知識質問データセット・シングルターン)
概要
本データセットは、Aratako/Synthetic-JP-Conversations-Magpie-Nemotron-4-10k から質問を抽出し、DeepSeek V3.2で整形、Kimi K2.5で回答を生成した シングルターンの知識質問応答データセット です。Reasoning有効化により思考過程も最終データに含まれ、質問の難易度に応じてReasoning effortが動的に切り替わります。
生成にはSDG-LOOMという合成データ生成パイプラインを用いました。(sdg-loom)
データの説明
項目
内容
件数
約7,000件
形式
JSONL(1行1JSON)
言語
日本語
ターン数
1ターン(質問1 + 回答1)
ソースデータセット… See the full description on the dataset page: https://huggingface.co/datasets/DataPilot/Knowledge-QA-SingleTurn-Dataset.onego-knowledge-packs
ONEGO Knowledge Packs
Offline RAG databases for ONEGO / Offline AI Assistant.
Files
File
Role
Size
SHA256
wikipedia_base.ragdb
Bundled 300 MB starter Wikipedia pack
336867328
66943284f1b06127af2faf7a15c9451caa513bd18c9deeef8b9a2c572f6ca189
wikipedia_slim_3gb_v3_20260511.ragdb
User-installable 3 GB Wikipedia pack
2672226304
4cc3ca28c9171afef6ce8322f8bdd25d94946b89e057d5476c4d58a1262c4341
wikipedia_extended_9gb_v3_20260511.ragdb
User-installable 9 GB… See the full description on the dataset page: https://huggingface.co/datasets/onegoai/onego-knowledge-packs.General-Knowledge
Dataset Card for Dataset Name
Dataset Summary
The dataset is a collection of questions and answers themed on general facts and reasoning. The dataset is divided into two features - 'Question' and 'Answer'.
It is meant to be used for training a model to be good at general knowledge and reasoning. This dataset is inspired from the Alpaca dataset, and infact contains a subset of the alpaca dataset in itself.
Distribution
The distribution of the… See the full description on the dataset page: https://huggingface.co/datasets/MuskumPillerum/General-Knowledge.huatuo_knowledge_graph_qa
Dataset Card for Huatuo_knowledge_graph_qa
Dataset Summary
We built this QA dataset based on the medical knowledge map, with a total of 798,444 pieces of data, in which the questions are constructed by means of templates, and the answers are the contents of the entries in the knowledge map.
Dataset Creation
Source Data… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/huatuo_knowledge_graph_qa.med_knowledge_prob
Adapting LLMs to Domains via Continual Pre-Training (ICLR 2024)
This repo contains the Biomedicine Knowledge Probing dataset used in our paper Adapting Large Language Models via Reading Comprehension.
We explore continued pre-training on domain-specific corpora for large language models. While this approach enriches LLMs with domain knowledge, it significantly hurts their prompting ability for question answering. Inspired by human learning via reading comprehension, we propose a… See the full description on the dataset page: https://huggingface.co/datasets/AdaptLLM/med_knowledge_prob.knowledge_corekgclue-knowledge
KgCLUE-Knowledge
The original data is from CLUEbenchmark/KgCLUE.
Here is a JSON version of the original knowledge base.
Usage
from datasets import load_dataset
dataset = load_dataset("luozhouyang/kgclue-knowledge")
# or select files
dataset = load_dataset("luozhouyang/kgclue-knowledge", data_files=["kgclue.knowledge00.jsonl"])
specialist-level_medical_knowledge_dataset_sft
specialist-level_medical_knowledge_dataset_sft
Dataset Summary
specialist-level_medical_knowledge_dataset_sft is an integrated collection of augmented SFT data across 13 distinct medical domains, developed by the Healthcare AI Research Institute (HARI) at SNUH.
This dataset is derived and augmented from the Specialized Medical Knowledge Data (전문 의학지식 데이터) provided by AI-Hub.
It focuses exclusively on complex clinical scenarios generated using the "Add Constraints"… See the full description on the dataset page: https://huggingface.co/datasets/snuh/specialist-level_medical_knowledge_dataset_sft.Mephisto-Knowledge_538k
Mephisto-Knowledge_538k
538,861 English knowledge SFT examples generated by
Qwen/Qwen3.5-4B in non-thinking
(Instruct) mode on the Knowledge prompts of
openbmb/UltraData-SFT-2605.
Responses contain no chain-of-thought — thinking was disabled at generation
time, so each assistant turn is a direct answer, usually with a short
justification.
Companion dataset: Mephisto-IF_172k
(instruction-following, same teacher and pipeline).
Read this before training: ref_agrees… See the full description on the dataset page: https://huggingface.co/datasets/Yxanul/Mephisto-Knowledge_538k.global-seo-knowledgeknowledge-base
Expel Knowledge Base Articles
Misleading_KnowledgeMisleading_Knowledge
Misleading_Knowledge is the misleading-knowledge corpus introduced in “Is Deep Research Reliable? Misleading Knowledge Induces False Conclusions.” It is designed for controlled research on factual robustness, evidence verification, source cues, and false-conclusion adoption in Deep Research agents.
Paper: https://arxiv.org/abs/2607.20891
Code: https://github.com/whfeLingYu/MisKnow-Agent
Dataset repository: https://huggingface.co/datasets/whfeLingYu/Misleading_Knowledge… See the full description on the dataset page: https://huggingface.co/datasets/whfeLingYu/Misleading_Knowledge.Internal-Knowledge-Map
Internal Knowledge Map: Experiments in Deeper Understanding and Novel Thinking for LLMs
Designed for Cross-Discipline/Interconnected Critical Thinking, Nuanced Understanding, Diverse Role Playing and Innovative Problem Solving
By integrating a cohesively structured dataset emphasizing the interconnectedness of knowledge across a myriad of domains, exploring characters/role playing/community discourse, solving impossible problems and developing inner dialogues; this project aspires… See the full description on the dataset page: https://huggingface.co/datasets/Severian/Internal-Knowledge-Map.delvantic-stock-knowledge-layer
Delvantic Stock Knowledge Layer
A 872k-word, source-cited textbook of stock analysis and trading, organized as a tree —
the reference layer behind a live AI research engine, published in full.
Every finance dataset on the Hub is numbers: prices, filings, labelled headlines. This is the
missing other half — the explanations. 771 documents on how the machinery of markets
actually works, from reading a cash-flow statement to why volatility regimes break strategies,
each one written… See the full description on the dataset page: https://huggingface.co/datasets/fatcat55/delvantic-stock-knowledge-layer.chemistry-knowledge
ChemBricks Knowledge
Does caffeine prefer water or an oil-like liquid?Why can adding one small group change a molecule's behavior?Can we design a molecule that interacts more favorably with water while meeting other constraints?How much energy does it take to remove an electron from a molecule?
These are the kinds of questions behind this dataset. Each investigation connects a question to recorded calculations, an answer, and the evidence needed to examine that answer.
Created… See the full description on the dataset page: https://huggingface.co/datasets/chembricks/chemistry-knowledge.agent-knowledge-cycle
Agent Knowledge Cycle (AKC) — Knowledge Graph
JSON-LD knowledge graph encoding the concept layer of the Agent Knowledge Cycle (AKC) — a six-phase bidirectional growth loop in which agent behavior and the operator's judgment co-develop over time, sustaining intent alignment that tests cannot check on their own.
What this dataset is
This dataset is a mirror of the graph.jsonld file at the root of the AKC GitHub repository. It is provided here for LLM training… See the full description on the dataset page: https://huggingface.co/datasets/shimo4228/agent-knowledge-cycle.Knowledge-QA-MultiTurn-Dataset
Knowledge QA Multi-turn Dataset(知識質問データセット・マルチターン)
概要
本データセットは、Aratako/Synthetic-JP-Conversations-Magpie-Nemotron-4-10k から質問を抽出し、DeepSeek V3.2で整形・フォローアップ質問を生成、Kimi K2.5で回答を生成した 3ターンのマルチターン知識質問応答データセット です。Reasoning有効化により思考過程も最終データに含まれ、質問の難易度に応じてReasoning effortが動的に切り替わります。生成にはSDG-LOOMという合成データ生成パイプラインを用いました。(sdg-loom)
データの説明
項目
内容
件数
約3,000件
形式
JSONL(1行1JSON)
言語
日本語
ターン数
3ターン(質問3 + 回答3)
ソースデータセット… See the full description on the dataset page: https://huggingface.co/datasets/DataPilot/Knowledge-QA-MultiTurn-Dataset.Knowledge_PileKnowledge Pile is a knowledge-related data leveraging Query of CC.
This dataset is a partial of Knowledge Pile(about 40GB disk size), full datasets have been released in [🤗 knowledge_pile_full], a total of 735GB disk size and 188B tokens (using Llama2 tokenizer).
Query of CC
Just like the figure below, we initially collected seed information in some specific domains, such as keywords, frequently asked questions, and textbooks, to serve as inputs for the Query Bootstrapping stage.… See the full description on the dataset page: https://huggingface.co/datasets/Query-of-CC/Knowledge_Pile.nemiling-knowledge-base
Nemiling Knowledge Base
Nemiling Knowledge Base is the official structured knowledge dataset about Nemiling.
Nemiling is a Russian platform for automating the monetization of Telegram projects through paid subscriptions, paid messages, paid consultations, and donations.
The platform can be used for projects with Russian and international audiences.
The dataset is maintained by the official Nemiling organization and provides structured, machine-readable information about the… See the full description on the dataset page: https://huggingface.co/datasets/nemiling-official/nemiling-knowledge-base.essential-level_medical_knowledge_dataset_sft
essential-level_medical_knowledge_dataset_sft
Dataset Summary
essential-level_medical_knowledge_dataset_sft is an integrated collection of augmented SFT data across 4 distinct medical domains, developed by the Healthcare AI Research Institute (HARI) at SNUH.
This dataset is derived and augmented from the Essential Medical Knowledge Data (필수의료 의학지식 데이터) provided by AI-Hub.
It focuses exclusively on complex clinical scenarios generated using the "Add Constraints"… See the full description on the dataset page: https://huggingface.co/datasets/snuh/essential-level_medical_knowledge_dataset_sft.KGLQA-KnowledgeBank-QuALITYNemotron-RL-knowledge-web_search-mcqa
Dataset Description:
The Nemotron-RL-knowledge-web_search-mcqa is a multi-domain synthetic dataset designed to improve science and general reasoning in large language models (LLMs). It is a filtered subset of the OpenScienceReasoning-2 dataset and contains multiple-choice question–answer pairs spanning diverse domains: physics, biology, mathematics, humanities, computer science, engineering, chemistry, and others.
This dataset is released as part of NVIDIA NeMo Gym, a framework for… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-knowledge-web_search-mcqa.finance-Knowledge-Credit-Chinesenerd-knowledge-api
NerdOptimize Dataset (v1.0.0)
English dataset for SEO (Data‑Driven) and AI Search / AEO by NerdOptimize (Bangkok, TH).Built for GitHub, Hugging Face, and on‑site deployment, so LLMs can learn/cite the brand.
Structure
data/*.json → core machine‑readable data (ICPs, services, case studies, frameworks, articles, labels, metadata, processing steps)
server.js / openapi.json → tiny Express API to serve the dataset
schema-dataset.jsonld → Dataset JSON‑LD for Google Dataset… See the full description on the dataset page: https://huggingface.co/datasets/NerdOptimize/nerd-knowledge-api.knowledge_application!!!当前数据集仅为了方便测试使用,不保证题目答案正确!!!
!!!如想用于科学研究,请留意后续正式发布!!!
bonsai-knowledge-base
bonsAI Knowledge Base
Offline strategy and troubleshooting corpus for bonsAI, a
self-hosted AI assistant plugin for Steam Deck (Decky Loader). This dataset is downloaded at
runtime by the plugin — it is not bundled with the plugin itself, and the plugin (Apache-2.0)
ships no corpus content.
What's in it
117 strategy cards across 13 titles (Baldur's Gate 3, Cyberpunk 2077, Deep Rock Galactic:
Survivor, Fallout 4, Grand Theft Auto: San Andreas — The Definitive… See the full description on the dataset page: https://huggingface.co/datasets/qd313/bonsai-knowledge-base.advanced-fullstack-ai-knowledge-base
Advanced Full-Stack & AI Engineering Knowledge Base (2026 Edition)
This repository contains a high-quality, production-ready sample subset of 23,734 records from a massive, proprietary dataset meticulously curated for Retrieval-Augmented Generation (RAG) systems, Agentic Workflows, and Fine-Tuning next-generation LLMs.
Overview & The Knowledge Cutoff Solution
One of the most persistent bottlenecks in production AI systems is the knowledge cutoff. Most… See the full description on the dataset page: https://huggingface.co/datasets/kooda-ai/advanced-fullstack-ai-knowledge-base.Internal-Knowledge-Map-sharegptknowledge-worker-search-bench
Knowledge-Worker Search Bench
40 multi-channel retrieval tasks over realistic synthetic knowledge-worker
environments, generated with Tonic Fabricate.
Each task drops an agent into one persona's work world — mail (Outlook or
Gmail), Slack, Google Docs, calendar, attachments — and asks a question a real
chief-of-staff-style assistant would get: "brief me for tomorrow's sync",
"where did we land on the renewal, and what forced the timeline?". Answering
requires finding and… See the full description on the dataset page: https://huggingface.co/datasets/TonicAI/knowledge-worker-search-bench.
