datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Misleading_KnowledgeMisleading_Knowledge
Misleading_Knowledge is the misleading-knowledge corpus introduced in “Is Deep Research Reliable? Misleading Knowledge Induces False Conclusions.” It is designed for controlled research on factual robustness, evidence verification, source cues, and false-conclusion adoption in Deep Research agents.
Paper: https://arxiv.org/abs/2607.20891
Code: https://github.com/whfeLingYu/MisKnow-Agent
Dataset repository: https://huggingface.co/datasets/whfeLingYu/Misleading_Knowledge… See the full description on the dataset page: https://huggingface.co/datasets/whfeLingYu/Misleading_Knowledge.chemistry-knowledge
ChemBricks Knowledge
Does caffeine prefer water or an oil-like liquid?Why can adding one small group change a molecule's behavior?Can we design a molecule that interacts more favorably with water while meeting other constraints?How much energy does it take to remove an electron from a molecule?
These are the kinds of questions behind this dataset. Each investigation connects a question to recorded calculations, an answer, and the evidence needed to examine that answer.
Created… See the full description on the dataset page: https://huggingface.co/datasets/chembricks/chemistry-knowledge.delvantic-stock-knowledge-layer
Delvantic Stock Knowledge Layer
A 872k-word, source-cited textbook of stock analysis and trading, organized as a tree —
the reference layer behind a live AI research engine, published in full.
Every finance dataset on the Hub is numbers: prices, filings, labelled headlines. This is the
missing other half — the explanations. 771 documents on how the machinery of markets
actually works, from reading a cash-flow statement to why volatility regimes break strategies,
each one written… See the full description on the dataset page: https://huggingface.co/datasets/fatcat55/delvantic-stock-knowledge-layer.agent-knowledge-cycle
Agent Knowledge Cycle (AKC) — Knowledge Graph
JSON-LD knowledge graph encoding the concept layer of the Agent Knowledge Cycle (AKC) — a six-phase bidirectional growth loop in which agent behavior and the operator's judgment co-develop over time, sustaining intent alignment that tests cannot check on their own.
What this dataset is
This dataset is a mirror of the graph.jsonld file at the root of the AKC GitHub repository. It is provided here for LLM training… See the full description on the dataset page: https://huggingface.co/datasets/shimo4228/agent-knowledge-cycle.nemiling-knowledge-base
Nemiling Knowledge Base
Nemiling Knowledge Base is the official structured knowledge dataset about Nemiling.
Nemiling is a Russian platform for automating the monetization of Telegram projects through paid subscriptions, paid messages, paid consultations, and donations.
The platform can be used for projects with Russian and international audiences.
The dataset is maintained by the official Nemiling organization and provides structured, machine-readable information about the… See the full description on the dataset page: https://huggingface.co/datasets/nemiling-official/nemiling-knowledge-base.bonsai-knowledge-base
bonsAI Knowledge Base
Offline strategy and troubleshooting corpus for bonsAI, a
self-hosted AI assistant plugin for Steam Deck (Decky Loader). This dataset is downloaded at
runtime by the plugin — it is not bundled with the plugin itself, and the plugin (Apache-2.0)
ships no corpus content.
What's in it
117 strategy cards across 13 titles (Baldur's Gate 3, Cyberpunk 2077, Deep Rock Galactic:
Survivor, Fallout 4, Grand Theft Auto: San Andreas — The Definitive… See the full description on the dataset page: https://huggingface.co/datasets/qd313/bonsai-knowledge-base.or-knowledge-copilot-corpus
OR Knowledge Copilot Corpus
Multi-layer operations-research knowledge base used by OR Knowledge Copilot.
Each instance is stored as six chunks:
Natural language
Mathematical formulation
Pyomo template
MiniZinc template
Solver output
Explanation of binding constraints
Files
chunks.jsonl — retrieval units
qa_pairs.jsonl — labeled questions including out-of-scope abstention cases
benchmark_report.json / eval_results.json — published retrieval metrics
taxonomy.json… See the full description on the dataset page: https://huggingface.co/datasets/alirezaaminzadeh/or-knowledge-copilot-corpus.aicoolies-developer-tools-knowledge-graph
aicoolies-developer-tools-knowledge-graph
Public catalog dump from aicoolies.com: tools, comparisons, and scored reviews as JSON.
This dataset is not a coding agent. It does not edit repositories, run tools, or execute code. It is a machine-readable snapshot of the public aicoolies Developer Tools Knowledge Graph so humans and research agents can reuse the catalog without scraping HTML.
Canonical open-data page: https://aicoolies.com/data
Homepage: https://aicoolies.com… See the full description on the dataset page: https://huggingface.co/datasets/rasitakyol/aicoolies-developer-tools-knowledge-graph.repro-understanding-lora-as-knowledge-memory-an-empirical-analysis-traces
Agent traces
Agent sessions published from a Trackio Logbook.
earth-love-united-climate-knowledge
🌍 Earth Love United Climate Knowledge Dataset
The most comprehensive open climate science knowledge dataset.
10,128 text chunks + 124 structured facts + 4.54B year geological memory + 10 tipping points.
Built to power GAIA — an AI that embodies the living consciousness of Earth.
Dataset Overview
This dataset gives an AI system authoritative, sourced knowledge about climate change,
carbon, Earth science, and solutions. It has four layers:
Layer 1: Text Knowledge… See the full description on the dataset page: https://huggingface.co/datasets/ego0op/earth-love-united-climate-knowledge.LLMpedia
LLMpedia
Encyclopedic articles generated entirely from the parametric memory of large
language models — no retrieval — released as a benchmark for studying LLM
factuality, unverifiability, and subject-choice behavior at scale.
This dataset accompanies the paper
"LLMpedia: A Transparent Framework to Materialize an LLM's Encyclopedic
Knowledge at Scale" (Saeed & Razniewski, 2026), arXiv:2603.24080.
Motivation
Benchmarks like MMLU suggest frontier models are near… See the full description on the dataset page: https://huggingface.co/datasets/Knowledge-aware-AI/LLMpedia.drtulu_v2_stepfun_scientific_knowledge_0415FactoryBench-KnowledgeGraphrag-demo-knowledgeFactoryBench-KnowledgeGraphknowledgeAnna
