datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
python-image-copilot-training-using-import-knowledge-graphs
Python Copilot Image Training using Import Knowledge Graphs
This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset.
Details
Each row contains a png file in the dbytes column.
Rows: 216642
Size: 211.2 GB
Data type: png
Format: Knowledge graph using NetworkX with alpaca text box
Schema
The png is in the dbytes column:
{
"dbytes": "binary"… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-image-copilot-training-using-import-knowledge-graphs.python-image-copilot-training-using-class-knowledge-graphs
Python Copilot Image Training using Class Knowledge Graphs
This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset.
Details
Each row contains a png file in the dbytes column.
Rows: 312277
Size: 304.3 GB
Data type: png
Format: Knowledge graph using NetworkX with alpaca text box
Schema
The png is in the dbytes column:
{
"dbytes": "binary"… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-image-copilot-training-using-class-knowledge-graphs.python-audio-copilot-training-using-function-knowledge-graphs
Python Copilot Audio Training using Global Functions with Knowledge Graphs
This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset.
Details
Each global function has a question and answer mp3 where one voice reads the question and another voice reads the answer. Both mp3s are stored in the parquet dbytes column and the associated source code file_path… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-audio-copilot-training-using-function-knowledge-graphs.python-audio-copilot-training-using-import-knowledge-graphs
Python Copilot Audio Training using Imports with Knowledge Graphs
This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset.
Details
Each imported module for each unique class in each module file has a question and answer mp3 where one voice reads the question and another voice reads the answer. Both mp3s are stored in the parquet dbytes column and the… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-audio-copilot-training-using-import-knowledge-graphs.python-audio-copilot-training-using-class-knowledge-graphs
Python Copilot Audio Training using Class with Knowledge Graphs
This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset.
Details
Each class method has a question and answer mp3 where one voice reads the question and another voice reads the answer. Both mp3s are stored in the parquet dbytes column and the associated source code file_path identifier.
Rows:… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-audio-copilot-training-using-class-knowledge-graphs.huatuo_knowledge_graph_qa
Dataset Card for Huatuo_knowledge_graph_qa
Dataset Summary
We built this QA dataset based on the medical knowledge map, with a total of 798,444 pieces of data, in which the questions are constructed by means of templates, and the answers are the contents of the entries in the knowledge map.
Dataset Creation
Source Data… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/huatuo_knowledge_graph_qa.python-audio-copilot-training-using-inheritance-knowledge-graphs
Python Copilot Audio Training using Inheritance and Polymorphism Knowledge Graphs
This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset.
Details
Each base class for each unique class in each module file has a question and answer mp3 where one voice reads the question and another voice reads the answer. Both mp3s are stored in the parquet dbytes column and… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-audio-copilot-training-using-inheritance-knowledge-graphs.python-audio-copilot-training-using-class-knowledge-graphs-2024-01-27
Python Copilot Audio Training using Class with Knowledge Graphs
This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset.
Details
Each class method has a question and answer mp3 where one voice reads the question and another voice reads the answer. Both mp3s are stored in the parquet dbytes column and the associated source code file_path identifier.
Rows:… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-audio-copilot-training-using-class-knowledge-graphs-2024-01-27.python-image-copilot-training-using-inheritance-knowledge-graphs
Python Copilot Image Training using Inheritance and Polymorphism Knowledge Graphs
This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset.
Details
Each row contains a png file in the dbytes column.
Rows: 259017
Size: 135.2 GB
Data type: png
Format: Knowledge graph using NetworkX with alpaca text box
Schema
The png is in the dbytes column:
{… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-image-copilot-training-using-inheritance-knowledge-graphs.python-image-copilot-training-using-class-knowledge-graphs-2024-01-27
Python Copilot Image Training using Class Knowledge Graphs
This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset.
Details
Each row contains a png file in the dbytes column.
Rows: 312836
Size: 294.1 GB
Data type: png
Format: Knowledge graph using NetworkX with alpaca text box
Schema
The png is in the dbytes column:
{
"dbytes": "binary"… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-image-copilot-training-using-class-knowledge-graphs-2024-01-27.python-image-copilot-training-using-function-knowledge-graphs
Python Copilot Image Training using Function Knowledge Graphs
This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset.
Details
Each row contains a png file in the dbytes column.
Rows: 134357
Size: 130.5 GB
Data type: png
Format: Knowledge graph using NetworkX with alpaca text box
Schema
The png is in the dbytes column:
{
"dbytes": "binary"… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-image-copilot-training-using-function-knowledge-graphs.nasa-eo-knowledge-graph
Dataset Summary
The NASA Knowledge Graph Dataset is an expansive graph-based dataset designed to integrate and interconnect information about satellite datasets, scientific publications, instruments, platforms, projects, data centers, and science keywords. This knowledge graph is particularly focused on datasets managed by NASA's Distributed Active Archive Centers (DAACs), which are NASA's data repositories responsible for archiving and distributing scientific data. In addition… See the full description on the dataset page: https://huggingface.co/datasets/nasa-gesdisc/nasa-eo-knowledge-graph.sec-xbrl-knowledge-graphs
SEC XBRL Knowledge Graph — one embedded LadybugDB file
The XBRL filings made with the SEC since January 2024 — annual and quarterly reports (10-K, 10-Q), foreign-filer annuals (20-F, 40-F), proxy statements (DEF 14A) and registration statements (S-1) — as a single
queryable property graph: 8,525 filers, 76,816 filings, 84.1 million facts on
294.5 million nodes, in one 130.3 GiB LadybugDB
file that ships as a 39.3 GiB zstd archive, written by LadybugDB 0.18.1 — open it with that… See the full description on the dataset page: https://huggingface.co/datasets/robosystems/sec-xbrl-knowledge-graphs.patent-legal-knowledge-graph
justicedao/patent-legal-knowledge-graph
JusticeDAO patent/legal knowledge_graph repository layout (schema patent-legal-hf-layout/v2, tag patent-legal-v2.0.0).
Generated by PatentHubLayoutV2. Publication is a separate operator-approved action; this card is layout metadata only.
Sources
cfr-title37-annual-2024
license: public-domain-US-government
official-edition cutoff: 2024-07-01
current-through: 2024-07-01
freshness: public legal corpus materialization
source… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/patent-legal-knowledge-graph.A-Dataset-for-Complex-Reasoning-over-Textual-Knowledge-Graphs-in-Medicine
RiTeK: Medical Textual Knowledge Graph QA Benchmark
RiTeK is a benchmark for complex reasoning over medical Textual Knowledge Graphs (medical TKGs). It evaluates whether retrieval systems and Large Language Models (LLMs) can answer realistic medical questions by using both relational paths and textual entity descriptions.
Dataset Overview
The benchmark contains three medical graph QA subsets:
Dataset
Directory
Splits
KG file
ADint
Adint
train / dev / test… See the full description on the dataset page: https://huggingface.co/datasets/ChenAI2015/A-Dataset-for-Complex-Reasoning-over-Textual-Knowledge-Graphs-in-Medicine.kaggle-knowledge-graph
Kaggle Knowledge Graph
A knowledge graph built from Kaggle's public
Meta Kaggle (and Meta Kaggle Code
datasets). It links competitions, teams, submissions, users, notebooks,
datasets, discussion forums, tags, organizations, and notebook code invocations.
See the project repository for build scripts and
documentation on the graph schema, data model, and usage examples.
Release
Field
Value
Version
2026-09-18
Meta Kaggle snapshot
2026-09-18
Build code… See the full description on the dataset page: https://huggingface.co/datasets/habedi/kaggle-knowledge-graph.omnimcp_graphrag_knowledge_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_graphrag_knowledge_teaser.oran_spec_knowledge_graph
🌐 Knowledge Graph for Open Radio Access Network (O-RAN)
A large-scale, semantically grounded knowledge graph built from O-RAN Alliance specifications,designed to enhance LLM reasoning and retrieval for next-generation telecom systems.
Overview • Motivation • Dataset Details • Getting Started • Use Cases
Overview
O-RAN (Open Radio Access Network) is an industry-driven paradigm for designing mobile networks with open, interoperable interfaces and intelligent… See the full description on the dataset page: https://huggingface.co/datasets/GSMA/oran_spec_knowledge_graph.modip-plastics-knowledge-graph
MoDiP Plastics Knowledge Graph
A standards-based knowledge graph built from the full open catalogue of the
Museum of Design in Plastics (MoDiP, Arts University Bournemouth), the UK's
only accredited museum devoted to plastics in design, as a worked example of
turning a small museum's raw "collections as data" into something computable.
11,865 object records (the complete MoDiP set), retrieved from the
Museum Data Service under CC BY 4.0.
485,013-triple CIDOC-CRM (Linked Art… See the full description on the dataset page: https://huggingface.co/datasets/fabsssss/modip-plastics-knowledge-graph.CS-Knowledge-Graph-Dataset
CS Knowledge Graph Dataset
A multi-scale heterogeneous knowledge graph of Computer Science scholarly data,
built from OpenAlex. Each scale is an independent,
self-contained subgraph centered on Computer Science papers, their authors,
publication venues, and concept tags, plus the relationships between them.
The dataset is intended for research on knowledge graph embeddings, link
prediction, node classification, scholarly recommendation, and graph neural
networks at varying scales… See the full description on the dataset page: https://huggingface.co/datasets/jugalgajjar/CS-Knowledge-Graph-Dataset.knowledge-graph-risk-engine-20260828-dataset
Knowledge Graph Risk Engine Synthetic Dataset
Summary
This dataset contains 14 training examples and 4
held-out examples for Risk teams need relationship-level explanations instead of opaque entity scores.
Every record is synthetic and includes:
input: query, event, or feature description
label: expected class, route, relation, or evidence category
context: synthetic supporting context
source: fictional source identifier
variant: generation pattern
synthetic:… See the full description on the dataset page: https://huggingface.co/datasets/RKB109/knowledge-graph-risk-engine-20260828-dataset.production-entity-knowledge-graph
SHAR Production Entity Knowledge Graph
Provenance-first bilingual graph published by SHAR Production — https://sharprod.com/.
The dataset contains public-safe nodes and relations for official SHAR Production technical surfaces and dated external page observations. Every record carries a public source_url and observed_at date. External observations are not endorsements. The graph contains no sameAs assertions.
Files
graph.json: full graph.
nodes.jsonl: one node… See the full description on the dataset page: https://huggingface.co/datasets/SHARProduction/production-entity-knowledge-graph.aicoolies-developer-tools-knowledge-graph
aicoolies-developer-tools-knowledge-graph
Public catalog dump from aicoolies.com: tools, comparisons, and scored reviews as JSON.
This dataset is not a coding agent. It does not edit repositories, run tools, or execute code. It is a machine-readable snapshot of the public aicoolies Developer Tools Knowledge Graph so humans and research agents can reuse the catalog without scraping HTML.
Canonical open-data page: https://aicoolies.com/data
Homepage: https://aicoolies.com… See the full description on the dataset page: https://huggingface.co/datasets/rasitakyol/aicoolies-developer-tools-knowledge-graph.knowledge-graph-risk-engine-20260907-dataset
Knowledge Graph Risk Engine Synthetic Dataset
Summary
This dataset contains 14 training examples and 4
held-out examples for Risk teams need relationship-level explanations instead of opaque entity scores.
Every record is synthetic and includes:
input: query, event, or feature description
label: expected class, route, relation, or evidence category
context: synthetic supporting context
source: fictional source identifier
variant: generation pattern
synthetic:… See the full description on the dataset page: https://huggingface.co/datasets/RKB109/knowledge-graph-risk-engine-20260907-dataset.knowledge-graph-rag-retrieval-artifacts
Knowledge Graph RAG Retrieval Artifacts
Pinned vector-retrieval artifacts for the Knowledge Graph RAG Assistant, a Washington State University capstone project combining knowledge-graph and dense-vector retrieval.
This repository is a project-owned, documented mirror of the two binary artifacts used by the maintained application. The files are byte-identical to the current artifacts originally hosted in miverson9/acme10-he-ragapp-embeddings at revision… See the full description on the dataset page: https://huggingface.co/datasets/ethanvillalovoz/knowledge-graph-rag-retrieval-artifacts.NLP-KnowledgeGraph
Dataset Card for Dataset Name
Dataset Summary
KG dataset created by using spaCy PoS and Dependency parser.
Supported Tasks and Leaderboards
Can be leveraged for token classification for detection of knowledge graph entities and relations.
Languages
English
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
Important fields for the token classification task are
tokens - tokenized text
tags - Tags… See the full description on the dataset page: https://huggingface.co/datasets/vishnun/NLP-KnowledgeGraph.huatuo_knowledge_graph_qa
Dataset Card for Huatuo_knowledge_graph_qa
Dataset Summary
We built this QA dataset based on the medical knowledge map, with a total of 798,444 pieces of data, in which the questions are constructed by means of templates, and the answers are the contents of the entries in the knowledge map.
Dataset Creation
Source Data… See the full description on the dataset page: https://huggingface.co/datasets/wuwu616/huatuo_knowledge_graph_qa.knowledge-graphsknowledge-graph-risk-engine-20260917-dataset
Knowledge Graph Risk Engine Synthetic Dataset
Summary
This dataset contains 14 training examples and 4
held-out examples for Risk teams need relationship-level explanations instead of opaque entity scores.
Every record is synthetic and includes:
input: query, event, or feature description
label: expected class, route, relation, or evidence category
context: synthetic supporting context
source: fictional source identifier
variant: generation pattern
synthetic:… See the full description on the dataset page: https://huggingface.co/datasets/RKB109/knowledge-graph-risk-engine-20260917-dataset.transformers-knowledge-graph
