datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
python-image-copilot-training-using-import-knowledge-graphs
Python Copilot Image Training using Import Knowledge Graphs
This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset.
Details
Each row contains a png file in the dbytes column.
Rows: 216642
Size: 211.2 GB
Data type: png
Format: Knowledge graph using NetworkX with alpaca text box
Schema
The png is in the dbytes column:
{
"dbytes": "binary"… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-image-copilot-training-using-import-knowledge-graphs.python-image-copilot-training-using-class-knowledge-graphs
Python Copilot Image Training using Class Knowledge Graphs
This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset.
Details
Each row contains a png file in the dbytes column.
Rows: 312277
Size: 304.3 GB
Data type: png
Format: Knowledge graph using NetworkX with alpaca text box
Schema
The png is in the dbytes column:
{
"dbytes": "binary"… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-image-copilot-training-using-class-knowledge-graphs.GraphQApython-audio-copilot-training-using-function-knowledge-graphs
Python Copilot Audio Training using Global Functions with Knowledge Graphs
This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset.
Details
Each global function has a question and answer mp3 where one voice reads the question and another voice reads the answer. Both mp3s are stored in the parquet dbytes column and the associated source code file_path… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-audio-copilot-training-using-function-knowledge-graphs.python-audio-copilot-training-using-import-knowledge-graphs
Python Copilot Audio Training using Imports with Knowledge Graphs
This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset.
Details
Each imported module for each unique class in each module file has a question and answer mp3 where one voice reads the question and another voice reads the answer. Both mp3s are stored in the parquet dbytes column and the… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-audio-copilot-training-using-import-knowledge-graphs.python-audio-copilot-training-using-class-knowledge-graphs
Python Copilot Audio Training using Class with Knowledge Graphs
This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset.
Details
Each class method has a question and answer mp3 where one voice reads the question and another voice reads the answer. Both mp3s are stored in the parquet dbytes column and the associated source code file_path identifier.
Rows:… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-audio-copilot-training-using-class-knowledge-graphs.AnonyRAG
AnnoyRAG Dataset
The AnnoyRAG dataset, introduced in Youtu-GraphRAG: Vertically Unified Agents for Graph Retrieval-Augmented Complex Reasoning, employs entity anonymization to isolate LLMs' parametric knowledge. This design enables more precise evaluation of how effectively LLMs integrate retrieved information in RAG systems.
Dataset Details
Dataset Description
The basic statistical information of the dataset is as follows:
Question Type
Difficulty Level… See the full description on the dataset page: https://huggingface.co/datasets/Youtu-Graph/AnonyRAG.sma-evidence-graph
SMA Evidence Graph
An open-source, evidence-first dataset for Spinal Muscular Atrophy (SMA) drug research.
Description
This dataset contains structured evidence extracted from PubMed papers, clinical trials
from ClinicalTrials.gov, computationally generated hypotheses, AI-designed molecules,
and DiffDock molecular docking results — all linking gene targets to potential
therapeutic interventions for SMA.
Built by a researcher who has SMA, this dataset aims to accelerate… See the full description on the dataset page: https://huggingface.co/datasets/SMAResearch/sma-evidence-graph.python-audio-copilot-training-using-inheritance-knowledge-graphs
Python Copilot Audio Training using Inheritance and Polymorphism Knowledge Graphs
This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset.
Details
Each base class for each unique class in each module file has a question and answer mp3 where one voice reads the question and another voice reads the answer. Both mp3s are stored in the parquet dbytes column and… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-audio-copilot-training-using-inheritance-knowledge-graphs.python-audio-copilot-training-using-class-knowledge-graphs-2024-01-27
Python Copilot Audio Training using Class with Knowledge Graphs
This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset.
Details
Each class method has a question and answer mp3 where one voice reads the question and another voice reads the answer. Both mp3s are stored in the parquet dbytes column and the associated source code file_path identifier.
Rows:… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-audio-copilot-training-using-class-knowledge-graphs-2024-01-27.python-image-copilot-training-using-inheritance-knowledge-graphs
Python Copilot Image Training using Inheritance and Polymorphism Knowledge Graphs
This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset.
Details
Each row contains a png file in the dbytes column.
Rows: 259017
Size: 135.2 GB
Data type: png
Format: Knowledge graph using NetworkX with alpaca text box
Schema
The png is in the dbytes column:
{… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-image-copilot-training-using-inheritance-knowledge-graphs.GraphRAG
Local GraphRAG Research Artifacts
This repository contains the data and output artifacts produced for my master's
thesis research on knowledge graph construction in local GraphRAG systems. It
is an evidence package for the study, rather than the code used to run the
experiment.
The collection includes the study corpus, model extraction outputs, final
knowledge graphs, retrieved evidence, generated answers, and evaluation results.
Together, these files make it possible to inspect… See the full description on the dataset page: https://huggingface.co/datasets/boblaros/GraphRAG.python-image-copilot-training-using-class-knowledge-graphs-2024-01-27
Python Copilot Image Training using Class Knowledge Graphs
This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset.
Details
Each row contains a png file in the dbytes column.
Rows: 312836
Size: 294.1 GB
Data type: png
Format: Knowledge graph using NetworkX with alpaca text box
Schema
The png is in the dbytes column:
{
"dbytes": "binary"… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-image-copilot-training-using-class-knowledge-graphs-2024-01-27.python-image-copilot-training-using-function-knowledge-graphs
Python Copilot Image Training using Function Knowledge Graphs
This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset.
Details
Each row contains a png file in the dbytes column.
Rows: 134357
Size: 130.5 GB
Data type: png
Format: Knowledge graph using NetworkX with alpaca text box
Schema
The png is in the dbytes column:
{
"dbytes": "binary"… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-image-copilot-training-using-function-knowledge-graphs.luatdo-graph
luatdo-graph
A knowledge graph over Vietnamese law, built from about 128,000 documents by luatdo.
This is the result of running the pipeline, published so that nobody has to run it again.
The pipeline takes days and several hundred dollars of model calls, and the output is the same for everyone.
What is in it
Nodes
8,175,346
Relationships
9,119,011
Node tables
14
Relationship tables
19
Node labels
13
Parquet
623MB across 47 files
Neo4j… See the full description on the dataset page: https://huggingface.co/datasets/open-index/luatdo-graph.Graph-R1-dataset-complete
Graph-R1 Complete Dataset
This dataset contains the complete Graph-R1 graph reasoning dataset with all difficulty levels (1-5).
Files
train_graph_all_levels.parquet: Combined training data from all levels (with level column)
test_graph_and_math_all_levels.parquet: Combined test data from all levels (with level column)
test_graph_mixedsize_cleaned.parquet: Mixed size test data (with level='mixed')
Usage
import pandas as pd
# Load combined data… See the full description on the dataset page: https://huggingface.co/datasets/HKUST-DSAIL/Graph-R1-dataset-complete.GTSQA
Dataset card for GTSQA
Dataset Summary
GTSQA is a synthetic Knowledge Graph Question Answering dataset constructed from Wikidata, using the SynthKGQA framework. It offers a challenging benchmark for GraphRAG models and KG-augmented LLMs, and enables the stand-alone evaluation of a KG retriever's performance, by providing the set of ground-truth KG edges that are required to reason over each question. It is specifically designed to test generalization abilities of KG… See the full description on the dataset page: https://huggingface.co/datasets/Graphcore/GTSQA.qualora-workforce-skills-graph
Qualora Workforce Skills Graph (Representative Sample)
Rights-clean, provenance-tracked vocational learning data, rebuilt from roughly $2B of U.S. Department of Labor funded open courseware into a labeled skills graph: cleaned courses and lessons, Bloom-tagged assessment items with answer rationales and learning objectives, and a content-grounded course to skill to career graph with salary context. Built for post-training and evaluation, not pretraining bulk.
This repository is… See the full description on the dataset page: https://huggingface.co/datasets/qualora-data-labs/qualora-workforce-skills-graph.islamic-corpus-graph
QuranLab — Qur'an & Hadith Structured Corpus and Knowledge Graph
A unified, verse- and ḥadīth-aligned structured corpus for the Qur'an and the canonical Sunnah,
assembled by volunteers under the QuranLab effort. It links Qur'anic verses, multilingual
translations, classical tafsīr, word-level morphology, and ḥadīth text with normalized authenticity
grades into one consistent graph, alongside retrieval passages, grounded question–answer pairs and a
held-out evaluation set. Every… See the full description on the dataset page: https://huggingface.co/datasets/quranlab/islamic-corpus-graph.fin-cfa-graphgen
Fin-CFA-GraphGen: 785K Knowledge-Guided Financial QA Examples
Fin-CFA-GraphGen is a large-scale English dataset for financial instruction tuning, financial question answering, and domain-specific language-model post-training. It contains 785,149 synthetic question–answer examples generated from CFA curriculum and exam-preparation books with the GraphGen knowledge-driven data-generation method.
The dataset and its role in the post-training pipeline are described in Data-Centric… See the full description on the dataset page: https://huggingface.co/datasets/whoisjiji/fin-cfa-graphgen.omnimcp_graphrag_knowledge_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_graphrag_knowledge_teaser.Graph-R1-RFT-COT-30K
Dataset Card: Graph-CoT-30k
Dataset Details
Dataset Name: Graph-CoT-30k
Dataset Creator: HKUST-DSAIL
Dataset Version: 1.0
Release Date: August 2025
Description
Graph-CoT-30k is a large-scale, high-quality instruction tuning dataset designed to enhance the reasoning capabilities of large language models (LLMs) on complex graph-theoretic problems. It contains 30,000 question-answer (QA) pairs, each featuring ultra-long chain-of-thought (CoT) reasoning… See the full description on the dataset page: https://huggingface.co/datasets/HKUST-DSAIL/Graph-R1-RFT-COT-30K.freebase-neo4j-graph
Freebase → Neo4j Graph
A property-graph conversion of the final Freebase RDF dump (English-filtered),
including proper resolution of Freebase's Compound Value Type (CVT) nodes,
ready for import into Neo4j or use as a general-purpose large knowledge graph.
Freebase was a large collaborative knowledge base, discontinued by Google in
2016. This dataset is derived from the last publicly available RDF dump
(freebase-rdf-latest.gz, 1.9B raw triples), filtered to English-language… See the full description on the dataset page: https://huggingface.co/datasets/ksk-1729/freebase-neo4j-graph.omnimcp_graphrag_grounded_answer_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_graphrag_grounded_answer_teaser.omnimcp_graphrag_neo4j_cypher_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_graphrag_neo4j_cypher_teaser.graphinfer
GraphInfer
A benchmark for evaluating an LLM's ability to infer over a graph — to produce an answer that
jointly leverages a node's attributes, its neighbours' attributes, and the edges connecting them.
GraphInfer probes this capability along two axes — Description (what is a region of the graph?)
and Comparison (how do regions of the graph differ?) — over five tasks and six structurally
distinct real-world graphs.
Dataset Summary
GraphInfer contains 42,000… See the full description on the dataset page: https://huggingface.co/datasets/graphinfer/graphinfer.Graph-R1-dataset-level-3
Graph-R1 Dataset Level 3
This dataset contains graph reasoning problems at difficulty level 3.
Files
train_graph_level_3.parquet: Training data for level 3
test_graph_and_math_level_3.parquet: Test data for level 3
Usage
import pandas as pd
# Load training data
train_df = pd.read_parquet('train_graph_level_3.parquet')
# Load test data
test_df = pd.read_parquet('test_graph_and_math_level_3.parquet')
Citation
If you use this… See the full description on the dataset page: https://huggingface.co/datasets/HKUST-DSAIL/Graph-R1-dataset-level-3.Graph-R1-dataset-level-2
Graph-R1 Dataset Level 2
This dataset contains graph reasoning problems at difficulty level 2.
Files
train_graph_level_2.parquet: Training data for level 2
test_graph_and_math_level_2.parquet: Test data for level 2
Usage
import pandas as pd
# Load training data
train_df = pd.read_parquet('train_graph_level_2.parquet')
# Load test data
test_df = pd.read_parquet('test_graph_and_math_level_2.parquet')
Citation
If you use this… See the full description on the dataset page: https://huggingface.co/datasets/HKUST-DSAIL/Graph-R1-dataset-level-2.GraphInstruct-RFT-72Komnimcp_graphrag_hybrid_rrf_rerank_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_graphrag_hybrid_rrf_rerank_teaser.
