datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MUTAG
Dataset Card for MUTAG
Dataset Summary
The MUTAG dataset is 'a collection of nitroaromatic compounds and the goal is to predict their mutagenicity on Salmonella typhimurium'.
Supported Tasks and Leaderboards
MUTAG should be used for molecular property prediction (aiming to predict whether molecules have a mutagenic effect on a given bacterium or not), a binary classification task. The score used is accuracy, using a 10-fold cross-validation.
External… See the full description on the dataset page: https://huggingface.co/datasets/graphs-datasets/MUTAG.GraphRAG-Bench
GraphRAG-Bench : A Comprehensive Benchmark for Evaluating Graph Retrieval-Augmented Generation Models
🎉News •
📖About •
🏆Leaderboards •
🧩Task Examples
🔧Getting Started •
📬Contact •
📝Citation
This repository is for the GraphRAG-Bench project, a comprehensive benchmark for evaluating Graph Retrieval-Augmented Generation models.
🎉 News
[2025-05-25] We release GraphRAG-Bench, the benchmark for evaluating GraphRAG… See the full description on the dataset page: https://huggingface.co/datasets/GraphRAG-Bench/GraphRAG-Bench.GraphQAhuatuo_knowledge_graph_qa
Dataset Card for Huatuo_knowledge_graph_qa
Dataset Summary
We built this QA dataset based on the medical knowledge map, with a total of 798,444 pieces of data, in which the questions are constructed by means of templates, and the answers are the contents of the entries in the knowledge map.
Dataset Creation
Source Data… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/huatuo_knowledge_graph_qa.MTID
TurnGate: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue
Overview
TurnGate is a response-aware defense mechanism designed to detect and mitigate hidden malicious intent in multi-turn dialogue systems. Defending state-of-the-art multi-turn malicious attacks like CKA-Agent.
MTID Dataset
We include the MTID (Multi-Turn Intent Dataset) here. This dataset contains a collection of… See the full description on the dataset page: https://huggingface.co/datasets/Graph-COM/MTID.GraphRAG
Local GraphRAG Research Artifacts
This repository contains the data and output artifacts produced for my master's
thesis research on knowledge graph construction in local GraphRAG systems. It
is an evidence package for the study, rather than the code used to run the
experiment.
The collection includes the study corpus, model extraction outputs, final
knowledge graphs, retrieved evidence, generated answers, and evaluation results.
Together, these files make it possible to inspect… See the full description on the dataset page: https://huggingface.co/datasets/boblaros/GraphRAG.szl-estate-graph
SZL Estate Graph v1
This is a deterministic, local publication bundle for SZLHOLDINGS/szl-estate-graph.
It turns the two receipted SZL Constellation topology documents into a typed,
trinity-connected graph suitable for graph-learning and drift comparison.
Publication target: SZLHOLDINGS/szl-estate-graph. Hub publication and its
immutable revision are provider evidence separate from this content bundle;
neither publication nor download implies model training, admission, or… See the full description on the dataset page: https://huggingface.co/datasets/SZLHOLDINGS/szl-estate-graph.trajectory-graph-monitoring
Trajectory Graph Monitoring v0.2; Research Release
Prospective boundary prediction from cumulative typed trajectory transitionsR.J. Sabouhi · Symbolic Suite · August 2026
What this is
Trajectory Graph Monitoring is a prospective runtime-evaluation benchmark. It asks whether structural information available in an incomplete execution trajectory can improve prediction of a boundary violation that occurs later. Every evaluated prefix ends before the violating action… See the full description on the dataset page: https://huggingface.co/datasets/rjsabouhi/trajectory-graph-monitoring.Nemotron-Problem-Graph-v2graphql-nplusone-trajectories
Graphql Nplusone Trajectories
Rights & intended use: legacy public research corpus / portfolio
artifact. Hosted frontier-model outputs are research-only inputs under
project policy (synthetic-factory#161):
intended_use: research_only, project_training_policy: blocked. Not
training data for any model-weight update. Machine-readable record:
rights.json.
Release status: The raw, uncurated payload is now published under
data/raw/. It is available for inspection and… See the full description on the dataset page: https://huggingface.co/datasets/rmems/graphql-nplusone-trajectories.bottleneck-oracle-graphsPROTEINS
Dataset Card for PROTEINS
Dataset Summary
The PROTEINS dataset is a medium molecular property prediction dataset.
Supported Tasks and Leaderboards
PROTEINS should be used for molecular property prediction (aiming to predict whether molecules are enzymes or not), a binary classification task. The score used is accuracy, using a 10-fold cross-validation.
External Use
PyGeometric
To load in PyGeometric, do the following:
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/graphs-datasets/PROTEINS.qualora-workforce-skills-graph
Qualora Workforce Skills Graph (Representative Sample)
Rights-clean, provenance-tracked vocational learning data, rebuilt from roughly $2B of U.S. Department of Labor funded open courseware into a labeled skills graph: cleaned courses and lessons, Bloom-tagged assessment items with answer rationales and learning objectives, and a content-grounded course to skill to career graph with salary context. Built for post-training and evaluation, not pretraining bulk.
This repository is… See the full description on the dataset page: https://huggingface.co/datasets/qualora-data-labs/qualora-workforce-skills-graph.GraphInstruct-Testfin-cfa-graphgen
Fin-CFA-GraphGen: 785K Knowledge-Guided Financial QA Examples
Fin-CFA-GraphGen is a large-scale English dataset for financial instruction tuning, financial question answering, and domain-specific language-model post-training. It contains 785,149 synthetic question–answer examples generated from CFA curriculum and exam-preparation books with the GraphGen knowledge-driven data-generation method.
The dataset and its role in the post-training pipeline are described in Data-Centric… See the full description on the dataset page: https://huggingface.co/datasets/whoisjiji/fin-cfa-graphgen.awesome-graph-engineering
Awesome Graph Engineering Resource Atlas
A versioned collection of research, standards, frameworks, protocols, reliability systems, evaluations, and critiques for graph-structured multi-agent systems and programmable AI-agent organizations.
This dataset mirrors Awesome Graph Engineering. The GitHub JSONL file is canonical; the Hub exposes the same records through Dataset Viewer, direct downloads, datasets, and pandas.
Working definition
Graph engineering is the… See the full description on the dataset page: https://huggingface.co/datasets/cy0307/awesome-graph-engineering.graph-captioning-train-onlyCATH4.2Graph2Counsel
Dataset Card
Paper: Graph2Counsel: Clinically Grounded Synthetic Counseling Dialogue Generation from Client Psychological Graphs
Language(s) (NLP): English
license: cc-by-sa-4.0
Dataset Summary
Graph2Counsel is a synthetic counseling session dataset generated from Client Psychological Graphs (CPGs). The dataset provides the CPG, generated diverse client profiles and dialogues from this CPG as well as counselor strategies corresponding to each CPG collected from the real… See the full description on the dataset page: https://huggingface.co/datasets/UKPLab/Graph2Counsel.GraphMasterGraphRAG-Bench
GraphRAG-Bench : A Comprehensive Benchmark for Evaluating Graph Retrieval-Augmented Generation Models
🎉News •
📖About •
🏆Leaderboards •
🧩Task Examples
🔧Getting Started •
📬Contact •
📝Citation
This repository is for the GraphRAG-Bench project, a comprehensive benchmark for evaluating Graph Retrieval-Augmented Generation models.
🎉 News
[2025-05-25] We release GraphRAG-Bench, the benchmark for evaluating GraphRAG models.… See the full description on the dataset page: https://huggingface.co/datasets/abhisi60/GraphRAG-Bench.federal-register-live-graphrag-research-20260810
Federal Register live GraphRAG (research)
Local LCR-071 live pipeline output for the 2026-08-10 cutoff (11,784 documents,
CUDA thenlper/gte-small). This Hub copy is a research snapshot.
It is not a current-bundle and does not replace
justicedao/ipfs_federal_register. LCR-084 remains open. Official Federal
Register publications remain the authority.
Hub git directories may contain at most 10,000 files. Document bodies beyond
that cap are stored under corpus/bodies-part2/ rather… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/federal-register-live-graphrag-research-20260810.GraphRAG-Bench
GraphRAG-Bench
This repository hosts the official website for the GraphRAG-Bench project, a comprehensive benchmark for evaluating Graph Retrieval-Augmented Generation models.
Website Overview
🎉 News
[2025-05-25] We release GraphRAG-Bench, the benchmark for evaluating GraphRAG models.
[2025-05-14] We release the GraphRAG-Bench dataset.
[2025-01-21] We release the GraphRAG survey.
📖 About
Introduces Graph Retrieval-Augmented Generation… See the full description on the dataset page: https://huggingface.co/datasets/wuchuanjie/GraphRAG-Bench.GraphInstructIMDB-BINARY
Dataset Card for IMDB-BINARY (IMDb-B)
Dataset Summary
The IMDb-B dataset is "a movie collaboration dataset that consists of the ego-networks of 1,000 actors/actresses who played roles in movies in IMDB. In each graph, nodes represent actors/actress, and there is an edge between them if they appear in the same movie. These graphs are derived from the Action and Romance genres".
Supported Tasks and Leaderboards
IMDb-B should be used for graph classification… See the full description on the dataset page: https://huggingface.co/datasets/graphs-datasets/IMDB-BINARY.graphext-qa
GQA: Graph Question Answering
Dataset Summary
This dataset is asks models to make use of embedded graph for question answering.
Stats:
train: 57,043
test: 2,890
An exmaple of the dataset is as follows:
{
"id": "mcwq-176119",
"question": "What was executive produced by Scott Spiegel , Boaz Yakin , and Quentin Tarantino , executive produced by My Best Friend's Birthday 's editor and star , and edited by George Folsey",
"answers": [
"Hostel: Part II"]… See the full description on the dataset page: https://huggingface.co/datasets/drt/graphext-qa.knowledge-graph-risk-engine-20260828-dataset
Knowledge Graph Risk Engine Synthetic Dataset
Summary
This dataset contains 14 training examples and 4
held-out examples for Risk teams need relationship-level explanations instead of opaque entity scores.
Every record is synthetic and includes:
input: query, event, or feature description
label: expected class, route, relation, or evidence category
context: synthetic supporting context
source: fictional source identifier
variant: generation pattern
synthetic:… See the full description on the dataset page: https://huggingface.co/datasets/RKB109/knowledge-graph-risk-engine-20260828-dataset.graphtrain91GraphInstruct-RFT-72KGraphSilo-Test
