datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
python-image-copilot-training-using-import-knowledge-graphs
Python Copilot Image Training using Import Knowledge Graphs
This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset.
Details
Each row contains a png file in the dbytes column.
Rows: 216642
Size: 211.2 GB
Data type: png
Format: Knowledge graph using NetworkX with alpaca text box
Schema
The png is in the dbytes column:
{
"dbytes": "binary"… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-image-copilot-training-using-import-knowledge-graphs.python-image-copilot-training-using-class-knowledge-graphs
Python Copilot Image Training using Class Knowledge Graphs
This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset.
Details
Each row contains a png file in the dbytes column.
Rows: 312277
Size: 304.3 GB
Data type: png
Format: Knowledge graph using NetworkX with alpaca text box
Schema
The png is in the dbytes column:
{
"dbytes": "binary"… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-image-copilot-training-using-class-knowledge-graphs.python-audio-copilot-training-using-class-knowledge-graphs
Python Copilot Audio Training using Class with Knowledge Graphs
This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset.
Details
Each class method has a question and answer mp3 where one voice reads the question and another voice reads the answer. Both mp3s are stored in the parquet dbytes column and the associated source code file_path identifier.
Rows:… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-audio-copilot-training-using-class-knowledge-graphs.python-audio-copilot-training-using-function-knowledge-graphs
Python Copilot Audio Training using Global Functions with Knowledge Graphs
This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset.
Details
Each global function has a question and answer mp3 where one voice reads the question and another voice reads the answer. Both mp3s are stored in the parquet dbytes column and the associated source code file_path… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-audio-copilot-training-using-function-knowledge-graphs.python-audio-copilot-training-using-import-knowledge-graphs
Python Copilot Audio Training using Imports with Knowledge Graphs
This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset.
Details
Each imported module for each unique class in each module file has a question and answer mp3 where one voice reads the question and another voice reads the answer. Both mp3s are stored in the parquet dbytes column and the… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-audio-copilot-training-using-import-knowledge-graphs.python-audio-copilot-training-using-inheritance-knowledge-graphs
Python Copilot Audio Training using Inheritance and Polymorphism Knowledge Graphs
This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset.
Details
Each base class for each unique class in each module file has a question and answer mp3 where one voice reads the question and another voice reads the answer. Both mp3s are stored in the parquet dbytes column and… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-audio-copilot-training-using-inheritance-knowledge-graphs.python-audio-copilot-training-using-class-knowledge-graphs-2024-01-27
Python Copilot Audio Training using Class with Knowledge Graphs
This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset.
Details
Each class method has a question and answer mp3 where one voice reads the question and another voice reads the answer. Both mp3s are stored in the parquet dbytes column and the associated source code file_path identifier.
Rows:… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-audio-copilot-training-using-class-knowledge-graphs-2024-01-27.python-image-copilot-training-using-inheritance-knowledge-graphs
Python Copilot Image Training using Inheritance and Polymorphism Knowledge Graphs
This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset.
Details
Each row contains a png file in the dbytes column.
Rows: 259017
Size: 135.2 GB
Data type: png
Format: Knowledge graph using NetworkX with alpaca text box
Schema
The png is in the dbytes column:
{… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-image-copilot-training-using-inheritance-knowledge-graphs.python-image-copilot-training-using-class-knowledge-graphs-2024-01-27
Python Copilot Image Training using Class Knowledge Graphs
This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset.
Details
Each row contains a png file in the dbytes column.
Rows: 312836
Size: 294.1 GB
Data type: png
Format: Knowledge graph using NetworkX with alpaca text box
Schema
The png is in the dbytes column:
{
"dbytes": "binary"… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-image-copilot-training-using-class-knowledge-graphs-2024-01-27.python-image-copilot-training-using-function-knowledge-graphs
Python Copilot Image Training using Function Knowledge Graphs
This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset.
Details
Each row contains a png file in the dbytes column.
Rows: 134357
Size: 130.5 GB
Data type: png
Format: Knowledge graph using NetworkX with alpaca text box
Schema
The png is in the dbytes column:
{
"dbytes": "binary"… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-image-copilot-training-using-function-knowledge-graphs.bottleneck-oracle-graphsclinical-trials-patient-graphstravel-fraud-graphs
TravelFraudBench (TFG)
The first publicly available labeled heterogeneous graph benchmark for GNN-based fraud ring detection in travel networks.
Paper
Dataset Structure
This dataset contains heterogeneous graph data split into 20 named configurations — one per node type and one per edge type — each with small, medium, and large splits.
Loading a specific node or edge table
from datasets import load_dataset
# Load user nodes (medium scale)
users =… See the full description on the dataset page: https://huggingface.co/datasets/bsajja7/travel-fraud-graphs.star_graphs_paths20_len20_nodes1000cbeeg-evidence-graphs
CBEEG — Compute-Budgeted Exploitability Evidence Graphs (derived artifacts)
Reproducible derived artifacts for the paper Compute-Budgeted Exploitability
Evidence Graphs for Prospective Vulnerability Triage (Alpay & Alpay). The paper
frames prospective CVE triage as a leakage-safe, compute-budgeted evidence
selection problem: for each CVE we admit only public evidence visible by a fixed
decision time, select a few documents under a budget, and attach an auditable
evidence… See the full description on the dataset page: https://huggingface.co/datasets/Lightcap/cbeeg-evidence-graphs.clinical-trials-eligibility-graphs-rerank
Clinical Trial Eligibility Graphs (rerank candidate set)
Parsed eligibility criteria for 13,229 ClinicalTrials.gov trials — the complete
candidate set retrieved across TREC Clinical Trials 2021–23 by two first-stage
retrievers. Each trial's inclusion and exclusion criteria are represented as
entities and typed relations.
This exists so the
CrossGAT eligibility model
can be run without an LLM. Generating these graphs took ~28 GPU-hours on a
20B-parameter model; this dataset is… See the full description on the dataset page: https://huggingface.co/datasets/2001jdev/clinical-trials-eligibility-graphs-rerank.star_graphs_paths25_len5_nodes1000graphcontainer-graphs
GraphContainer Graph Artifacts
Overview
This repository contains preconstructed graph artifacts released with GraphContainer: A Unified Platform for Comparing and Debugging Graph RAG Methods.
Links:
Paper: https://arxiv.org/abs/2607.19362
Hugging Face paper page: https://huggingface.co/papers/2607.19362
Code: https://github.com/asmath472/GraphContainer
YouTube demo: https://youtu.be/O02eNJLwkU0
The Hugging Face datasets interface loads the artifact catalog in… See the full description on the dataset page: https://huggingface.co/datasets/hchaejeong/graphcontainer-graphs.neural-mis-training-graphs
Synthetic Maximum Independent Set training graphs
500 synthetic Erdos-Renyi graphs (20-1,972 nodes, density 0.05-0.5), each labelled by
20 restarts of randomized-greedy construction plus (1,2)-exchange local search. Built
to train Bauxitiego/neural-mis, a GCN
evaluated against QOBLIB's Maximum Independent
Set benchmark — code, evaluation, and honest results (including a documented failure)
at github.com/Bauxitiego/neural-mis.
Generated, not collected: free, unlimited… See the full description on the dataset page: https://huggingface.co/datasets/Bauxitiego/neural-mis-training-graphs.vg_coco_overlap_for_graphormer_processed_amr_graphsvg-captions-graphs-processed-image-graphsZ3-Verified-Reasoning-Graphs
Z3-Verified Constraint Reasoning Dataset
5k Baseline · Production-Ready · Zero Label Noise
The Problem This Solves
Most synthetic reasoning datasets only show the "happy path". Real reasoning requires knowing when to backtrack.
Open-source LLMs hallucinate on constraint satisfaction problems because they are trained on fluent-sounding but logically inconsistent traces. This dataset is different:
❌ No LLM-generated reasoning — zero hallucinations, zero label noise
✅… See the full description on the dataset page: https://huggingface.co/datasets/nagygabor/Z3-Verified-Reasoning-Graphs.g-retriever-scene-graphsdiscrete_prompting_scene_graphscayley-graphs-384-to-447
Cayley Graphs — orders 384–447
This dataset contains Cayley graphs of finite groups: one row per group, covering groups whose order is between 384 and 447. Each graph is the Cayley graph built from the group's minimal generating set (see Provenance below).
Rows (groups): 21,814
Group orders covered: 384–447 (64 distinct orders)
Task: binary graph classification (default label: IsMonolithic).
About the CayleyNet collection
This dataset is part of a census of 131… See the full description on the dataset page: https://huggingface.co/datasets/Engrima18/cayley-graphs-384-to-447.vg_coco_overlap_for_graphormer_processed_filtered_image_graphscayley-graphs-513-to-576
Cayley Graphs — orders 513–576
This dataset contains Cayley graphs of finite groups: one row per group, covering groups whose order is between 513 and 576. Each graph is the Cayley graph built from the group's minimal generating set (see Provenance below).
Rows (groups): 9,782
Group orders covered: 513–576 (64 distinct orders)
Task: binary graph classification (default label: IsMonolithic).
About the CayleyNet collection
This dataset is part of a census of 131… See the full description on the dataset page: https://huggingface.co/datasets/Engrima18/cayley-graphs-513-to-576.cayley-graphs-641-to-767
Cayley Graphs — orders 641–767
This dataset contains Cayley graphs of finite groups: one row per group, covering groups whose order is between 641 and 767. Each graph is the Cayley graph built from the group's minimal generating set (see Provenance below).
Rows (groups): 6,160
Group orders covered: 641–767 (127 distinct orders)
Task: binary graph classification (default label: IsMonolithic).
About the CayleyNet collection
This dataset is part of a census of 131… See the full description on the dataset page: https://huggingface.co/datasets/Engrima18/cayley-graphs-641-to-767.cayley-graphs-577-to-639
Cayley Graphs — orders 577–639
This dataset contains Cayley graphs of finite groups: one row per group, covering groups whose order is between 577 and 639. Each graph is the Cayley graph built from the group's minimal generating set (see Provenance below).
Rows (groups): 1,119
Group orders covered: 577–639 (63 distinct orders)
Task: binary graph classification (default label: IsMonolithic).
About the CayleyNet collection
This dataset is part of a census of 131… See the full description on the dataset page: https://huggingface.co/datasets/Engrima18/cayley-graphs-577-to-639.cayley-graphs-all-wo-256-or-512
Cayley Graphs — all orders (excl. 256, 512)
This dataset contains Cayley graphs of finite groups: one row per group, covering all available group orders, excluding orders 256, 512. Each graph is the Cayley graph built from the group's minimal generating set (see Provenance below).
Rows (groups): 75,314
Group orders covered: 1–767 (765 distinct orders)
Task: binary graph classification (default label: IsMonolithic).
About the CayleyNet collection
This dataset is… See the full description on the dataset page: https://huggingface.co/datasets/Engrima18/cayley-graphs-all-wo-256-or-512.
