datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Graphite_Past_ProblemsGenPoster100K
Dataset Card for GenPoster100K
Dataset Summary
GenPoster-100K is a large-scale dataset for content-aware graphic layout generation introduced in the SEGA paper.
The paper describes it as a high-quality poster dataset with layer-parseable source materials and rich metadata.
This repository provides a Hugging Face datasets loader implementation that reads the source release (BruceW91/GenPoster-100K) and exposes normalized examples with:
poster background image… See the full description on the dataset page: https://huggingface.co/datasets/creative-graphic-design/GenPoster100K.graphkind-universos
Español · Page (ES) · Page (EN) · Paper · Proof · Reproduce
Abstract
We classify the operators under which the color refinement (1-WL)
profile of a graph is invariant, inside a parametric family of
observation universes (f, ι) — f compresses neighborhood counts
and ι is an involution. The engine proposed and verified that the
color refinement partition is complement-invariant (call it T4; the
conjecture mechanism was human-built), verified it on 2,131… See the full description on the dataset page: https://huggingface.co/datasets/Jose-dev/graphkind-universos.neural-graphics-dataset
Neural Graphics Dataset
A compact collection of reference image sequences with accompanying motion, depth, and related rendering data, designed for training, validation, and evaluation of neural graphics models.
It is intended for use with models in the Neural Graphics Model Gym, including:
Neural Super Sampling (NSS)
Neural Frame Rate Upscaling (NFRU)
The dataset is intended as a small, practical example dataset for tutorials and experimentation. For best performance in… See the full description on the dataset page: https://huggingface.co/datasets/Arm/neural-graphics-dataset.graph_ghost_datavisem-tracking-graphs
VISEM-Tracking-graphs - HuggingFace Repository
This HuggingFace repository contains the pre-generated graphs for the sperm video dataset called VISEM-Tracking (https://huggingface.co/papers/2212.02842) . The graphs represent spatial and temporal relationships between sperm in a video. Spatial edges connect sperms within the same frame, while temporal edges connect sperms across different frames.
The graphs have been generated with varying spatial threshold values: 0.1, 0.2, 0.3, 0.4… See the full description on the dataset page: https://huggingface.co/datasets/SimulaMet-HOST/visem-tracking-graphs.GQA-Scene-Graph
Dataset Card for GQA-35k
The GQA (Visual Reasoning in the Real World) dataset is a large-scale visual question answering dataset that includes scene graph annotations for each image.
This is a FiftyOne dataset with 35000 samples.
Note: This is a 35,000 sample subset which does not contain questions, only the scene graph annotations as detection-level attributes.
You can find the recipe notebook for creating the dataset here
Installation
If you haven't already… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/GQA-Scene-Graph.graphmemix-benchmarks
GraphMemix Benchmarks
Unified multimodal memory benchmark bundles used by
GraphMemix
(arXiv:2608.26983) — four long-term
personalized memory benchmarks with their raw media assets, packaged together
for reproducible evaluation.
Benchmark
Questions
Memories
Track
Upstream license
ATM-Bench (default + hard)
1,044
11,034
memory QA over one multimodal archive
MIT
Mem-Gallery
1,711
7,944
multimodal gallery memory QA
MIT
MemEye
1,855
3,392
comics-derived memory QA… See the full description on the dataset page: https://huggingface.co/datasets/oking0197/graphmemix-benchmarks.graphwalks
GraphWalks: a multi hop reasoning long context benchmark
In Graphwalks, the model is given a graph represented by its edge list and asked to perform an operation.
Example prompt:
You will be given a graph as a list of directed edges. All nodes are at least degree 1.
You will also get a description of an operation to perform on the graph.
Your job is to execute the operation on the graph and return the set of nodes that the operation results in.
If asked for a breadth-first… See the full description on the dataset page: https://huggingface.co/datasets/openai/graphwalks.PKU-PosterLayout
Dataset Card for PKU-PosterLayout
Dataset Summary
PKU-PosterLayout is a content-aware visual-textual poster layout benchmark released with PosterLayout: A New Benchmark and Approach for Content-aware Visual-Textual Presentation Layout. The paper defines the task as arranging predefined text, logo, and underlay elements on a non-empty poster canvas while considering both inter-element and inter-layer relationships. The original benchmark contains 9,974… See the full description on the dataset page: https://huggingface.co/datasets/creative-graphic-design/PKU-PosterLayout.Mintaka_Graph_Features_T5-xl-ssm
Dataset Card for "Mintaka_Graph_Features_T5-xl-ssm"
More Information needed
PubLayNet
Dataset Card for PubLayNet
Dataset Summary
PubLayNet is a large document layout analysis dataset built by automatically matching XML representations and PDF content from more than one million PubMed Central Open Access articles. It contains more than 360,000 document images with COCO-style annotations for common layout elements such as text, title, list, table, and figure regions.
Supported Tasks and Leaderboards
The dataset supports document… See the full description on the dataset page: https://huggingface.co/datasets/creative-graphic-design/PubLayNet.stochastic_block_model_graphsCreativePSD
Dataset Card for CreativePSD
Dataset Summary
CreativePSD is the PSD-derived graphic design dataset released with PSDesigner. Each example is a poster archive containing PSD tree text, structured layer metadata, tool-call trajectories, source image resources, and stepwise rendered images.
This loader keeps the contents of each poster_*.zip archive: all metadata text/JSON files, all raw_resource images, all rendering_imgs images, and a manifest of every member in… See the full description on the dataset page: https://huggingface.co/datasets/creative-graphic-design/CreativePSD.graphical-bootstrap-correlator-dataset
Graphical Bootstrap Correlator Dataset
This dataset contains large-scale graph-structured data arising from high-order perturbative computations of four-point correlators in planar $\mathcal{N}=4$ super Yang--Mills theory.
The data consists of denominator graphs (d-graphs) appearing in the graphical bootstrap formulation of correlators. Each graph is associated with a binary label indicating whether it contributes to the correlator at a given perturbative order.
The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/Gabriele-dian/graphical-bootstrap-correlator-dataset.Compact_OpenAIRE_citation_graph
📚 Compact OpenAIRE Citation Graph
Based on OpenAIRE Graph v11.1.1 (source on Zenodo).
The complete OpenAIRE citation graph, distilled into a handful of compact, analysis-ready files — the full scholarly citation network of the open-science ecosystem, small enough to actually work with.
Citation graphs at this scale are usually locked behind multi-terabyte dumps and heavyweight infrastructure. This dataset makes the entire OpenAIRE citation network loadable… See the full description on the dataset page: https://huggingface.co/datasets/Zmeos/Compact_OpenAIRE_citation_graph.graph-pannuke
Graph-PanNuke: A Cell-Graph Dataset for Nucleus Classification from PanNuke
Graph-PanNuke is a node-level classification dataset derived from the PanNuke pan-cancer histology dataset. We use all slides at 40× magnification. Each tissue patch is converted into a cell-graph where nodes represent detected cell nuclei and edges encode spatial proximity. The task is predicting the cell type of each nucleus across 5 classes. Note that node features describe cell morphology, texture… See the full description on the dataset page: https://huggingface.co/datasets/ogutsevda/graph-pannuke.Graph200K
VisualCloze: A Universal Image Generation Framework via Visual In-Context Learning
[Paper] [Project Page] [Github]
[🤗 Online Demo]
[🤗 Full Model Card (Diffusers)] [🤗 LoRA Model Card (Diffusers)]
Graph200k is a large-scale dataset containing a wide range of distinct tasks of image generation. If you find Graph200k is helpful, please consider to star ⭐ the Github Repo. Thanks!
📰 News
[2025-5-15] 🤗🤗🤗 VisualCloze has been merged into the… See the full description on the dataset page: https://huggingface.co/datasets/VisualCloze/Graph200K.Rico
Dataset Card for Rico
Dataset Summary
Rico is a mobile app UI dataset for building data-driven design applications. The original dataset mines Android apps at runtime and exposes visual, textual, structural, and interactive design properties from more than 9.3k apps across 27 categories and more than 66k unique UI screens. This packaging provides metadata, screenshots, view hierarchies, and semantic annotations as separate configs.
Supported Tasks… See the full description on the dataset page: https://huggingface.co/datasets/creative-graphic-design/Rico.Text-Attributed-Graphs
Overview
This dataset covers the encoder embeddings and prediction results of LLMs of paper 'Model Generalization on Text Attribute Graphs: Principles with Lagre Language Models', Haoyu Wang, Shikun Liu, Rongzhe Wei, Pan Li.
Dataset Description
The dataset structure should be organized as follows:
/dataset/
│── [dataset_name]/
│ │── processed_data.pt # Contains labels and graph information
│ │── [encoder]_x.pt # Features extracted by different encoders
│… See the full description on the dataset page: https://huggingface.co/datasets/Graph-COM/Text-Attributed-Graphs.MUTAG
Dataset Card for MUTAG
Dataset Summary
The MUTAG dataset is 'a collection of nitroaromatic compounds and the goal is to predict their mutagenicity on Salmonella typhimurium'.
Supported Tasks and Leaderboards
MUTAG should be used for molecular property prediction (aiming to predict whether molecules have a mutagenic effect on a given bacterium or not), a binary classification task. The score used is accuracy, using a 10-fold cross-validation.
External… See the full description on the dataset page: https://huggingface.co/datasets/graphs-datasets/MUTAG.genvsr-video-benchmarksHaystackCraft@article{li2025haystack,
title={Haystack Engineering: Context Engineering for Heterogeneous and Agentic Long-Context Evaluation},
author={Mufei Li and Dongqi Fu and Limei Wang and Si Zhang and Hanqing Zeng and Kaan Sancak and Ruizhong Qiu and Haoyu Wang and Xiaoxin He and Xavier Bresson and Yinglong Xia and Chonglin Sun and Pan Li},
journal={arXiv preprint arXiv:2510.07414},
year={2025}
}
refusal-lens-graphsGraphRAG-Bench
GraphRAG-Bench : A Comprehensive Benchmark for Evaluating Graph Retrieval-Augmented Generation Models
🎉News •
📖About •
🏆Leaderboards •
🧩Task Examples
🔧Getting Started •
📬Contact •
📝Citation
This repository is for the GraphRAG-Bench project, a comprehensive benchmark for evaluating Graph Retrieval-Augmented Generation models.
🎉 News
[2025-05-25] We release GraphRAG-Bench, the benchmark for evaluating GraphRAG… See the full description on the dataset page: https://huggingface.co/datasets/GraphRAG-Bench/GraphRAG-Bench.GBC10M
Graph-based captioning (GBC) is a new image annotation paradigm that combines the strengths of long captions, region captions, and scene graphs
GBC interconnects region captions to create a unified description akin to a long caption, while also providing structural information similar to scene graphs.
** The associated data point can be found at demo/water_tower.json
Description and data format
The GBC10M dataset, derived from the original images in CC12M, is… See the full description on the dataset page: https://huggingface.co/datasets/graph-based-captions/GBC10M.GraphRL_2room_50-99_0310Graph-Algorithmspcod_graph_idsgraphical-bootstrap-correlator-dataset
Graphical Bootstrap Correlator Dataset
This dataset contains large-scale graph-structured data arising from high-order perturbative computations of four-point correlators in planar $\mathcal{N}=4$ super Yang--Mills theory.
The data consists of denominator graphs (d-graphs) appearing in the graphical bootstrap formulation of correlators. Each graph is associated with a binary label indicating whether it contributes to the correlator at a given perturbative order.
The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/anonymous314/graphical-bootstrap-correlator-dataset.
