datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
physical-ai-bench-generation
Physical AI Bench - Generation
Paper | Code
Dataset Description
The PAI-Bench is a benchmark to measure the progress of world models quantitatively.
The predict task contains a list of 1044 samples of text prompts, conditioning images, and qa pairs, covering Physical AI target domains including autonomous vehicle (AV) driving, robotics, industry (smart space), physics, human, and common sense. All the questions are binary questions, and the answer is either Yes or No. Our… See the full description on the dataset page: https://huggingface.co/datasets/shi-labs/physical-ai-bench-generation.GMAI-VL-5.5M
GMAI-VL-5.5M Dataset
GMAI-VL-5.5M is a comprehensive, large-scale medical General Medical AI Vision-Language (GMAI-VL) dataset built specifically for training multimodal foundation models in the medical domain. It contains an extraordinary scale of high-quality instructions encompassing over 5.5 million multimodal question-answering pairs, carefully constructed based on hundreds of medical classification, segmentation, and detection datasets.
This repository… See the full description on the dataset page: https://huggingface.co/datasets/General-Medical-AI/GMAI-VL-5.5M.GMAI-Reasoning10K
GMAI-Reasoning10K
Medical Reasoning dataset used in GMAI-VL-R1
Data description
GMAI-Reasoning10K is a high-quality medical image reasoning dataset containing 10,000 carefully selected samples. The data was collected from 95 medical datasets from reliable sources such as Kaggle, GrandChallenge, and Open-Release, covering 12 imaging modalities including X-ray, CT, and MRI.
Data preprocessing followed the standardization methods from SAMed-20M: 3D data (CT/MRI) had… See the full description on the dataset page: https://huggingface.co/datasets/General-Medical-AI/GMAI-Reasoning10K.KnowGen-Bench
KnowGen Benchmark
Project Page | Paper | Code
This repository contains the KnowGen benchmark data for Gen-Searcher: Reinforcing Agentic Search for Image Generation.
👀 Intro
We introduce Gen-Searcher, as the first attempt to train a multimodal deep research agent for image generation that requires complex real-world knowledge. Gen-Searcher can search the web, browse evidence, reason over multiple sources, and search visual referencesbefore generation, enabling… See the full description on the dataset page: https://huggingface.co/datasets/GenSearcher/KnowGen-Bench.OS-Genesis-mobile-data🫶 If you are interested in our work or find this data helpful, please consider using the following .bib when referencing our paper:
@article{sun2024genesis,
title={OS-Genesis: Automating GUI Agent Trajectory Construction via Reverse Task Synthesis},
author={Sun, Qiushi and Cheng, Kanzhi and Ding, Zichen and Jin, Chuanyang and Wang, Yian and Xu, Fangzhi and Wu, Zhenyu and Jia, Chengyou and Chen, Liheng and Liu, Zhoumianze and others},
journal={arXiv preprint arXiv:2412.19723}… See the full description on the dataset page: https://huggingface.co/datasets/OS-Copilot/OS-Genesis-mobile-data.Valen-Eval-General-5k
Valen-Eval-General-5k
GitHub · 中文 README · Preview model · Technical notes
✨ Introduction
Valen brings visual perception to System One decision-making: text, images or video in, candidate probabilities out. This dataset provides 5,000 image-based decision records for held-out evaluation, spanning visual question answering, interfaces, games and documents.
Each record contains one decision question, a target probability distribution, local image… See the full description on the dataset page: https://huggingface.co/datasets/Valen-Team/Valen-Eval-General-5k.synthetic-manuscript-generator
Synthetic Manuscript Generator
Synthetic Indic manuscript folios (paper + palm-leaf backgrounds) for OCR
training. Three scripts are produced as separate subsets/configs:
devanagari — 100 folios (85/10/5)
modi — 100 folios (85/10/5)
sharada — 100 folios (85/10/5)
Layout
Each subset is structured as a Hugging Face imagefolder:
<subset>/
train/
0000.png 0000.md metadata.jsonl
...
validation/
...
test/
...
metadata.jsonl rows look like:… See the full description on the dataset page: https://huggingface.co/datasets/Sampada22/synthetic-manuscript-generator.GENIUS
🧠 GENIUS
Generative Fluid Intelligence Evaluation Suite
[Paper] [Code] [Blog] [Dataset]
Leaderboard
GENIUS evaluates every generated image on three complementary axes, each scored 0 (fail), 1 (partial), or 2 (perfect):
Rule Compliance (RC): follows the newly defined rule, grounded by expert-written evaluation hints.
Visual Consistency (VC): preserves required identities, objects, and contextual visual attributes.… See the full description on the dataset page: https://huggingface.co/datasets/HankYang428/GENIUS.image-generation-flux1-schnell
Dataset Card for Dataset Name
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/davidberenstein1957/image-generation-flux1-schnell.fashion_model_generatordataset-viber-image-generation-preference-inference-endpoints-battle-flux
Dataset Card for Dataset Name
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/davidberenstein1957/dataset-viber-image-generation-preference-inference-endpoints-battle-flux.SCS_data
[NeurIPS 2025] Enhancing the Outcome Reward-based RL Training of MLLMs with Self-Consistency Sampling
A simple, general sampling method for RLVR with multi-choice dataset to solve unfaithful reasoning phenomenon!
SCS Resouces
📖 Paper | 🤗 Dataset | 💻 Code
🔔News
🔥[2025-11-9] Release the eval codes! 🚀
🔥[2025-10-13] Release the dataset the codes! 🚀
🔥[2025-9-17] Our SCS paper is accepted by NeurIPS 2025! 🚀
To-do
Release the eval codes… See the full description on the dataset page: https://huggingface.co/datasets/GenuineWWD/SCS_data.nft-females-generated-cli-test-3-clone
nft females generated cli test 3
Automated NFT generation report
Number of characters: 10
Model: midorimae-characters-female
Guidance: 6.0
LORA scale: 0.95
Total time taken: 0h 0m 51s
Average time per prompt: 2.39 seconds
Average time per image: 2.74 seconds
Average time per image: 5.13 seconds
gpt-generated-image-collection
GPT-Generated Image Collection
A growing collection of contributor-provided GPT-generated PNG images. The first release contains 216 original PNG files. Exact generation-model versions have not been independently verified; the generation_model field is therefore null rather than an inferred model name.
For the Chinese version of this card, see README.zh-CN.md.
Layout
README.md English dataset card
README.zh-CN.md Chinese… See the full description on the dataset page: https://huggingface.co/datasets/winrisef/gpt-generated-image-collection.gensin-instruct-read-to-traingenerated-stanford-dogs
Generated Stanford Dogs Dataset
Description
This repository contains the dataset used for the generative-data-augmentation project. The dataset is organized as follows:
Dataset Structure
analysis/: This directory contains analysis related to the dataset.
metadata/: This directory contains the list of file path used for the Synthetic (Noisy) and Synthetic (Clean) datasets.
synthetic/: This directory contains the image files. Each folder represents a class.… See the full description on the dataset page: https://huggingface.co/datasets/czl/generated-stanford-dogs.celeba-lora-generatedgenerated-imagenette
Generated Imagenette Dataset
Description
This repository contains the dataset used for the generative-data-augmentation project. The dataset is organized as follows:
Dataset Structure
analysis/: This directory contains analysis related to the dataset.
metadata/: This directory contains the list of file path used for the Synthetic (Noisy) and Synthetic (Clean) datasets.
synthetic/: This directory contains the image files. Each folder represents a class.… See the full description on the dataset page: https://huggingface.co/datasets/czl/generated-imagenette.genesis-magazinegenerated-imagewoof
Generated Imagewoof Dataset
Description
This repository contains the dataset used for the generative-data-augmentation project. The dataset is organized as follows:
Dataset Structure
analysis/: This directory contains analysis related to the dataset.
metadata/: This directory contains the list of file path used for the Synthetic (Noisy) and Synthetic (Clean) datasets.
synthetic/: This directory contains the image files. Each folder represents a class.… See the full description on the dataset page: https://huggingface.co/datasets/czl/generated-imagewoof.
